Sora is OpenAI’s text-to-video model, announced in February 2024. The announcement described it as “an AI model that can create realistic and imaginative scenes from text instructions” and stated that all of the example videos on the page “were generated directly by Sora without modification.” The description below reflects that announcement; the product has changed since.

Capabilities described at announcement

  • Clip length and prompt adherence. OpenAI stated that Sora “can generate videos up to a minute long while maintaining visual quality and adherence to the user’s prompt.”
  • Image and video inputs. Beyond text prompts, OpenAI said the model can take an existing still image and animate its contents, and can take an existing video and extend it or fill in missing frames.
  • Diffusion transformer architecture. OpenAI described Sora as a diffusion model that uses a transformer architecture, representing videos and images as “collections of smaller units of data called patches, each of which is akin to a token in GPT,” which it said allowed training across a wider range of durations, resolutions and aspect ratios.

OpenAI also published a list of the model’s weaknesses at announcement, including difficulty simulating the physics of complex scenes, confusion of spatial details such as left and right, and trouble with events unfolding over time.

Access and safety controls at announcement

OpenAI said Sora was “becoming available to red teamers to assess critical areas for harms or risks” — describing those red teamers as domain experts in misinformation, hateful content and bias — and that it was “also granting access to a number of visual artists, designers, and filmmakers to gain feedback.”

On safety, OpenAI said it was building a detection classifier to identify Sora-generated video, that its existing text classifier would reject prompts violating its usage policies, and that image classifiers would review the frames of every generated video before it was shown to the user. It said it planned “to include C2PA metadata in the future if we deploy the model in an OpenAI product” — a stated intention rather than a shipped feature at that time.

Credited team

OpenAI’s announcement credited Bill Peebles and Tim Brooks as research leads and Connor Holmes as systems lead, alongside a list of contributors and separate communications, legal and external engagement groups.

Political media significance

Sora’s announcement placed instruction-following video generation, safety review and content-provenance plans in a single public release. Because generated video is difficult to distinguish from recorded footage, the detection and provenance measures OpenAI described bear directly on how such tools can be used and identified in political contexts.

Sources

  1. 01.

    OpenAI. Sora: Creating video from text. OpenAI's announcement page as captured on 2024-02-20. Source for the one-minute clip length, the statement that all videos on the page were generated without modification, still-image animation and video extension, the patches and diffusion transformer description, the red-teamer and invited-creator access model, the text and image classifiers, the stated plan to include C2PA metadata, and the credited research, systems, communications and legal leads.

Related Entities

research-lead
bill-peebles
Bill Peebles is credited as a Sora research lead on OpenAI's announcement
research-lead
tim-brooks
Tim Brooks is credited as a Sora research lead on OpenAI's announcement