Information processing system

CN122802734APending Publication Date: 2026-09-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610319151.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-03-19
Filing Date
2026-03-16
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

现有的自动化工具普遍无法充分理解和解析用户对目标播放时长、应采用镜头片段等具体要求,无法灵活组合素材并优化镜头节奏与结构布局,从而难以满足用户在短篇视频场景下对内容紧凑性和传播效果的需求

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802734A_ABST
    Figure CN122802734A_ABST
Patent Text Reader

Abstract

The present application provides an information processing system, comprising: a processor; an input means, the input means being configured to input pre-edited video material, image material and audio material to the processor to enable the processor to receive the material; an information input means, the information input means being configured to input information about a video expected to be generated to the processor to enable the processor to receive the information; a prompt generation unit, the prompt generation unit being configured to enable the processor to analyze the input material and information, and generate a prompt for indicating a generative artificial intelligence model to automatically generate a video based on the analysis result, and enable the processor to input the prompt to the generative artificial intelligence model to generate the video based on the prompt; an emotion analysis means, the emotion analysis means being configured to input a user's emotional state to the processor to enable the processor to analyze the user's emotion, and enable the processor to adjust the content of the video according to the emotion analysis result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology disclosed herein relates to an information processing system. Background Technology

[0002] Japanese Patent Application Publication No. 2022-180282 discloses a method for controlling a role-based chatbot executed by at least one processor. The method includes the following steps: receiving a user's speech; adding the user's speech to a prompt word, the prompt word containing instruction statements associated with an explanation of the chatbot's role; encoding the prompt word; and inputting the encoded prompt word into a language model to generate a chatbot response to the user's speech.

[0003] In current technology, video production and editing largely rely on manual operation. Users typically need to use complex non-linear editing software to manually cut, arrange, and adjust video, image, and audio materials. This not only requires high levels of professional skills and a significant time investment, but also makes it difficult for general users lacking video editing experience to efficiently generate video content that meets their desired style and purpose based on abstract creative intentions and emotional needs. Even though some automatic editing tools have emerged in recent years, these tools are mostly based on preset templates or simple rules, with limited ability to understand user input. They struggle to adaptively adjust to the user's specific creative intentions and emotional state, resulting in videos that are inadequate in terms of content matching, emotional expression, and compatibility with the target platform.

[0004] Furthermore, with the rapid popularization of short videos, users often have specific publishing purposes, such as publishing videos of a certain length on short video platforms and hoping to highlight key shots or information within a limited time. Existing automated tools generally cannot fully understand and analyze users' specific requirements regarding target playback length, appropriate shot clips, etc., and cannot flexibly combine materials and optimize shot rhythm and structural layout, thus failing to meet users' needs for concise content and effective dissemination in short video scenarios.

[0005] Furthermore, existing systems typically ignore the user's emotional state during the creation and viewing process, or only support style selection based on simple tags. They lack the ability to intelligently analyze and adjust the feedback of the user's true emotions, and cannot dynamically optimize the generated video content according to the user's emotional changes, resulting in insufficient performance in terms of emotional resonance and personalized experience.

[0006] Therefore, there is an urgent need for a system that can receive various media materials and information about the desired video, automatically generate videos using generative artificial intelligence models, and adaptively adjust video content based on user sentiment analysis results. This would improve the intelligence and personalization of video generation, while lowering the barrier to video production for users, especially enhancing the generation effect and efficiency in short video creation and publishing scenarios. Summary of the Invention

[0007] To address the aforementioned issues, this invention proposes an information processing system. This system includes modules such as a processor, input methods, information input methods, a prompt generation unit, and sentiment analysis methods. Through the collaborative work of these modules, intelligent analysis of user-input materials and information is achieved, and video content is automatically generated and adjusted through a generative artificial intelligence model.

[0008] Specifically, the processor is electrically connected to an input means configured to input unedited video footage, image footage, and audio footage to the processor, enabling the processor to receive and store various types of media footage. The processor is also electrically connected to an information input means configured to input information about the desired video output to the processor. This information may include the video's target style, purpose, target duration, target platform, and user preferences for key content, thereby enabling the processor to obtain an abstract or concrete description of the video generation objective.

[0009] The prompt generation unit is connected to the processor. The prompt generation unit is configured to control the processor to parse the input material and the information, and based on the parsing results, generate prompts to instruct the generative artificial intelligence model to automatically generate video. These prompts can be in natural language or structured form, and at least include instructions regarding material selection, shot composition, pacing, style settings, and duration control. The prompt generation unit is further configured to cause the processor to input the prompts into the generative artificial intelligence model, and based on the model's output, generate corresponding video content, thereby achieving automated generation from raw material to the target video.

[0010] To further achieve adaptive response to the user's emotional state, the emotion analysis method is connected to the processor. This method is configured to acquire input information related to the user's emotions, such as evaluation information, voice content, facial expressions, or other signals reflecting emotions expressed by the user in the interactive interface, and provide this information to the processor for emotion analysis. Based on the emotion analysis results, the processor determines the user's current emotional tendency or desired emotional style, and accordingly adjusts the already generated or being generated video content. This includes adjusting shot selection and order, changing music rhythm and intensity, adding or removing certain emotional scenes, or modifying text descriptions, thereby making the final generated video more in line with the user's emotional needs and aesthetic preferences.

[0011] In one embodiment, to adapt to the requirements of publishing short videos on social media platforms, the processor can parse the necessary playback duration related to the short video and the user-specified shot segments to be used, when the purpose is to publish the short video. Based on this, the prompt generation unit generates prompts containing target duration constraints and mandatory shot constraints, enabling the generative AI model to reasonably arrange mandatory shots and other candidate materials within a limited duration when automatically generating videos, thereby generating compact video content suitable for publishing on short video platforms.

[0012] In another embodiment, the processor can utilize artificial intelligence technology to perform feature extraction and semantic understanding on the input materials and information, including using models such as image recognition, speech recognition, and natural language processing to comprehensively analyze the material content and user intent from a multimodal perspective. The prompt generation unit automatically constructs high-quality prompts based on the analyzed content tags, emotional features, and user needs, and drives the generative artificial intelligence model to output a video solution that highly matches the user intent. By combining the generated results with feedback adjustments using sentiment analysis methods, the system of this invention can achieve high-quality, personalized automatic video generation for different scenarios (including but not limited to family anniversary videos, product promotional videos, and short social media videos) with fewer human-computer interaction steps.

[0013] The "system" refers to a whole consisting of components such as a processor, input means, information input means, prompt generation unit, emotion analysis means, and necessary storage devices and communication interfaces. It is a hardware and software combination used to receive materials and information, generate prompts, call generative artificial intelligence models to automatically generate videos, and make adjustments.

[0014] A "processor" refers to an electronic computing unit used to execute program instructions, parse and process input materials and information, generate prompts, and invoke generative artificial intelligence models. It can be a single central processing unit (CPU), graphics processing unit (GPU), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or any combination thereof.

[0015] "Input means" refers to hardware and / or software components used to input pre-edited video, image, and audio materials into the processor, including but not limited to touch screen interfaces, keyboards, mice, cameras, microphones, file upload interfaces, sensors, or network receiving modules.

[0016] "Information input means" refers to a device or module used to input information about the desired video output to a processor, including but not limited to text input boxes, option buttons, slider controls, voice input interfaces, question-and-answer interactive interfaces, or network interfaces for receiving pre-configured parameters in a graphical user interface.

[0017] "Pre-edit video footage" refers to raw or semi-finished video data that has not yet undergone final editing or compositing and can be used as the basic material for generating the target video. This includes video clips shot by the user, downloaded or imported raw video files, etc.

[0018] "Image material" refers to static image data used in video generation or compositing, including photographs, illustrations, icons, posters, caption images, brand logo images, and other static images that can be overlaid into a video.

[0019] "Audio material" refers to sound data that can be used in combination with video footage, including background music, ambient sound effects, voice narration, dubbing, and other forms of audio files or audio streams.

[0020] "Information about the video to be generated" refers to various parameters and descriptions provided by the user or external system that characterize the features and requirements of the target video, including but not limited to video style, purpose, target duration, target platform, aspect ratio, pacing preference, key subjects, and the emotional atmosphere to be presented.

[0021] The "prompt generation unit" refers to a functional module that controls the processor to parse the input materials and information, and generates prompts based on the parsing results to instruct the generative artificial intelligence model to automatically generate videos. It can be implemented through software programs, microservices, or specific hardware logic.

[0022] "Generation prompts" refer to the instruction information used to input into a generative artificial intelligence model to guide the model to generate video content according to specific requirements. The instruction information can be a natural language description, a structured parameter set, or a combination of both, and at least includes control information regarding material selection, shot arrangement, style setting, duration control, and emotional tendency.

[0023] "Generative AI model" refers to an AI model that can automatically generate video content or video structure scheme corresponding to a given prompt based on trained parameters. This includes, but is not limited to, deep learning-based text-to-video models, multimodal generation models, or other machine learning models with generative capabilities.

[0024] "Emotional analysis methods" refer to modules used to acquire and analyze data related to user emotions in order to determine the user's emotional state or emotional tendency. These include hardware and software components implemented using technologies such as facial expression recognition, voice emotion recognition, text emotion analysis, or physiological signal analysis.

[0025] "User's emotions" refers to a user's emotional state or tendency within a specific time period, such as happiness, sadness, excitement, calmness, or being moved. These emotions can be indirectly inferred from a user's facial expressions, tone of voice, word choice, physiological signals, or interactive behaviors.

[0026] "Sentiment analysis results" refer to structured information about user emotion categories, intensity, or trends of emotion changes obtained by processing and analyzing user-related data using sentiment analysis methods. This information is used to guide adjustments to video content.

[0027] "Short videos" refer to short video content with a total playback time limited to a predetermined range, typically used for posting on social media or short video platforms, such as videos ranging from a few seconds to tens of seconds.

[0028] "Necessary playback duration" refers to the total playback time set or required by the user or system for the target video in a scenario where the purpose is to publish a short video, which is used to constrain the overall duration range of the generated video.

[0029] "Shot clips that should be used" refers to video clips that are pre-specified by the user or determined by the system according to logic, and that must be included or preferentially included when generating the target video. Its characteristics include the material identifier and the start and end times in the material.

[0030] "Artificial intelligence technology" refers to a set of technologies that use methods such as machine learning, deep learning, pattern recognition, natural language processing, or computer vision to extract features, analyze patterns, and make inference decisions from input data, thereby achieving automatic understanding and processing of material content and user needs. Attached Figure Description

[0031] Figure 1 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the first embodiment.

[0032] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.

[0033] Figure 3 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the second embodiment.

[0034] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.

[0035] Figure 5 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the third embodiment.

[0036] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing device and head-mounted terminal according to the third embodiment.

[0037] Figure 7 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the fourth embodiment.

[0038] Figure 8 This is a conceptual diagram illustrating an example of the main functions of the data processing device and robot according to the fourth embodiment.

[0039] Figure 9 This represents an emotion map that maps multiple emotions.

[0040] Figure 10 This represents an emotion map that maps multiple emotions.

[0041] Figure 11 This is a sequence diagram illustrating the processing flow of the data processing system of the first embodiment.

[0042] Figure 12 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 1.

[0043] Figure 13 This is a sequence diagram illustrating the processing flow of the data processing system of the second embodiment.

[0044] Figure 14 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 2. Detailed Implementation

[0045] Hereinafter, an example of an implementation of the system according to the present disclosure will be described with reference to the accompanying drawings.

[0046] First, let me explain the terminology used in the following instructions.

[0047] In the following embodiments, the processor (hereinafter referred to as "processor") with reference numerals may be a single computing device or a combination of multiple computing devices. Furthermore, the processor may be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.

[0048] In the following embodiments, RAM (Random Access Memory), as indicated in the figures, is a memory that temporarily stores information and is used as working memory by the processor.

[0049] In the following embodiments, the memory, as indicated by the reference numerals, is one or more non-volatile storage devices that store various programs and parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disks (e.g., hard disks), or magnetic tapes.

[0050] In the following embodiments, the communication I / F (Interface) with reference numerals is an interface that includes a communication processor and an antenna, etc. The communication I / F is responsible for communication between multiple computers. As an example of a communication specification applicable to the communication I / F, wireless communication specifications such as 5G (5th Generation Mobile Communication System), Wi-Fi (wireless fidelity) (registered trademark), or Bluetooth (registered trademark) can be listed.

[0051] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" connects to express more than three items, the same interpretation as "A and / or B" applies.

[0052] First Implementation Method Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.

[0053] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. A server can be cited as an example of the data processing device 12.

[0054] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0055] The smart device 14 includes a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiving device 38, output device 40, camera 42, and communication I / F 44 are also connected to the bus 52.

[0056] The receiving device 38 includes a touchscreen 38A and a microphone 38B, and receives user input. The touchscreen 38A receives user input via touch by detecting contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input via sound by detecting the user's voice. The control unit 46A in the processor 46 sends data representing the user input received by the touchscreen 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data representing the user input.

[0057] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting data in a form perceptible to the user 20 (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0058] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.

[0059] Figure 2 The diagram shows an example of the main functions of the data processing device 12 and the smart device 14.

[0060] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0061] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).

[0062] In the smart device 14, the processor 46 performs the acceptance output processing. The memory 50 stores the acceptance output program 60. The acceptance output program 60 is used in conjunction with the data processing system 10 and the specific processing program 56. The processor 46 reads the acceptance output program 60 from the memory 50 and executes the read acceptance output program 60 on the RAM 48. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48. Furthermore, the smart device 14 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48.

[0063] Alternatively, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. Furthermore, the data processing device 12 may be a server device or a user-held terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing of the data processing system 10 of the first embodiment will be described.

[0064] Example 1 The flow of a specific process in Example 1 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. Furthermore, the data processing device 12 is referred to as the "server," and the smart device 14 is referred to as the "terminal."

[0065] Existing video generation and editing technologies typically employ two approaches: one involves users manually editing, adding effects, dubbing, and color grading using a graphical user interface; the other involves automatically splicing footage based solely on simple templates or fixed rules. The former heavily relies on users' professional editing skills and extensive interactive operations, resulting in a lengthy and inefficient process that struggles to adapt to the nuanced requirements of different publishing scenarios (such as short-duration dissemination). While the latter lowers the operational barrier, its automatic processing logic often relies on pre-defined, fixed rules, failing to fully understand the semantic content of multimodal materials and hindering personalized editing based on users' abstract intentions and emotional preferences, thus limiting the quality and relevance of the generated videos.

[0066] On the other hand, with the rapid development of generative AI models in text, image, and audio fields, existing technologies mostly use simple text prompts input by users as model prompts for coarse-grained video generation or editing. They fail to construct a system-level integrated processing architecture encompassing raw multimodal material analysis, feature extraction, structured prompt generation, and precise timeline synthesis based on editing plan information. This direct invocation method, lacking an "intermediate representation layer," suffers from the following problems: (1) The server cannot fully utilize the temporal features, semantic features and structural features obtained from fine-grained analysis of video, image and audio materials. Therefore, generative artificial intelligence models have difficulty accurately controlling the composition of the final video at the shot level and frame level. (2) The prompts are usually just simple transcriptions of the user's natural language, and fail to systematically embed editing parameters such as duration constraints, shot selection rules, transition modes, subtitle content and audio mixing strategies, resulting in insufficient controllability and stability of the generated results; (3) The lack of an automatic duration control and segment selection mechanism for specific application scenarios such as short-duration videos limits the adaptation efficiency and technical effectiveness in applications such as short video platforms and advertising. (4) The lack of an automatic adjustment mechanism based on user emotional state and iterative feedback means that the server cannot optimize video content in a closed loop without increasing the user's operational burden, resulting in a weak adaptive ability of the system to user subjective satisfaction.

[0067] Therefore, it is necessary to provide a new system and processing method that, on the server side, standardizes and extracts multimodal features from unedited video, image, and audio materials to automatically generate structured prompts, drives generative artificial intelligence models to output executable editing plan information, and automatically adjusts the generated results based on user emotional state and iterative feedback. This improves the performance, controllability, and user experience of video generation and editing at the computer architecture and multimedia processing workflow levels.

[0068] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 1 is achieved by the following means.

[0069] In this invention, the server includes: an input unit for acquiring unedited video footage, image footage, and audio footage; a feature extraction unit for standardizing the format of the unedited video footage, image footage, and audio footage and extracting visual features, acoustic features, and text information; a condition receiving unit for acquiring abstract conditional information including the target video's purpose, atmosphere, usage scenario, and duration; a prompting statement generation unit for generating prompting statements that provide a layered description of the video's composition, duration allocation, scene selection, and effect additions based on the abstract conditional information and the multimodal features; and a generative artificial intelligence model for invoking the prompting statements and the multimodal features and acquiring information containing elements. The system includes an editing plan acquisition unit (containing editing plan information such as material identifiers, time intervals, scene transition types, visual effects, subtitle information, and acoustic effects), a video compositing unit (for segmenting, splicing, effects overlaying, subtitle rendering, and audio mixing of raw multimedia data based on the editing plan information to generate complete video data), an iterative generation unit (for mapping feedback information to new abstract condition information and regenerating prompts after receiving user evaluation information or modification requests, and for re-acquiring editing plan information and completing video data), and a sentiment analysis unit (for automatically adjusting the content of the completed video data by correcting the abstract condition information or the editing plan information based on information representing the user's emotional state). This allows for a closed-loop processing flow within the server, from multimodal feature extraction, structured prompt generation, editing plan inference to precise timeline compositing. This enables the computer system to more efficiently and precisely control the video generation process while meeting duration constraints and content strategies in different scenarios, and to continuously optimize the generation results based on user emotions and interactive feedback, thereby achieving substantial improvements in multimedia processing efficiency, generation quality, and human-computer interaction experience.

[0070] "Unedited video footage" refers to video data collected or acquired by users that has not undergone editing operations such as editing, special effects processing, dubbing mixing, or color correction.

[0071] "Image material" refers to still image data that has been collected or acquired by users and has not undergone editing operations such as retouching, filtering, cropping, or layout.

[0072] "Audio material" refers to sound data collected or acquired by users that has not yet undergone editing operations such as cutting, mixing, noise reduction, or effects processing, including speech, background music, and ambient sounds.

[0073] "Input unit" refers to a hardware or software module used to receive unedited video, image, and audio materials from a terminal or storage device, including communication interfaces, file upload interfaces, and their control programs.

[0074] "Format standardization" refers to the process of uniformly converting multimedia data with different encoding formats, resolutions, frame rates, or sampling rates into predetermined encoding methods, spatial resolutions, and temporal resolutions.

[0075] "Visual features" refer to numerical or symbolic information calculated from video or image data to represent the content and attributes of a scene, including color distribution, brightness, texture, motion, scene category, object category and its spatial location.

[0076] "Acoustic features" refer to numerical or symbolic information calculated from audio data to represent the content and attributes of sound, including energy, loudness, spectral distribution, fundamental frequency, rhythm, beat, and timbre-related parameters.

[0077] “Text information” refers to text data extracted from audio materials through speech recognition or other means, or descriptive text data associated with multimedia materials.

[0078] "Abstract conditional information" refers to non-specific timeline parameters provided or inferred by the user to describe the overall requirements of the target video, including purpose, atmosphere, usage scenario, target duration, style preference, and content focus.

[0079] A "conditional receiving unit" refers to a hardware or software module used to receive and store abstract conditional information, including a user interface handler and a backend data receiving program.

[0080] "Prompt statements" refer to text data constructed for generative artificial intelligence models, which describe instructions such as video composition, scene selection, duration control, and effect settings in natural language or structured text.

[0081] The "prompt statement generation unit" refers to a functional module that automatically generates prompt statements based on abstract conditional information and multimodal features, including rule engines, template filling logic, or artificial intelligence-based text generation programs.

[0082] "Generative AI model" refers to a model that uses generative AI technology to process input prompts and multimodal features and output video content or video editing solutions, including multimodal generative models based on deep learning.

[0083] "Editing plan information" refers to structured data output by generative artificial intelligence models that guides subsequent video compositing. It includes at least material identifiers, time intervals, transition types between scenes, visual effects, subtitle information, and acoustic effects.

[0084] The "Edit Plan Acquisition Unit" refers to the functional module used to send prompts and feature data to the generative artificial intelligence model and receive editing plan information, including the model call interface and the result parsing program.

[0085] The "video compositing unit" refers to the processing module that cuts, splices, overlays special effects, overlays subtitles, and mixes audio tracks from original video materials, image materials, and audio materials according to the editing plan information to generate complete video data.

[0086] "Completed video data" refers to the target video file or video data stream that is available for playback or distribution by a terminal after the unedited material has been processed in the video compositing unit according to the editing plan information.

[0087] The "iterative generation unit" refers to the control module used to receive user evaluation information or modification requests, transform them into updated abstract condition information, regenerate prompt statements, obtain editing plan information again, and regenerate completed video data.

[0088] "User emotional state" refers to information obtained through sensor data, user input, or model inference that represents a user's current emotional tendency or preference, including emotional categories or related parameters such as pleasure, calmness, tension, and sadness.

[0089] The “sentiment analysis unit” refers to a module used to acquire and analyze users’ emotional state information, and to modify abstract conditional information or editing plan information based on the analysis results, thereby adjusting the atmosphere, rhythm or content emphasis of the completed video data.

[0090] In one embodiment of the invention, the server is deployed on a computing device in a data center. This computing device may include a multi-core general-purpose processor, a graphics processor, main memory, non-volatile storage, and a network interface. The server preferably runs in an environment based on a general-purpose operating system, such as a Unix-like operating system, utilizing a containerized environment to deploy various service modules. The server may use, but is not limited to, the following software components: FFmpeg for multimedia encoding / decoding and filtering, OpenCV for image and video frame processing, PyTorch for deep learning inference, and an audio processing library for audio analysis, as well as a service framework for implementing generative artificial intelligence model inference.

[0091] In one embodiment of the present invention, the terminal may be a mobile computing device or other information processing device, such as a smart terminal equipped with a display device, a touch input device, a microphone, and a camera. The terminal runs an application program or browser script to provide a user interface, collect unedited video footage, image footage, and audio footage provided by the user, and collect abstract conditional information described by the user in natural language.

[0092] In one embodiment of the present invention, a user selects a locally stored multimedia file through the terminal's interface and inputs descriptive information such as the desired video's purpose, atmosphere, duration, and content focus. The terminal sends the above data to a server via a network, and after receiving the results returned by the server, presents and plays the generated video.

[0093] In one embodiment of the present invention, the server includes multiple modules such as an input unit, a feature extraction unit, a condition receiving unit, a prompt statement generation unit, an editing plan acquisition unit, a video synthesis unit, an iterative generation unit, and a sentiment analysis unit. Each module runs on the server's processor in the form of a process, thread, or service, and interacts with data through shared storage or a message queue.

[0094] The server uses a network interface control program and a file receiving program in the input unit to receive and parse data streams from the terminal. After receiving unedited video, image, and audio materials, the server stores them in non-volatile storage and registers corresponding metadata such as material identifiers, file paths, encoding formats, durations, and resolutions in the database. In this way, the server utilizes structured data to manage multimedia resources, enabling searchable storage of large amounts of material, thereby reducing the number of random disk accesses and lowering I / O latency during subsequent queries and processing.

[0095] In the feature extraction unit, the server calls FFmpeg to perform demultiplexing and transcoding operations on unedited video and audio footage. The server can uniformly convert videos to a predetermined encoding method (e.g., H.264-based video encoding), a uniform spatial resolution (e.g., full HD resolution or resolution suitable for vertical short videos), and a uniform temporal resolution (e.g., 25 frames per second or 30 frames per second). Through this standardization process, the server unifies footage from different sources and in different formats into a unified internal representation, reducing branching judgments on multiple formats in subsequent processing logic, thereby improving the stability and efficiency of the overall processing pipeline.

[0096] The server uses OpenCV in the feature extraction unit to decode the standardized video frame by frame, calculating the color histogram differences, edge intensity changes, and optical flow information between adjacent frames to detect shot boundaries and motion intensity. The server further selects keyframes from each shot and calls an image recognition sub-model based on a convolutional neural network or visual Transformer to perform scene classification (e.g., outdoor beach, city night scene, indoor scene), object detection (e.g., people, animals, vehicles), and face detection. The server encodes these recognition results into visual feature quantities in vector form and generates a structured feature record for each shot, including shot start time, end time, scene type label, main object category, brightness range, and motion intensity level. By performing this feature vector-based representation internally on the server, the system can quickly filter shot segments that meet the constraints during subsequent editing plan inference through vector comparison and filtering operations, significantly reducing the number of candidate segments requiring deep inference and thus improving the computational efficiency of the generation process.

[0097] The server also uses a speech recognition engine in its feature extraction unit to transcribe the audio material, mapping the audio waveform into a text sequence and its corresponding timestamp. The speech recognition engine can employ a sequence-to-sequence structure model, such as an acoustic-language joint model consisting of multiple convolutional layers, recurrent neural network layers, or a Transformer encoder and decoder. During speech recognition, the server performs frame segmentation, windowing, Fourier transform, and Mel-frequency cepstral coefficient calculation on the audio signal to generate acoustic feature vectors. These vectors are then used for acoustic model inference via a neural network, outputting probability distributions at the phoneme or word level. Finally, a language model decodes the data to obtain the text sequence. The server stores the recognized text and timestamps in a database as textual information for subsequent subtitle generation and semantic analysis.

[0098] In the feature extraction unit, the server uses an audio analysis library to calculate the beat, rhythm intensity, spectral centroid, and energy peak position of the audio material, generating acoustic feature quantities. The server can calculate the number of beats per minute using short-time Fourier transform and beat detection algorithms. In subsequent editing control, the server aligns shot transitions based on the beat position, ensuring the editing rhythm matches the music rhythm, thereby improving the perceived smoothness of the final cut. This approach of writing beat detection results into structured feature data and aligning them with the timeline facilitates optimization decisions based on time and rhythm constraints when generating the editing plan, thus achieving better audiovisual matching effects with the same hardware resources.

[0099] The server parses abstract conditional information from the terminal in the conditional reception unit, including the user-entered target purpose, atmosphere, target duration, and style description. The server can convert the natural language description into semantic vectors using a text encoding model, such as a Transformer-based text encoder, embedding each word as a fixed-dimensional vector, and encoding the entire sentence using multi-head self-attention and a feedforward network. The server stores the encoded semantic vector along with numerical constraints (e.g., target duration) and classification labels (e.g., usage scenario type) as a conditional vector, providing the foundation for subsequent prompt generation.

[0100] In the prompt generation unit, the server utilizes the aforementioned abstract conditional information and multimodal features to jointly generate structured prompts. The server can employ a two-level generation strategy: first, a rule engine filters out a set of candidate shots that meet constraints related to scene category, brightness, and motion intensity based on the target duration and purpose; then, a submodule based on a text generation model inputs the candidate shot set along with the conditional vectors to generate prompts in natural language. These prompts not only contain the user's original abstract description but also specific instructions obtained from server analysis, such as prioritizing certain shot types, controlling shot duration, and specifying transition methods. Therefore, they more completely express the constraints on the generative AI model than simple text manually written by the user, thereby improving the controllability and consistency of the model's output.

[0101] For example, the server can generate one of the following prompt statements: "Use the following materials to create a short, vertical travel recap video of approximately 60 seconds. The overall atmosphere should be relaxed and enjoyable, suitable for posting on short video platforms."

[0102] 1. Prioritize beach scene shots (ID: s1, s3, s5), as these shots are bright and dynamic.

[0103] 2. Insert appropriate city night scene shots (ID: s2, s4) as a transition.

[0104] 3. Switch camera angles between every two strong beats, following the rhythm of the background music, to maintain a relatively fast pace.

[0105] 4. Use a city wide shot for the first 3 seconds to create the feeling of arriving in the city; use a slow-motion shot of the sunset over the sea for the last 4 seconds to create a sense of ending and lingering impression.

[0106] 5. Add appropriate text descriptions below the video, based on key information from the audio-to-text transcription, such as 'Arrived at the beach on day one' or 'Visiting the night market in the evening.' The server can also generate different prompt message examples for other applications, such as: "Please create a 30-second vertical short video based on the following materials. The overall style should be lively and fun, suitable for posting on short video platforms. Prioritize clips containing smiling faces and jumping actions, and switch shots on each beat of the music rhythm. Material list: Video 1 (walking on the beach), Video 2 (riding a roller coaster at an amusement park), Video 3 (eating snacks at a night market), Image 1 (group photo), Image 2 (night scene)." or "Please use the uploaded family gathering videos and photos to automatically generate a heartwarming commemorative video of approximately 90 seconds. The visual style should be soft and warm, suitable for playing at family gatherings. Please automatically recognize people and try to ensure that each family member appears multiple times in the video. Choose soothing piano music for the background music, and display the names of the people and simple blessings in the subtitles." In one embodiment of the invention, the server invokes a generative artificial intelligence model in the editing plan acquisition unit. This model can be a multimodal Transformer structure, comprising a text encoding subnetwork, a visual feature encoding subnetwork, and an audio feature encoding subnetwork. The server inputs prompts into the text encoding subnetwork and inputs visual and acoustic features into the other modality encoding subnetworks respectively. Each subnetwork aligns different modalities in a high-dimensional feature space through multi-layer self-attention and cross-attention mechanisms, outputting a unified representation. In the model's decoding part, the server uses a sequence generation mechanism to generate editing decisions on a time-step basis, including selecting a material identifier, determining start and end times, specifying transition types, and visual effect parameters. The server can introduce a specific loss function into the model during the training phase, such as a weighted loss function composed of shot sorting error, duration deviation, and visual matching degree, and repeatedly update the model parameters using a gradient descent algorithm, enabling it to learn the statistical patterns of human editing behavior on a large-scale labeled dataset.

[0107] During the inference phase, the server executes the generative AI model, deploying computations on GPUs to perform matrix multiplication and addition operations and attention weight calculations in parallel. This hardware acceleration, combined with a unified feature vector representation, allows the same hardware resources to handle more project requests per unit time, increasing the overall video generation throughput of the platform. Simultaneously, by introducing duration and rhythm alignment constraints during generation, the server limits the search space to a set of editing schemes that meet these hard constraints, thereby reducing the evaluation of invalid candidates and improving inference efficiency.

[0108] The server generates specific compositing instructions within the video compositing unit based on the editing plan information. The server uses FFmpeg's filter graph mechanism to construct a video processing pipeline, chaining operations such as cutting, transitions, color adjustments, text overlays, and audio mixing into one or more processing paths. Since the editing plan information already specifies the start and end times, transition types, and subtitle time intervals for each segment, the server only needs to map these parameters to FFmpeg's filter configuration to automatically generate multimedia processing commands. This hierarchical structure—first generating a high-level editing plan in the feature and symbol spaces, and then having the underlying compositing module perform specific pixel-level operations—reduces the complexity of control logic compared to traditional frame-by-frame script control, improves the system's adaptability to different hardware environments, and reduces redundant decoding and encoding operations, thereby shortening processing time while maintaining image quality.

[0109] In its iterative generation unit, the server receives user feedback and modification requests uploaded by the terminal, converts this feedback into new constraint signals, and updates the abstract condition information. The server can detect the user's physiological feedback or explicit evaluations (such as satisfaction ratings) during playback through the sentiment analysis unit, mapping them to changes in atmosphere, rhythm, or content highlights. After updating the abstract condition information, the server regenerates prompts and again calls the generative AI model to obtain a new editing plan. Through this feedback loop, the server can gradually converge to an editing scheme that better suits user preferences within a finite number of iterations. Compared to the method of users manually modifying the timeline multiple times, this iterative approach achieves "batch optimization" within the server through adjustments to the feature space and condition space, reducing the complexity of the user interface and lowering the latency and error rate caused by human-computer interaction.

[0110] In one embodiment of the invention, the server can use a dedicated emotion recognition model in the emotion analysis unit. This model, based on a neural network structure, classifies user emotions using features extracted from audio intonation, user text feedback, and even facial expression images captured by the terminal's camera. After obtaining the emotion prediction result, the server incorporates it as an additional condition into the abstract conditional information. For example, when it detects that the user prefers a more cheerful video style, the server proactively increases the priority of selecting high-brightness, high-motion-intensity shots and increases the weight of faster-paced editing schemes during the prompt generation process. In this way, the user's emotional state directly affects feature selection and parameter settings, forming a causal link from input, processing to output, thereby enabling the system to have adaptive adjustment capabilities at the technical level, rather than simply remaining at the level of static template replacement.

[0111] In another embodiment of the invention, the server can employ different generative artificial intelligence model structures. For example, the server can use a multimodal generator containing a variational autoencoder structure or a diffusion model structure, changing the model's internal generation mechanism without altering the overall framework of feature extraction and prompt generation. This replacement method does not affect the server's external interface and data flow structure, facilitating the selection of appropriate models under different hardware resources and application scenarios, thereby fully utilizing hardware acceleration features while ensuring system stability.

[0112] In another embodiment of the invention, the server can divide the editing plan information into two levels: coarse-grained and fine-grained. At the coarse-grained level, the server first generates shot-level sequencing and duration, and then at the fine-grained level, it generates frame-level transition and effects parameters. The server can use two cascaded generation sub-models; the first sub-model outputs a coarse editing plan, and the second sub-model refines the parameters based on the received coarse plan. Through this hierarchical generation structure, the server can perform local optimization within a smaller search space, improving the quality of the final synthesized video in terms of transition smoothness and visual consistency.

[0113] In summary, by constructing a complete processing system within the server—from multimodal feature extraction and structured prompt generation to editing plan inference based on generative artificial intelligence models and precise synthesis using multimedia processing software—this invention not only automates video generation but also improves data representation, algorithm flow, and hardware resource utilization within the computer. The server utilizes a hierarchical representation of feature vectors, semantic vectors, and plan structures to effectively reduce redundant computation and improve the overall efficiency of model inference and multimedia synthesis. By explicitly incorporating the user's abstract intent and emotional state into the model input, the matching degree and stability of the generated results are improved, technically enhancing generation accuracy and user satisfaction. Therefore, the system of this invention can be used in the real world to produce short-duration video content, promotional videos, and commemorative videos, providing considerable performance optimization and quality improvement in multimedia production pipelines.

[0114] use Figure 11 The processing flow is explained.

[0115] Step 1: Users select materials and enter abstract condition information on the terminal.

[0116] In the terminal interface, the user selects unedited video, image, and audio files from local storage and enters abstract conditional information such as intended use, atmosphere, target duration, and style preference in the text input area. The input consists of multimedia file data and natural language text. The terminal reads the file path, size, and basic metadata of the selected file, packages the natural language text into a string, and encapsulates it into request data via network protocols. The output is an upload request containing multimedia binary data and abstract conditional text.

[0117] Step 2: The terminal uploads materials and abstract condition information to the server.

[0118] The terminal uses its network communication module to send the upload request generated in step 1 to the server via a secure communication channel. The input consists of a request message containing binary streams of video, images, and audio, as well as abstract conditional text. During transmission, the terminal performs data fragmentation, retransmission control, and progress management. The output consists of the upload data stream transmitted over the network to the server, and the upload status information displayed on the user's side.

[0119] Step 3: The server receives and stores unedited multimedia materials.

[0120] The server receives data streams from the terminal in the input unit, parses multi-part form or streaming data, writes video, images, and audio to storage devices respectively, and generates a unique material identifier for each file in the database, registering metadata such as path, format, and duration. Input consists of multimedia data streams and abstract conditional text from the terminal. The server maps unstructured data to file system and database records through parsing and file writing operations. Output consists of the actual multimedia files stored in the file system and database records containing material identifiers and metadata.

[0121] Step 4: The server performs format standardization processing on video and audio materials.

[0122] In the feature extraction unit, the server invokes multimedia processing software to transcode and resample materials with different encoding formats and resolutions. The inputs are the original video and audio file paths and their metadata stored in step 3. The server invokes multimedia processing commands to demultiplex and re-encode the video, convert the sampling rate and number of channels for the audio, and unify the timeline to a predetermined frame rate. The outputs are standardized video and audio files conforming to internal standards (unified encoding method, resolution, and frame rate), as well as updated standardized file path information.

[0123] Step 5: The server extracts visual features and structured information from videos and images.

[0124] The server uses an image processing library to decode standardized video frame by frame and extract keyframes. Simultaneously, it reads uploaded image files and performs scene classification, object detection, face detection, and color histogram calculation on each frame or image. Inputs include: standardized video files, image files, and their paths. The server transforms image data into numerical feature vectors and label sets through convolution operations, attention calculations, and classification decisions, and aggregates them into shot-level feature records based on the timeline. Outputs include: structured information such as visual feature quantities organized by shot or image unit, scene type labels, object category labels, and shot start and end times.

[0125] Step 6: The server extracts acoustic features and text information from the audio.

[0126] The server invokes a speech recognition engine to perform frame segmentation, spectral analysis, and acoustic modeling of the audio signal, decoding it to obtain text content. Simultaneously, it uses an audio analysis library to calculate rhythm, beat position, energy peaks, and spectral features. The input is a standardized audio file and its path. The server uses Fourier transform, feature extraction, and classification inference to map the continuous waveform into a text sequence and a series of acoustic feature vectors. The output is text information including timestamps and a list of acoustic features describing the beat, energy, and spectrum, and these results are written as records to a database.

[0127] Step 7: The server parses and encodes abstract conditional information.

[0128] The server reads the abstract conditional text received in step 3 in the conditional receiving unit, converts it into a semantic vector using a text encoding model, and parses out numerical parameters (e.g., target duration) and categorical parameters (e.g., purpose, atmosphere type). The input is the user's natural language conditional text. The server obtains a fixed-length semantic vector through word segmentation, embedding mapping, and multi-layer attention operations, and normalizes parameters such as target duration into internal numerical representations. The output is a set of abstract conditional vectors and structured parameter records for subsequent inference.

[0129] Step 8: The server generates structured prompts.

[0130] In the prompt generation unit, the server combines visual features, acoustic features, textual information, and abstract conditional vectors to generate prompts for a generative artificial intelligence model. The inputs are: shot-level visual feature records, audio feature records, text transcription results, and abstract conditional vectors. The server first filters candidate shots that meet constraints related to scene category, brightness, motion intensity, and duration using rule-based logic. Then, it uses a text generation algorithm to construct detailed descriptions and an overall description based on the candidate shots and conditional parameters, converting this information into prompts in natural language. The output is a prompt text containing information on the priority of source material used, shot length, transition rules, and subtitle source information.

[0131] Step 9: The server calls a generative artificial intelligence model to obtain editing plan information.

[0132] In the editing plan acquisition unit, the server takes the prompt statement as text input and visual and acoustic features as multimodal feature inputs, and passes them to the multimodal generative AI model. The input consists of the prompt statement generated in step 8 and the multimodal feature vectors from steps 5 and 6. Within the model, the server performs vector transformations through text encoding, visual encoding, and audio encoding sub-networks, then fuses the multimodal representations through an attention layer. Finally, in the decoding layer, it generates editing decisions in chronological order, including material selection, start and end times, transition types, filter parameters, and subtitle content. The output is structured editing plan information, including a list of shots, time intervals, transition effects, visual effects settings, subtitle content, and audio synthesis instructions.

[0133] Step 10: The server generates compositing instructions based on the editing plan and executes the video compositing.

[0134] The server parses the editing plan information in the video compositing unit, constructs a multimedia processing pipeline, and configures cutting, splicing, transitions, filters, text overlays, and audio mixing into specific operation sequences. The inputs are: the original or standardized multimedia file and the editing plan information output from step 9. The server generates corresponding cutting commands based on the source material identifier and time interval for each shot; generates filter configurations based on transition types; generates text overlay configurations based on subtitle text and timestamps; and adjusts volume and fade-in / fade-out parameters based on audio instructions. The output is: a completed video data file generated by the multimedia processing software, i.e., the target video containing images, sound, transitions, and subtitles.

[0135] Step 11: The server provides the completed video to the terminal and receives user feedback.

[0136] The server generates an access address for the completed video in the output module and returns this address to the terminal via a response message. The inputs are: the completed video data path and project identifier generated in step 10. The server carries video metadata in the response and records video version information in the database. After receiving the response, the terminal loads and plays the video, and provides an evaluation button or text input box on the interface. The output is: the completed video stream requested and played by the terminal, and the evaluation text or modification request entered by the user on the terminal.

[0137] Step 12: The server updates the abstract conditions based on user feedback and emotional state, and then generates the video again.

[0138] The server reads user feedback text and possible sentiment signals in the iterative generation and sentiment analysis units, performs sentiment recognition and intent analysis, and maps changes in satisfaction and preferences into incremental updates of abstract conditional information. Inputs include user feedback text, rating data, and sentiment-related signals. The server uses text analysis and sentiment classification models to convert feedback into adjustment parameters for atmosphere, rhythm, or content weighting, updates the conditional vector, and re-executes the prompt generation and editing plan acquisition processes. Outputs include updated prompts, new editing plan information, and newly synthesized finished video data based on this information, thus forming an iteratively optimizeable video generation process.

[0139] Application Example 1 The process flow corresponding to the specific processing in Use Case 1 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. Furthermore, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".

[0140] In existing multimedia content generation technologies, video generation typically relies on human editors who, based on experience, select, edit, and arrange raw video clips, still images, and audio. To meet the requirements of different publishing platforms regarding duration, pacing, and visual style, editors need to manually sift through large amounts of material and design timeline structures and transitions according to the intended theme and purpose. While the enhanced acquisition capabilities of personal terminals allow users to easily access massive amounts of video and image material, computer systems still face the following technical challenges when automatically processing this complex, multimodal data: (1) At the computer level, traditional systems often use fixed templates or simple rules to splice video materials. They cannot dynamically generate a suitable timeline structure and shot selection strategy based on the abstract themes and creative purposes described by the user in natural language, resulting in a large deviation between the generated results and the user's intentions. (2) Although existing generative artificial intelligence models can generate videos or perform style transformations on videos based on text, they usually only use text as an additional input condition. They lack a unified feature modeling and multimodal fusion mechanism for the original video, image, and audio, making it difficult for the model to consider the content of the picture, the semantics of the scene, and the emotion of the audio at the computational level, and to finely sort and crop the importance of different materials. (3) In terms of terminal and server collaborative processing, traditional solutions often directly upload user text and original materials to the server for one-time processing. They lack a pre-judgment and conversion mechanism for material format and size, which can easily lead to large network transmission overhead and redundant server-side preprocessing, affecting the overall computing resource utilization efficiency and system response time. (4) For specific publishing scenarios such as short-duration videos, existing systems usually require manual editing by human based on platform restrictions (such as target duration and recommendation rhythm). There is a lack of a processing flow in which the computer automatically analyzes the user's intent, target duration and material content, and automatically derives the editing strategy and number of shots within the generative artificial intelligence model. As a result, it is impossible to achieve adaptive optimization of the structure of short-duration videos at the algorithm level. (5) In terms of the linkage between user experience feedback and video content generation, traditional systems mostly use user feedback for content recommendation or simple scoring, without introducing a dynamic closed-loop adjustment mechanism for emotional state in the generation pipeline. They cannot automatically adjust the timeline information and screen switching method based on the user's emotional changes during viewing or interaction, thereby improving the adaptability of subsequent generated content.

[0141] Therefore, an improved computer implementation method and system architecture are needed, enabling the server to: efficiently format and extract features from multimodal materials from the terminal; convert user-inputted strings in natural language into prompts and generation prompts understandable by the generative AI model; perform joint processing of multimodal features and semantic vectors within the generative AI model to automatically generate timeline information and editing strategies; and dynamically adjust the generation process based on the user's emotional state. Through these improvements, it is hoped that the intelligence and processing efficiency of automatic video generation can be enhanced at the computer technology level, achieving end-to-end computational optimization from "materials to finished product".

[0142] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is achieved by the following means.

[0143] In this invention, the server includes an input unit for receiving pre-edited video information, still image information, and audio information from a terminal device, determining their format and size, and converting them into a predetermined format; a text processing unit for obtaining string information representing the theme and purpose of the video content to be generated from the user and structuring it into prompt statements that can be received by a generative artificial intelligence model; a material parsing unit for parsing the video information, still image information, and audio information to extract multimodal features and extract candidate time intervals and candidate scenes; and a natural language processing unit for performing natural language processing on the prompt statements to generate semantic vectors and inputting the semantic vectors and the multimodal features together into the generative artificial intelligence model. The intelligent model includes a generation instruction unit that generates generation prompts containing video composition control information; an editing control unit that automatically determines timeline information such as scene selection, sequence, duration, and scene switching methods based on the generation prompts, candidate time intervals, and candidate scenes, and performs video, image, and audio editing accordingly; a video generation unit that performs video and audio encoding processing according to the timeline information to generate output video data that meets predetermined compression methods and image quality conditions; and an emotion analysis unit that infers the user's emotional state based on operation history information or biometric information from the terminal device and adjusts the timeline information or scene switching methods according to the emotional state. This creates a closed-loop processing chain within the computer, from raw multimodal materials and natural language descriptions to structured prompts and timeline-level editing control. On the server side, a generative artificial intelligence model jointly calculates multimodal features and semantic vectors, automatically determining shot selection and pacing suitable for different release scenarios (including short-duration videos), and dynamically adjusting the generation strategy based on user emotional feedback. This improves the intelligence level of automatic video generation, the efficiency of computing resource utilization, and the overall technical performance of the computer processing process.

[0144] "Information processing device" refers to a computing device used to execute programs, process multimedia data, and control the operation of various functional units, including but not limited to servers, computer terminals, or cloud computing nodes.

[0145] "Terminal device" refers to an electronic device operated by a user for collecting, storing, and sending multimedia data to a server, as well as receiving and playing out video data, including but not limited to smartphones, tablets, personal computers, or wearable devices.

[0146] "Video information" refers to digital data representing the content of moving images, including but not limited to video frame sequences and their related metadata contained in multimedia files stored in a predetermined encoding format.

[0147] "Still image information" refers to digital image data representing the content of a single frame, including but not limited to photographs, illustrations, or other image data stored in the form of image files and their metadata.

[0148] "Audio information" refers to digital audio data that represents sound content, including but not limited to speech, background music, ambient sound, and their sampled data and related parameters stored in the form of audio files.

[0149] "Input means" refers to hardware, software, or a combination thereof used to receive data from external devices or users, including communication interfaces, network modules, application programming interfaces, and logical functions for processing the format and capacity of the received data.

[0150] “String information” refers to text data consisting of a sequence of characters used to express a user’s intent, theme, or purpose, including natural language sentences, phrases, or keywords.

[0151] "Prompt statements" refer to text instructions that are structured based on string information input by the user and used as input for generative artificial intelligence models to describe the theme, style, length, or other generation requirements of the video to be generated.

[0152] "Generation prompt statements" refer to extended text instructions that are generated based on prompt statements and processed by the generation instruction unit and generative artificial intelligence model. These instructions contain video composition control information and multiple control parameters and are used to guide the subsequent video editing and generation process.

[0153] "Multimodal features" refers to numerical feature vectors or feature sets extracted from video information, still image information, and audio information to represent image content, scene semantics, audio characteristics, etc.

[0154] "Candidate time interval" refers to the set of time periods that can be selected to constitute the target video, based on the analysis results of multimodal feature quantities, on the timeline of video information.

[0155] "Candidate scenes" refer to a set of segments or images that are relatively consistent in semantics or image content, obtained by analyzing the content of video frames or still images and clustering features, and are used as candidate units when composing a video.

[0156] Natural Language Processing (NLP) refers to the process of segmenting, syntactically analyzing, semantically parsing, and vectorizing text data such as prompts to generate numerical representations that can be computed by models.

[0157] "Numerical vectors" refer to arrays of numbers obtained by mapping text or multimedia content to a high-dimensional space through natural language processing or feature extraction processes. These vectors are used for computation and reasoning in generative artificial intelligence models.

[0158] "Generative artificial intelligence models" refer to artificial intelligence models trained using machine learning or deep learning methods that can automatically generate or guide the generation of video content based on input prompts and multimodal features, including but not limited to text-video generation models and multimodal fusion models.

[0159] The "Generation Instruction Unit" refers to a functional module that receives prompt statements, performs natural language processing, generates semantic vectors, and calls a generative artificial intelligence model to output generated prompt statements or control parameters.

[0160] The "material analysis unit" refers to a functional module used to analyze video information, still image information, and audio information, extract multimodal features, and extract candidate time intervals and candidate scenes.

[0161] The "Editing Control Unit" is a functional module that automatically determines timeline information such as scene selection, scene order, scene duration, and screen switching method based on the generated prompts, candidate time intervals, and candidate scenes, and edits the original multimedia materials accordingly.

[0162] "Timeline information" refers to structured descriptive data that defines the order, start and end times, duration, and relationships between scenes of each clip in the target video along the time dimension.

[0163] "Scene transition method" refers to the type of display effect used to transition between adjacent scenes or segments, including but not limited to fade-in, fade-out, cross-dissolve, and fast switching transition effects.

[0164] The “video generation unit” refers to a functional module used to perform video encoding and audio encoding processing on edited multimedia data according to timeline information and scene switching methods to generate output video data.

[0165] "Output video data" refers to video files or video data streams that are encoded by the video generation unit, meet predetermined compression methods and image quality conditions, and can be played or stored on a terminal device.

[0166] The “emotion analysis unit” refers to a functional module used to infer the user’s emotional state based on the operation history information or biometric information obtained from the terminal device, and to adjust the timeline information or screen switching method according to the emotional state.

[0167] "Operation history information" refers to data related to user interaction recorded on the terminal device, including but not limited to interaction logs such as playback duration, number of pauses, fast forward or rewind actions, and repeated viewing actions.

[0168] "Bioinformation" refers to sensor data that can reflect a user's physiological or psychological state, including but not limited to heart rate, skin conductance, facial expression features, and gaze duration.

[0169] "Output means" refers to the hardware, software, or combination thereof used to convert output video data into a distribution format that can be played or saved on a terminal device and send the data to the terminal device via a network.

[0170] In various embodiments of this invention, the server, terminal, and user each assume different functional roles, collaboratively completing the automatic video generation process from multimodal raw materials and natural language descriptions to a structured timeline. The system of this invention is not limited to a specific hardware platform or operating system; however, for ease of understanding, the following description uses the server as a cloud-based information processing device and the terminal as a user-side portable information terminal as an example.

[0171] I. System Overall Structure A server includes: a multi-core central processing unit, a graphics processing unit, main memory, non-volatile memory, a network interface, and a bus connecting these components. The software aspects of a server include: an operating system (e.g., a Linux-based server operating system), web server programs, backend application frameworks, a database management system, multimedia processing libraries (e.g., general-purpose multimedia processing libraries for audio and video processing), and a generative artificial intelligence model runtime environment based on machine learning frameworks (e.g., tensor operation-based frameworks).

[0172] The terminal includes: a processor, memory, a display screen, a touch input device, a camera, a microphone, a speaker, a wireless communication module, and local storage. The software components of the terminal include: a mobile operating system, a graphical user interface framework, network communication libraries, multimedia playback components, and dedicated applications.

[0173] Users operate the application using a terminal, providing multimedia materials and text descriptions to the server; the terminal is responsible for performing local preprocessing and network transmission; the server is responsible for multimodal feature extraction, prompt generation and expansion, timeline structure derivation, encoding output, and adjustment based on the user's emotional state.

[0174] II. Server-side functional modules and data structures 1. Input Unit and Data Management The server receives video, still image, and audio information from the terminal via a network interface. In the backend application, the server creates a "project record" for each uploaded project, which includes at least: project identifier, user identifier, media list, target parameters, and status flags. The server generates a "media record" for each media item, with fields including: media identifier, type (video / image / audio), storage path, duration, resolution, sampling rate, file size, encoding format, and parsing status.

[0175] The server uses a multimedia processing library to read the header information of each source file, determine the encoding format and parameters, and perform format normalization when necessary. For example, it uniformly converts videos to a specified encoding format and resolution, and images to a specified pixel format. Through this centralized normalization, the server can adopt a unified processing pipeline in the subsequent feature extraction and encoding stages, reducing the number of different format branches to judge, thereby reducing the instruction branch misprediction rate, improving the cache hit rate, and improving the overall computational efficiency.

[0176] 2. Text processing unit and prompt statement generation The terminal sends the user-input string to the server, where the server cleans, segments, and normalizes the string in its text processing unit. The server performs word segmentation, noise removal, and synonym normalization on the text to reduce the sparsity of the semantic space. Based on preset templates and rules, the server expands the user input into structured prompts, including information fields such as topic, expected duration, style preference, and key content type.

[0177] For example, when a user enters "Help me make a relaxing vlog documenting my weekend camping trip" in the terminal, the server might generate the following message: "Based on the weekend camping videos and photos uploaded by users, please automatically generate a relaxed and natural-style vlog. Please highlight the camping scene, interactions with friends, and natural scenery. The vlog should be about 2 to 4 minutes long, with soothing and pleasant background music, and appropriate subtitles or titles." When a user enters: "A 30-second promotional video for a new mobile phone, with a high-tech feel," the server can generate the following message: "Please use user-uploaded photos and demonstration videos of the new mobile phone product to automatically generate a short promotional video of approximately 30 seconds. The overall style should be modern and technologically advanced, highlighting the product's appearance details, screen performance, and key features. Please use upbeat background music, paired with clean and crisp transitions, and add brief text descriptions at appropriate points, such as 'ultra-clear screen' and 'long battery life'." Through this structured processing, the server transforms free text into more information-dense and parsable prompts, facilitating subsequent model condition control and parameter decoding. This reduces the impact of language ambiguity on the generated results, thereby improving the consistency between the generated behavior and the user's intent.

[0178] 3. Text Vectorization and Semantic Representation The server encodes the prompts using a deep learning-based natural language encoding model (e.g., a text encoder based on a transformer structure). The server segments the prompts into sub-word units, maps each sub-word to a vector through an embedding layer, and then computes the semantic vector of the entire sentence through a multi-layer self-attention network and a feedforward network.

[0179] The server uses contrastive learning or language modeling objectives when training the encoder to make semantically similar prompts more similar in the vector space. In this way, the numerical vectors obtained by the server not only contain word-level information, but also encode high-level semantic features such as "relaxed," "technological," "memories," and "promotion," providing a compact, high-dimensional, and computable representation for subsequent multimodal fusion and control parameter generation.

[0180] 4. Material Analysis Unit and Multimodal Feature Extraction The server performs scene-level structural analysis on the video information. First, the server uses a multimedia processing library to extract keyframes from the video or samples frame sequences at fixed time intervals. These frames are then input into a pre-trained image recognition network (e.g., using a residual network or an efficient network) to obtain the visual feature vector for each frame. The server calculates the cosine or Euclidean distance between the feature vectors of adjacent frames and, combined with changes in the brightness histogram, identifies scene transition points, dividing the entire video into multiple candidate time intervals.

[0181] The server performs similar feature extraction on still image information and uses clustering algorithms to divide the images into several semantically similar groups to serve as supplementary images in candidate scenes. For audio information, the server obtains time-frequency features through short-time Fourier transform or Mel-frequency transform, and extracts audio emotion and rhythm features using convolution or recurrent structures.

[0182] The server stores the modal features in a structured data table. Each candidate time interval or candidate scene corresponds to a set of feature fields, including: visual semantic vectors, action intensity indicators, image stability scores, sharpness scores, audio energy distribution, and rhythm indicators. These specifically designed feature fields enable the server to evaluate the suitability of segments across different dimensions numerically without relying on explicit human rules in subsequent calculations.

[0183] III. Generative Artificial Intelligence Models and Generative Prompt Statements 1. Multimodal generative model structure The server uses a multimodal generative artificial intelligence model in the instruction generation unit. This model includes a text encoding subnetwork, a multimodal fusion subnetwork, and a control parameter decoding subnetwork. The text encoding subnetwork receives the semantic vector of the prompt statement; the multimodal fusion subnetwork receives the visual features, image features, and audio features of the candidate segments; and the control parameter decoding subnetwork outputs multiple continuous and discrete parameters used to control the time axis.

[0184] The server uses a multi-head attention mechanism in the multimodal fusion sub-network, enabling the model to weight and aggregate features of different candidate segments based on different semantic components in the prompt (e.g., "highlighting the camping scene," "relaxed pace," "30-second duration"). Through this computational approach, the model no longer simply relies on a fixed template, but dynamically determines which segments are more important based on the text's semantics.

[0185] 2. Generate prompt statements and control parameters The server decodes the output of the multimodal generative model into generation prompts. Besides retaining the original text information, these prompts also include control parameters indicating video length, tempo density, average shot length, close-up to long shot ratio, transition type distribution, and music volume curves. For example, generation prompts might implicitly convey structured meanings such as: total duration 120 seconds, approximately 30 shots, average shot length 4 seconds, slightly longer opening and closing shots, and fast-paced segments concentrated in the middle.

[0186] In this way, the server transforms the output of the deep network, which is difficult to interpret directly, into a set of parameters that can be parsed by the editable control unit. This gives the timeline generation process a clear control interface and also makes it easy to replace different generation models or strategies without changing the subsequent pipeline.

[0187] IV. Editing Control Unit and Timeline Information Generation In the editing and control unit, the server utilizes control parameters and the characteristics of candidate time intervals to execute a specific timeline construction algorithm. First, based on the target duration and the number of shots, the server employs heuristic search or dynamic programming to select a set of segments from the candidate interval set, ensuring the total duration is close to the target duration while also considering content diversity and semantic coverage.

[0188] The server calculates a comprehensive score for each candidate segment, which is a weighted sum of multiple features such as text relevance, visual quality, and rhythm matching. The weights are determined by control parameters in the generated prompts. For example, when the prompt emphasizes "relaxed and natural," the server increases the weight of segments with stable visuals, soft colors, and smooth audio energy; when the prompt emphasizes "high-tech feel," the server increases the weight of segments with high brightness, high contrast, and close-up shots of devices.

[0189] The server assigns specific transition effects, such as fade-in / fade-out, cross-dissolve, or fast transitions, to adjacent segments based on the proportion settings for transition methods in the control parameters, and records the effective time and parameters (such as duration and curve function) of each transition in the timeline data structure. This numerically optimized editing strategy differs from traditional manual experience-based operations; it is a machine-driven internal decision-making process driven by multi-objective functions and constraints.

[0190] V. Video Generation Unit and Encoding Processing In the video generation unit, the server drives the multimedia processing library to perform actual image compositing and encoding based on timeline information. The server reads the original video and image materials in timeline order and constructs a frame sequence with uniform resolution through cropping and scaling operations. The server uses transition parameters to insert transition frames between the frame sequences, such as generating cross-dissolve effects through linear interpolation or filter-based blending algorithms.

[0191] At the audio level, the server uses envelope control to achieve volume gradation based on the time alignment information between the music track and the original audio track, and reduces the background music volume in sections with narration. The server inputs the processed video frames and audio samples into the encoder, which performs compression encoding according to pre-set parameters such as bitrate, GOP length, and reference frame structure. Through centralized and unified encoding control, the server can reduce redundant information while ensuring image quality, thereby reducing the size of the generated file and the network transmission burden.

[0192] VI. Emotional Analysis Unit and Dynamic Adjustment When a user watches a generated video on their device, the device can record actions such as playback progress, pausing, fast-forwarding, and replaying. With the user's consent, the device can also collect specific biometric indicators. The device sends this data to a server, where the server uses a rule-based model or a lightweight classification model in its sentiment analysis unit to infer the user's emotional state, such as high interest, moderate interest, or low interest.

[0193] The server adjusts the timeline information and transition strategies of subsequent generated items based on the user's emotional state. For example, when the server detects that the user frequently fast-forwards through a certain type of fast-paced segment, the server appropriately increases or decreases the average shot length under similar prompts; when the user spends a long time on "flashback" videos, the server increases the weight of slow-paced segments and gentle transitions.

[0194] This parameter update based on emotional feedback is not a simple scoring of the generated results, but rather changes the behavior of the entire generation system by modifying the statistical distribution of control parameters and the weight settings of the editing algorithm. This allows for adaptive optimization of the timeline construction logic at the algorithm level, thereby improving the technical performance across multiple generation tasks.

[0195] VII. Terminal-side preprocessing and communication load control Before sending materials, the terminal can perform basic compression and format unification, such as compressing the original high-bitrate video to a predetermined maximum resolution and converting incompatible encodings to standard encodings. The terminal calculates summary information for each file, such as resolution, frame rate, and file size, and sends this information to the server. The server then returns instructions on whether further compression or cropping is needed. This interactive preprocessing mechanism allows the server to control the amount of data uploaded by the terminal, avoiding unnecessary large file transfers, reducing network bandwidth consumption, and shortening overall processing time.

[0196] VIII. Technical Effects and Improvements in Computer Technology Through the above structure and process, the server internally implements a unified data flow from raw multimodal footage to timeline-level editing control. In this data flow, prompts are transformed into explicit control parameters, candidate footage is represented as computable multimodal feature vectors, generative artificial intelligence models perform joint optimization in the multimodal space, and the editing control unit constructs the timeline and transition schemes using numerical methods. Because these processes are executed in a vectorized and batch manner on the server's processor and graphics processing unit, this invention improves computer technology in the following ways: The server reduces branching decisions and redundant decoding by using a unified format and feature structure, thereby improving cache utilization and pipeline efficiency. The server reduces the participation of irrelevant materials in encoding by using multimodal fusion, thereby reducing the amount of encoding computation and storage usage. The server reduces the number of trial-and-error generation attempts by using parameterized control based on target duration and user emotions, thereby improving the first-time success rate of the generated results and reducing overall energy consumption.

[0197] Furthermore, the processing logic of this invention does not simply simulate the manual editing process. Instead, it introduces mechanisms such as multimodal feature scoring, objective function optimization, and automatic parameter decoding within the model, forming an independent decision-making system different from manual rules. By learning the statistical relationships among "text description—multimodal material—timeline composition" in a large number of samples, the server automatically discovers editing patterns that are difficult for humans to explicitly write, thus surpassing traditional template-based or fixed-rule-based automatic generation tools in terms of accuracy and efficiency.

[0198] IX. Other Implementation Methods and Variations In other implementations, the server can adopt different generative artificial intelligence model structures, such as using a diffusion-based video generation model or a temporal attention network; the server can adjust the feature fields and control parameter sets according to specific application scenarios, such as adding layout parameters in educational videos and adding highlight event detection features in live replays.

[0199] The terminal can also perform partial feature extraction functions, such as locally estimating image clarity or scene type and uploading simplified features, thereby further reducing server load. Without departing from the spirit of this invention, the functional division between the server and terminal can be flexibly adjusted according to network conditions and hardware capabilities; these variations all fall within the scope of this invention.

[0200] use Figure 12 The processing flow is explained.

[0201] Step 1: Users select materials and enter text descriptions using the terminal.

[0202] Input: Raw video files, still image files, audio files, and free text entered by the user in the input box, all stored locally on the terminal.

[0203] Output: The terminal displays a "list of materials" (containing the selected file paths and basic attributes) and "raw string information".

[0204] Users browse photo albums or file lists in the terminal interface, select multiple videos, photos, and audio files, and the terminal records the local path, file size, duration, etc. of each file; users type text descriptions such as "Help me make a relaxing vlog to record my weekend camping" in the text input box, and the terminal saves the text in memory for subsequent processing.

[0205] Step 2: The terminal performs format checks and necessary local preprocessing on the materials.

[0206] Input: The list of materials obtained in step 1 and the corresponding local files.

[0207] Output: A "standardized list of materials" that meets the predefined format requirements, along with a format description for each material.

[0208] The terminal reads the header information of each file to determine the encoding format (video encoding, audio encoding, resolution, sampling rate, etc.) and file size. The terminal then uses a local multimedia processing library to transcode or scale files that do not meet the server's requirements; for example, compressing high-resolution videos to a preset maximum resolution or converting unsupported image formats to a universal format. Based on the results of this data processing, the terminal generates new files and updates the paths and parameters in the media list, thus outputting a set of uniformly formatted, uploadable media.

[0209] Step 3: The terminal organizes the raw string information into a basic prompt statement and prepares an upload request.

[0210] Input: The raw string information entered by the user and a list of standardized materials.

[0211] Output: Structured basic prompts and upload request data containing material metadata.

[0212] The terminal performs cleaning operations on the original string, such as removing extra spaces and deleting invalid control characters. Then, it completes the missing information according to the preset template. For example, it automatically adds descriptions such as "duration is about 2 to 4 minutes" and "style is relaxed and natural" to "easy Vlog", forming a more semantically complete primary prompt statement. The terminal packages the material type, quantity, estimated total duration and other information together with the prompt statement into a request body, which is then sent to the server.

[0213] Step 4: The terminal uploads the source files and basic prompts to the server via the network.

[0214] Input: Standardized source files and upload request data.

[0215] Output: Upload completion confirmation message and project identifier generated by the server.

[0216] The terminal establishes a connection with the server via a secure network protocol, reads local file content in chunks, and sends binary data and text fields together. During transmission, the terminal calculates the upload progress based on the number of bytes sent and the total number of bytes, and displays it on the interface. Based on the information returned by the server, the terminal obtains the project identifier and associates it with the local task record, thereby completing the mapping from "local task" to "server project".

[0217] Step 5: The server receives materials and basic prompts and performs storage management.

[0218] Input: binary data of the source material from the terminal, basic prompt statements, and related metadata.

[0219] Output: Material records and project records stored on the server.

[0220] The server parses the upload request, writes the video, image, and audio streams to the backend storage system, and assigns a unique media identifier and storage path to each file. The server also creates project records, writing the project identifier, user identifier, media identifier list, and initial prompts to the database. Through this series of data writes and index updates, the server constructs the basic data structure required for subsequent processing.

[0221] Step 6: The server performs multimodal feature extraction and candidate fragment extraction on the source material.

[0222] Input: Original video files, still image files, audio files, and their metadata stored on the server.

[0223] Output: A list of candidate time intervals and a list of candidate scenes containing multimodal features.

[0224] The server uses a multimedia processing library to extract keyframes or periodically sampled frames from the video and inputs these frames into a pre-trained image feature extraction network to calculate a high-dimensional visual feature vector for each frame. By calculating the differences in feature vectors between adjacent frames, the server detects locations of significant changes in the image, divides the video into several time intervals, and generates comprehensive features (e.g., sharpness score, motion intensity, scene type) for each interval. The server performs the same feature extraction on still images and uses a clustering algorithm to group similar images into candidate scene groups. For audio, the server calculates time-frequency features and extracts parameters such as rhythm and energy distribution. The server populates these feature values ​​into a structured list, which serves as one of the inputs to subsequent generative artificial intelligence models.

[0225] Step 7: The server performs natural language processing on the initial prompt statement and generates a semantic vector.

[0226] Input: The initial prompt text.

[0227] Output: A numerical vector representing the semantic information of the prompt statement.

[0228] The server segments or divides the prompt into sub-word units, maps each unit to a low-dimensional vector using an embedding table, and then inputs the sequence into a transformer-based text encoding network to perform multi-head attention, feedforward network, and other computations to obtain a sentence-level semantic representation. During the computation, the server performs matrix multiplication, addition, and non-linear activation to transform discrete text into fixed-dimensional vectors, enabling subsequent generative AI models to numerically understand abstract concepts such as "weekend camping," "relaxed style," and "30-second advertisement."

[0229] Step 8: The server inputs multimodal features and semantic vectors into the generative artificial intelligence model and generates control parameters.

[0230] Input: Multimodal features of candidate time intervals and candidate scenes, and semantic vectors corresponding to prompt statements.

[0231] Output: A set of control parameters including lens selection strategy, time allocation, transition preferences, etc.

[0232] The server inputs the feature vector of each candidate segment along with the semantic vector of the prompt statement into a multimodal fusion network. An attention mechanism is used to calculate the weights of textual semantics on different segments, and based on these weights, it outputs a relevance score and recommended duration for each segment. Simultaneously, the server outputs global control parameters, such as the target total duration, expected number of shots, average shot duration range, emphasized scene types, and suggested transition distribution. These parameters are generated by the network's final fully connected layer and normalization operations, representing the numerical optimization results performed by the server based on features and semantics.

[0233] Step 9: The server generates timeline information based on control parameters and candidate segments.

[0234] Input: The set of control parameters output from step 8, as well as the characteristics and identifiers of candidate time intervals and candidate scenes.

[0235] Output: Complete video timeline information, including segment order, start and end times, duration, and transition methods.

[0236] The server internally executes a timeline construction algorithm. Based on the target total duration and shot number constraints, it selects a set of segments from the candidate range, ensuring the total duration is close to the target value and the scene content is diverse. The server calculates a comprehensive score for each candidate segment (composed of relevance, quality, and pacing matching), selects the set of segments with the highest scores and whose sum satisfies the constraints, and sorts them according to the logical order suggested by the text semantics (e.g., "intro overview – process details – ending summary"). The server assigns a specific transition type and transition duration to each pair of adjacent segments and writes all of this into the timeline data structure as output.

[0237] Step 10: The server edits and synthesizes multimedia materials based on timeline information.

[0238] Inputs: timeline information, raw video files, still image files, and audio files.

[0239] Output: Unencoded or intermediate format synthesized audio and video stream.

[0240] The server trims segments from the original video according to the start and end times on the timeline, and extends the display duration of still images as needed. The server inserts transition effects based on positions defined on the timeline, such as generating a cross-dissolve sequence by linearly blending adjacent frames or achieving fade-in / fade-out by inserting completely black frames. On the audio track, the server aligns and overlays the background music and the original audio track according to the timeline, performing volume envelope operations to achieve effects such as crescendo, diminuendo, and speech priority, ultimately forming an audio and video frame sequence consistent with the timeline.

[0241] Step 11: The server encodes the synthesized audio and video streams to generate output video data.

[0242] Input: The synthesized audio and video frame sequence and target encoding parameters (resolution, bit rate, frame rate, etc.).

[0243] Output: Output video data file or data stream conforming to a predefined encoding format.

[0244] The server sends video frames to the video encoding module, where they are compressed and encoded according to the set frame rate and bit rate strategy. Simultaneously, audio samples are sent to the audio encoding module and compressed using a specified encoding scheme. During the encapsulation phase, the server multiplexes the encoded video and audio streams into the same container and generates file header information and an index table. Through this encoding process, the server compresses data volume while maintaining image and audio quality as much as possible, providing efficient data for transmission and playback on the terminal.

[0245] Step 12: The server infers the emotional state based on user feedback data and updates the generation strategy accordingly.

[0246] Input: Feedback data such as playback behavior logs and biological information from the terminal, as well as corresponding video project records.

[0247] Output: Updated sentiment assessment results and new control parameter configuration.

[0248] The server analyzes user behaviors such as pausing, fast-forwarding, and replaying during viewing, and combines this with possible changes in physiological indicators. It then uses an emotion classification model or rule system to score the user's current experience, categorizing it into levels such as "slightly satisfied," "neutral," and "slightly dissatisfied." Based on these emotion results, the server adjusts the distribution of control parameters. For example, it reduces the weight of unpopular fast transitions and increases the priority of more popular shot durations. The updated configuration is saved as the default parameters for subsequent similar prompts, thus iteratively optimizing the generation strategy based on data.

[0249] Step 13: The terminal receives and outputs video data and provides playback or save functions.

[0250] Input: The address of the output video data returned by the server or the video data itself, as well as related metadata (duration, resolution, thumbnail, etc.).

[0251] Output: A visual and controllable playback interface on the terminal and locally stored video files (such as when the user chooses to download).

[0252] The terminal requests video data via the network, uses its built-in player component to parse the video container and decode and display it frame by frame, while simultaneously decoding audio and outputting it through the speaker. The terminal provides control buttons on the interface for play, pause, progress bar dragging, sharing, and downloading. When the user selects to download, the entire video is written to the local storage system. During playback, the terminal records user interactions and uploads this data to the server as feedback, providing a basis for subsequent task generation.

[0253] Alternatively, an emotion engine for inferring user emotions can be combined. That is, the specific processing unit 290 can also use the emotion-specific model 59 to infer user emotions and perform specific processing using user emotions.

[0254] Example 2 The flow of a specific process in Example 2 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. The data processing device 12 will be referred to as the "server," and the smart device 14 as the "terminal."

[0255] While existing systems can automatically synthesize videos based on user-input natural language prompts, several technical shortcomings remain at the computer implementation level, despite advancements in generative AI model-based automatic video generation technology. First, current technologies often simply input prompts as text descriptions into the model, directly outputting a list of video clips or scripts. This lack of mechanisms for structured management of source data, duration constraint verification, and explicit modeling of the editing process leads to frequent errors in timeline consistency, source data validity, and image attribute matching, increasing the workload of subsequent manual corrections and reducing overall system efficiency. Second, existing systems typically do not uniformly record and analyze structured intermediate data from prompts, model outputs, and rendering processes. This prevents adaptive optimization of the generative AI model's input format and video composition rules based on historical generation behavior, hindering continuous improvement in parsing accuracy and generation quality during system operation. Third, for scenarios with strict requirements on rhythm, information content, and scene transition intervals, such as short-duration distribution videos, current technologies lack computer-level automatic time allocation and interval adjustment mechanisms, making it difficult to guarantee a high-density and rhythmically reasonable editing strategy within the target playback duration constraint.

[0256] Therefore, it is necessary to provide a technical solution that improves the computer system architecture and data processing flow. On the server side, unedited material data and natural language prompts are structured and parsed, duration and material attribute are verified, editing flow definitions are generated, and historical information-driven rule updates are performed. This improves the accuracy of timeline control, the effectiveness of material utilization, and the stability and scalability of generative artificial intelligence model calls during automatic video generation, thereby achieving a substantial improvement in computer video processing capabilities.

[0257] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 2 is achieved by the following means.

[0258] In this invention, the server includes: a device for receiving unedited video data, image data, and audio data via a communication network through a processing device and parsing their types, recording durations, and image attributes; a device for receiving natural language prompts containing target playback duration, aspect ratio, type of publishing media, and presentation guidelines, and converting them into structured video design information; a device for performing natural language processing on the prompts and video design information using a generative artificial intelligence model to generate video composition data containing time intervals, material types, text information content, image compositing information, and audio processing information; a device for automatically adjusting the time intervals based on the video composition data in the processing device to make their total duration consistent with the target playback duration, and for verifying and correcting the reference range according to the recording duration and attributes of the material; a device for generating editing flow definition data corresponding to the video composition data and controlling the video processing program or video editing device to automatically perform material time extraction, image compositing, text information overlay, and audio level adjustment to generate completed video data; and a device for recording the prompts, video composition data, and generating processing history information, and updating the input format of the generative artificial intelligence model or the video composition data generation rules based on the history information. This allows for structured management of materials and instructions on the server side, ensuring the consistency and legality of automatically generated videos in terms of timeline, material references, and image attributes. Furthermore, by continuously optimizing the calling strategy and composition rules of the generative artificial intelligence model through feedback updates of historical data, the processing efficiency, robustness, and scalability of computers in automatic video generation tasks are improved, resulting in a substantial improvement in computer video processing technology.

[0259] "System" refers to a computer implementation consisting of at least one processing device, storage device, and communication device, used to perform a series of processing flows such as video data reception, parsing, generation, and distribution.

[0260] "Processing device" refers to a computing device or its logical unit that includes at least one processor and its control program, used to perform computational processing such as data parsing, model calling, duration verification, editing process generation, and result output.

[0261] "Communication network" refers to wired or wireless data communication infrastructure used to transmit video data, prompts, control information, and result data between servers and information processing terminals.

[0262] "Unedited video data" refers to video files or video data streams formed from raw video signals that have not been edited, composited, or processed with special effects.

[0263] "Image data" refers to still image information, including photographs, image frames, or other visual data stored in the form of a pixel matrix.

[0264] "Audio data" refers to digital signals that reflect sound information, including audio files or audio data streams such as background music, ambient sounds, or voice recordings.

[0265] "Data type" refers to the attribute information that distinguishes whether the input data belongs to video, image, audio, or other media types.

[0266] "Recording duration" refers to the length of time from the start to the end of recording video or audio data.

[0267] "Image attributes" refers to a set of parameters that characterize the features of an image, such as resolution, aspect ratio, frame rate, and color space, which are related to video or image data.

[0268] "Information processing terminal" refers to a user-side computing device used to interact with a server, input prompts, upload materials, and receive generated results, such as a smartphone, tablet computer, or personal computer.

[0269] "Target playback duration" refers to the total playback time of the video that the user or system presets when generating the video.

[0270] "Aspect ratio" refers to the aspect ratio of a video frame, used to characterize the display aspect ratio of a video, such as landscape or portrait ratio.

[0271] "Publication Media Type" refers to the platform category or usage scenario for the scheduled release of the video, such as online platforms, mobile applications, or display terminals.

[0272] "Performance guidelines" refer to the overall design requirements or guiding principles for a video in terms of style, rhythm, emotional tone, or information density.

[0273] "Natural language prompts" refer to text instructions entered by users in natural language to describe the content, structure, style, and constraints of a video.

[0274] "Structured video design information" refers to a set of video design parameters represented by a predetermined data structure, obtained by parsing natural language prompts, including duration requirements, segment divisions, and content elements.

[0275] "Generative artificial intelligence models" refer to artificial intelligence models that can automatically generate text, structured information, or other outputs based on input data, especially natural language processing models based on deep learning.

[0276] Natural Language Processing (NLP) refers to the process of parsing, understanding, and transforming natural language text using algorithms and models, including operations such as word segmentation, semantic analysis, and structured output.

[0277] "Time interval" refers to a continuous time period with a start time and an end time that is divided within the target playback duration.

[0278] "Material type" refers to the category identifier used to distinguish whether the data selected in each time interval is video material, image material, or audio material.

[0279] "Textual information content" refers to textual information presented in the form of subtitles, titles, or descriptions in a video.

[0280] "Compositing information" refers to a set of parameters used to control how multiple videos, images, or layers are superimposed, arranged, transformed, and transitioned on the timeline.

[0281] "Audio processing information" refers to a set of parameters used to control audio processing behaviors such as volume, mixing, fade-in / fade-out, and channel allocation.

[0282] "Video composition data" refers to structured control data generated based on prompts and material information, used to determine time intervals, material selection, subtitle content, image compositing, and audio processing.

[0283] "Total duration of time intervals" refers to the cumulative length of all time intervals in the video data.

[0284] "The scope of material reference" refers to the time segment range selected from the original video or audio data for editing when generating the video.

[0285] "Verification and correction" refers to the process of checking the consistency between the time interval and the target playback duration, as well as the matching relationship between the scope of material references and the duration and attributes of the material records, and automatically adjusting when inconsistencies exist.

[0286] "Editing workflow definition data" refers to data structures or scripts used to describe the sequence and parameters of each step in automatic video editing (including editing, compositing, overlaying, and audio processing).

[0287] "Video processing program" refers to a software module or application that runs on a computing device and is used to perform operations such as video editing, compositing, encoding, and transcoding.

[0288] "Video editing device" refers to a computing device or special-purpose device with video editing functions, including the editing program and control interface running on it.

[0289] "Material time extraction" refers to the process of extracting corresponding segments from the original video or audio according to a specified start and end time, based on the video composition data.

[0290] "Image compositing" refers to the process of overlaying multiple videos, images, or layers on the same screen or arranging them sequentially on a timeline according to compositing information.

[0291] "Text overlay" refers to the process of overlaying subtitles, titles, or explanatory text onto a video frame at a specified time period and position.

[0292] "Audio level adjustment" refers to the operation of boosting, lowering, or balancing the volume of an audio signal according to set parameters.

[0293] "Completed video data" refers to the playable target video data obtained after all editing processes, including material time extraction, image compositing, text information overlay, and audio level adjustment.

[0294] "Summary information" refers to the summary attributes related to the completed video data, including brief descriptive information such as duration, resolution, file size, and cover image.

[0295] "Storage area" refers to the physical or logical storage space used to store source data, composition data, complete video data, and related metadata.

[0296] "Viewing or acquiring" refers to the act of playing, previewing, or downloading and saving completed video data on the terminal side.

[0297] "Historical information of generation and processing" refers to the collection of relevant data generated and recorded during the video generation process, such as prompts, video composition data, editing process, running status, and result quality.

[0298] "Input format" refers to the organization and encoding of instructions, parameters, and context information provided to generative artificial intelligence models.

[0299] "Video composition data generation rules" refers to the set of mapping logic, constraints, and priority strategies used when generating video composition data from prompt statements and material information.

[0300] "Short-duration distribution videos" refer to video content that is mainly used for distribution on online or mobile platforms and has a short playback duration and a fast pace.

[0301] "Material time interval" refers to the start and end time intervals pre-selected from the material data as candidate segments.

[0302] "Screen transition interval" refers to the time interval or editing rhythm between two adjacent screen segments in a video.

[0303] "Information content" refers to the density or richness of visual and textual content presented within a specific time interval of a video.

[0304] "Relevance" refers to the degree of semantic, thematic, or contextual relevance between unedited video, image, and audio data and prompts.

[0305] "Material priority" refers to the order in which each material is selected and used when generating video composition data, based on its relative importance.

[0306] The embodiments of the present invention will be described below with reference to a typical server-terminal structure. Each embodiment can be used individually or in any combination, as long as they do not contradict each other.

[0307] In this description, the "server" performs the main data processing and video generation operations, the "terminal" performs the user interface display and data input / output, and the "user" interacts with the server through the terminal.

[0308] I. Overall System Composition The server includes: at least one multi-core central processing unit (CPU), at least one graphics processing unit (GPU), main memory, non-volatile storage device (e.g., solid-state drive), network interface, and operating system. The server runs several software modules on the operating system, including: a web server module, an application server module, a database management system, object storage service, video processing program, and generative artificial intelligence model inference service.

[0309] The terminal includes: a processor, a display device, an input device (touchscreen, keyboard, mouse, etc.), local storage, and a network interface. The terminal runs a browser or a dedicated application and communicates with the server securely via Hypertext Transfer Protocol (HTTP).

[0310] The server can use general-purpose web server software as a front-end access component, such as using HTTP server software to receive HTTPS requests from terminals; it can use a relational database management system to store structured data; and it can use an object storage system to store large media files. The server can use a video processing program to perform material editing, transcoding, compositing, and subtitling, such as command-line-based video processing tools, or control desktop video editing software via a script interface. The server can deploy generative artificial intelligence models in an inference environment that supports CUDA or other GPU-accelerated frameworks.

[0311] II. Program Module Composition and Data Structure 1. Terminal-side interface module The browser or application running on the terminal loads the front-end page provided by the server. The terminal's interface modules include: a media upload interface, a prompt input interface, a parameter setting interface, a generation status display interface, and a video preview interface.

[0312] The terminal guides users to select local media files through a graphical user interface and uploads them to the server through a multi-part form; the terminal receives natural language prompts from users through text input boxes; the terminal receives parameters such as target playback duration, aspect ratio, and type of publishing media through drop-down boxes, radio buttons, etc.

[0313] After the terminal completes the material upload, it packages the material identifier obtained along with the prompts and parameters entered by the user into a structured request and sends it to the server via HTTPS.

[0314] 2. Server-side media management module When the server receives materials uploaded by the terminal, it calls the operating system interface to create files on the disk or object storage. The server extracts metadata of the video files through video processing programs, such as resolution, frame rate, encoding format, and recording duration; it extracts audio sampling rate, number of channels, and duration through audio analysis modules; and it parses image resolution and color space through image processing libraries.

[0315] The server stores the parsed structured metadata, such as "data type," "recording duration," and "image attributes," in a database. The server assigns a unique identifier to each piece of material and associates it with the user account, upload time, and other information. This structured management approach allows subsequent generative AI models to directly reference the searchable metadata when constructing video component data, avoiding the need to repeatedly scan the original files each time a new video is generated. This reduces computational overhead and improves response speed.

[0316] 3. Server-side prompt statement parsing and structured representation module After receiving the prompts and parameters from the terminal, the server stores them in the "Generation Request" table in the database. The prompts are saved in raw text form, and the server also creates a "Video Design Information" data structure to store the target playback duration, target aspect ratio, expected segment structure, key content, target audience type, etc., parsed from the prompts.

[0317] The server can use rule-based preprocessing components to initially segment the prompt statements, such as dividing them into multiple clauses according to line breaks, periods, or conjunctions, so that generative AI models can better recognize the structure. The server can also override or constrain ambiguous expressions in the prompt statements based on parameters explicitly input by the terminal (such as the target playback duration).

[0318] 4. Server-side generative artificial intelligence model module The generative AI model deployed on the server is a deep learning-based natural language processing model. This model can employ a multi-layered Transformer structure, including word embedding layers, multi-head self-attention layers, feedforward network layers, and an output layer. During the training phase, the server uses a large amount of video caption text and corresponding timeline annotation data as training data to supervise the model's learning.

[0319] During training, the server encodes the prompts and source metadata together into an input sequence. For example, the server converts the words in the prompts into vector representations and encodes the recording duration, aspect ratio, and other data of the source material into numerical features, which are then input into the model through feature concatenation or additional labeling. During training, the server uses the cross-entropy loss function to measure the difference between the predicted sequence and the standard video labels, and uses gradient descent algorithms combined with adaptive learning rate optimization methods to update the weights in the model parameters. The server can employ data augmentation strategies during training, such as making small perturbations to the time interval or using synonyms to rewrite the prompts, to improve the model's robustness to different instruction expressions.

[0320] During the inference phase, the server uses a generative AI model to parse new prompts. The model outputs structured video composition data, which can be described internally using a specific output format. For example, the model sequentially provides segment tags, time start offsets, durations, required material types, subtitle text, transition types, and audio volume tracks in the output sequence, which the server then reconstructs into its internal data structure.

[0321] The server can control the model's sampling parameters during inference, such as temperature, top-k, and top-p, to achieve a balance between diversity and determinism. For predictions of time intervals and media types, the server can employ constrained decoding strategies, forcing the total duration to not exceed the upper limit of the target playback duration, or prioritizing media with sufficient duration.

[0322] This generative AI model-based parsing method differs from traditional rule engines. By introducing timeline and material attribute labels during training, the server enables the model to learn how to allocate time and materials reasonably under different prompts. The model internally forms implicit representations of abstract concepts such as "rhythm," "information density," and "transition position," thereby automatically generating composition schemes that conform to video editing habits but are not limited to fixed templates during inference.

[0323] 5. Server-side duration consistency verification and adjustment module After receiving the video composition data, the server calculates the start and duration of each time interval. The server can then accumulate the lengths of each time interval segment by segment, calculate the total, and compare it with the target playback duration. If a discrepancy occurs, the server can automatically scale or offset a portion of the time interval according to a pre-set adjustment strategy.

[0324] The server's adjustment strategies can include: prioritizing the adjustment of the duration of non-critical information segments; maintaining the relative order between segments; and limiting the minimum and maximum duration of a single time interval. Through this numerical calculation and constraint solving method, the server eliminates potential small deviations in the output of the generative artificial intelligence model, ensuring that the final playback duration of the video precisely meets the target value within a given error range.

[0325] The server also verifies the "reference range of the material" based on the "recording duration" of the material. The server reads the total duration of the video or audio material and determines whether the start and end times of the references in the data fall within a valid range. If a reference exceeds the actual material's range, the server can automatically shorten that reference or search for similar alternative materials in the material library. Through this automatic verification and correction, the server avoids generating invalid editing instructions, reducing rendering failures and manual intervention.

[0326] 6. Server-side editing workflow definition and video processing module The server converts the video composition data into "editing workflow definition data." This data can be in the form of an internal JSON structure or an external project file, specifying the input materials, output time positions, subtitles, images, logos to be overlaid, and transitions, filters, volume curves, etc. to be applied for each time interval.

[0327] The server inputs the editing workflow definition data into the video processing program. This program can be a command-line tool or a desktop editing program that can be controlled by scripts. The server issues a series of specific instructions through command-line parameters or a script interface, such as cutting source files, merging timelines, overlaying subtitle tracks, overlaying logo layers, and mixing multiple audio tracks. In this process, the server does more than simply automate human editing steps; it automatically generates a large number of detailed timelines that are difficult for human editors to plan quickly and manually, following a structured workflow determined by a generative artificial intelligence model and a numerical verification module.

[0328] The server can enable multi-threading or parallel processing during video processing, parallelizing audio processing, subtitle rendering, and video encoding to fully utilize the computing power of multi-core CPUs and GPUs, achieving a significant speed improvement compared to manual editing and serial encoding.

[0329] 7. Server-side historical information recording and rule update module After each video generation task is completed, the server saves the prompts, video composition data, editing workflow definition data, rendering time, error logs, etc., as "generation processing history information". The server can periodically analyze this history information, for example, by using statistical methods or re-invoking the training pipeline to retrain or fine-tune the generative artificial intelligence model.

[0330] The server can identify mapping patterns between certain types of prompts and user-satisfied video compositions based on historical data, and adjust the input format accordingly. For example, it can add explicit identifiers, and the prompt model can more clearly distinguish fields such as "duration requirements," "style requirements," and "structural requirements." The server can also automatically update video composition data generation rules, such as shortening the maximum duration of a single shot and increasing the shot switching frequency in short video scenarios, thereby adapting to constantly changing application needs.

[0331] This adaptive update mechanism based on historical information enables the server to continuously improve the timeline accuracy and content matching of the generated videos over a long period of time, thereby improving the performance of the entire computer system in automatic video generation tasks, rather than just executing a fixed algorithm once.

[0332] 3. Specific use cases 1. Examples of short product promotion videos Users access the system page via the terminal and upload product demonstration video files, factory environment video files, brand logo image files, and background music / audio files. Users enter the following prompt in the terminal's text input box: "Generate a 60-second product promotional video:" The first 10 seconds showcase the factory environment, accompanied by the brand slogan subtitles; The middle 40 seconds are dedicated to a product demonstration video, divided into three parts: 'appearance design', 'core functions', and 'after-sales service'. The last 10 seconds showcase the brand logo and promotional text, making it suitable for posting on short video platforms. The terminal sends the prompt message along with parameters such as the target playback duration of 60 seconds, portrait screen ratio, and type of media to the server. The server's generative AI model analyzes the prompt message and, combined with the material metadata, generates three main time intervals: 0-10 seconds for the opening segment, 10-50 seconds for the product description segment, and 50-60 seconds for the ending segment; the product description segment is further subdivided into three sub-segments, each approximately 13 seconds long. The server then performs a duration consistency check, proportionally shortening each sub-segment if the total exceeds 60 seconds. The server also checks whether each referenced segment is within the recorded duration of the corresponding material; if any segment exceeds the material's duration, it automatically shortens or replaces the material.

[0333] The server generates the corresponding editing workflow definition data, calls the video processing program to complete editing, image compositing, and audio mixing, and generates the final video data. Users receive the video link provided by the server on their terminals, which they can then play or download for use.

[0334] 2. Examples of short travel videos The user uploads three travel video clips on the terminal and enters a prompt message: "Please use the three travel videos I uploaded to create a 30-second travel vlog." The destination title and date are displayed for the first 5 seconds. The program automatically selects the most representative landscape and portrait shots to alternate between each other for 20 seconds, and adds short captions for the corresponding locations (such as 'Sunrise at the Seaside' or 'Street in the Ancient City'). The last 5 seconds include a caption saying "Follow me for more travel tips," accompanied by light and cheerful background music. When parsing prompts, the server outputs structured video composition data through a generative artificial intelligence model, automatically assigning source material and subtitle content to each time interval. During video processing, it selects segments with rich visual changes and frequent movement from different materials as "representative shots." Compared to manual editing, the server's ability to process large amounts of material in a short time is fully utilized, allowing for more shot combinations to be tried within the same time budget, thus improving the overall quality of the video.

[0335] 3. Examples of product unboxing short videos The user enters the following prompt in the terminal: "Generate a 15-second product unboxing video."

[0336] The brand logo is displayed for the first 2 seconds; The middle 10 seconds should be a quick cut of the key footage from the unboxing process, with the editing points following the rhythm of the background music as much as possible, and the pace of the footage should be fast. The final 3 seconds feature an overlay of the "Buy Now" text in a CTA format, with a slight enlargement of the product image. The server incorporates rhythm-related features into the generative AI model, such as generating a list of candidate cut points based on the energy peaks or beat detection results of the audio track, and using these time points as candidate cut points in the composition data. During the composition stage, the server uses non-linear optimization or heuristic algorithms to align the camera cuts as close to the beat as possible, thereby generating visually more rhythmic edits. This approach goes beyond the traditional method of simple average segmentation, demonstrating a technological improvement in computer rhythm optimization.

[0337] IV. Technical Effects and Improvements in Computer Technology The server achieves the following technical improvements through the aforementioned methods, including structured data representation, generative artificial intelligence model parsing, duration consistency verification, and automatic editing process generation: 1. Improved processing speed: The server avoids repeated file scanning by extracting and caching source metadata; it reduces disk I / O and process creation overhead by using parallel execution and batch command calls during video processing; and it uses constraint decoding when generating data to reduce the exploration space of invalid candidates, thus significantly shortening the overall computation time for generating a complete video.

[0338] 2. Precision improvement and error control: The server precisely controls the sum of time intervals through numerical calculations and performs boundary checks on the scope of material references to prevent blank segments or out-of-bounds access. Generative AI models incorporate timeline labels and material attributes during training, making it easier to generate solutions that meet technical constraints during inference. Compared to systems that only predict descriptive scripts, timeline accuracy and structural consistency are improved.

[0339] 3. Data Management and Scalability Improvements: The server employs a unified data structure to manage prompts, video design information, video composition data, and editing workflow definition data. This allows the processing results at each layer to be recorded and analyzed, providing a reliable data foundation for subsequent model retraining and rule optimization. This layered structure facilitates the introduction of new generative AI models or the replacement of video processing programs, enabling the system to maintain overall stability with only minor adjustments at the interface layer.

[0340] 4. Communication and computing load optimization: The server transmits only necessary metadata and the completed video on the client side, while a large amount of intermediate and constituent data remains internal to the server, reducing the scale of network transmission. When scheduling generative AI model inference tasks, the server can use task queues and batch processing strategies to process multiple requests at once on the GPU, improving hardware utilization and reducing the marginal computational cost of a single request.

[0341] 5. Unique processing methods that differ from human work: The server uses a generative artificial intelligence model to parse prompts, moving beyond pre-written templates to learn complex timelines and material selection rules during training, achieving compositional generation based on statistical learning. The server automatically corrects timeline deviations and material citation errors through mathematical constraints and optimization strategies, forming a set of "machine-defined rules" independent of human experience. In scenarios such as rhythm alignment, the server generates editing schemes based on signal features through audio feature analysis and editing point optimization—a process that traditional subjective human decision-making struggles to achieve in a timely manner.

[0342] In summary, this invention, by introducing generative artificial intelligence model analysis, structured video composition data, automatic verification and correction, editing workflow definition, and historical information-driven rule updates on the server side, transforms the automatic video generation process from simple human task automation into a systematic improvement of computer video processing capabilities. This achieves an overall improvement in processing speed, timeline accuracy, material utilization efficiency, and model invocation stability. Users can leverage the system's technological advantages to generate high-quality videos that meet specific constraints simply by inputting prompts and uploading materials at the terminal.

[0343] use Figure 13 The processing flow is explained.

[0344] Step 1: Users access the system interface and upload materials using a terminal.

[0345] The user opens a browser or application on the terminal and accesses a web address provided by the server. Within the terminal interface, the user selects unedited video, image, and audio files from their local device using file selection controls. The terminal takes these files as input, packages them into multi-part form data, and sends it to the server via HTTPS. During the sending process, the terminal displays an upload progress bar and, upon completion, receives a list of media identifiers returned by the server as output.

[0346] Step 2: The server receives the materials and parses the metadata.

[0347] The server receives upload requests from terminals and writes video, image, and audio data from multiple forms into a storage device. Using the source file data and request header information as input, the server calls video processing programs and image processing libraries to perform data parsing operations on each file: the server parses resolution, frame rate, and recording duration from the video stream; it parses sampling rate, number of channels, and recording duration from the audio stream; and it parses resolution and color space from the image data. The server writes the parsing results along with the source file paths to the database, generating metadata records containing source IDs, data types, recording durations, and image attributes as output.

[0348] Step 3: Users enter prompts and video generation parameters on the terminal.

[0349] Users input natural language prompts in the text input boxes on the terminal interface, such as "Generate a 60-second product promotional video: the first 10 seconds showcase the factory environment with brand slogan subtitles; the middle 40 seconds feature a product demonstration video, divided into three segments explaining 'appearance design,' 'core functions,' and 'after-sales service'; the last 10 seconds display the brand logo and promotional copy, suitable for publishing on short video platforms." Users select the target playback duration, aspect ratio, and publishing media type in the terminal's parameter settings area. The terminal uses the above prompt text and parameter values ​​as input to construct a request data structure containing the prompt text, a list of material IDs, and video parameters, and sends it to the server via HTTPS. The terminal receives the task ID or confirmation information returned by the server as output.

[0350] Step 4: The server stores the prompts and generates video design information.

[0351] The server receives a request containing prompts and parameters, and writes the original prompt text along with the user-specified target playback duration, aspect ratio, and media type into a generation task record in the database. Taking the prompt text and parameters as input, the server performs preliminary text parsing: splitting the video into clauses based on line breaks, periods, or semicolons, identifying time-related words (e.g., "first 10 seconds," "last 10 seconds") and structural expressions (e.g., "three segments," "beginning," "middle," "end"). Based on these parsing results, the server generates a structured video design information data structure, including the expected number of segments, a draft of the target duration for each segment, and key content tags, and stores this structured data along with the original prompt text as output.

[0352] Step 5: The server invokes a generative artificial intelligence model to generate initial video composition data.

[0353] The server reads prompts and video design information from the database, using these as input to construct the model's input sequence. The server segments the prompts into words and embeds them as vectors, encoding the target playback duration, aspect ratio, and available material metadata as numerical features, forming a unified input tensor. This input tensor is fed into a generative AI model deployed on a GPU, performing matrix operations with multi-layered self-attention and feedforward networks to infer the start time, duration, material type, subtitle text, image compositing information, and audio processing information for each time interval. The server parses and decodes the model's output labeled sequence, converting it into a structured video composition data object, and writes this object into the database as the initial video composition data output.

[0354] Step 6: The server performs timeline consistency checks and duration adjustments.

[0355] The server takes the initial video composition data and the target playback duration as input, accumulates the duration of each time interval, and calculates the total duration. The server compares the total duration with the target playback duration and performs adjustments according to a preset strategy: for example, the server proportionally scales the duration of non-critical segments or simultaneously shortens the duration of all sub-segments; after adjustment, the server recalculates the start time to ensure that time intervals are continuous and non-overlapping on the timeline. The server saves the adjusted video composition data that meets the target playback duration constraint as output and updates the composition records in the database.

[0356] Step 7: The server verifies and corrects the citation range based on the duration of the material records.

[0357] The server reads the media ID and start and end times referenced for each time interval. Using video composition data and media metadata as input, it retrieves the corresponding media record duration from the database. The server performs boundary checks to determine if the end time of each reference is less than or equal to the media record duration. If a segment is found to be outside the range, the server automatically calculates the excess, shortens the reference duration, adjusts the start time, and replaces it with similar media if necessary. The server writes the corrected time intervals back to the video composition data structure, forming composition data that has passed media boundary checks as output.

[0358] Step 8: The server generates the editing workflow definition data.

[0359] The server takes the verified video data as input and generates specific editing instructions for each time interval. For each segment, the server generates corresponding cutting command parameters (source file path, start time, end time), overlay command parameters (subtitle content, font style, position, appearance time), compositing command parameters (position and transparency of overlaid images or logos), and audio filtering parameters (volume gain, fade-in / fade-out time). The server organizes these parameters into a unified editing workflow definition data structure, such as a list of sequentially executed nodes or a graphical workflow description, and saves this definition data to a storage device as input for subsequent video processing programs and output for this step.

[0360] Step 9: The server performs video editing and compositing.

[0361] The server takes the editing workflow definition data and original source files as input and sequentially sends editing and compositing commands to the video processing program. Based on the instructions for each time interval, the server calls the video processing program to perform video segment extraction, timeline splicing, subtitle rendering, image overlay, and audio mixing operations. Internally, the server uses queue scheduling and multi-threading control to distribute the processing tasks of different video segments in parallel to multiple sub-processes or threads, thereby shortening the overall processing time. Under the server's control, the video processing program outputs a completed video file that has undergone editing, compositing, subtitle overlay, and audio processing. The server writes the storage path and related metadata of this file into the database as the output of the completed video data.

[0362] Step 10: The server provides the complete video data to the terminal.

[0363] The server takes the storage path and task ID of the completed video data as input and generates an accessible download or playback link. When the server receives a query request from the terminal, it retrieves the status and video link of the corresponding task from the database and returns the link address, video duration, resolution, and thumbnail information as output to the terminal. Based on the information returned by the server, the terminal displays a "Generation Complete" status, a play button, and a download button on the interface. When the user clicks play, the terminal loads the video stream and plays it, using the video link as input.

[0364] Step 11: The server records the generation history and updates the model and rules.

[0365] After completing a task, the server takes the prompts, final video composition data, editing workflow definition data, rendering time, and error logs as input and writes this data to a historical information storage table. The server periodically reads this historical information and performs statistical analysis, such as calculating the average duration deviation, failure rate, and user feedback ratings for different prompt patterns. Based on the analysis results, the server adjusts the input format and hyperparameters of the generative AI model, or updates the video composition data generation rules (such as the maximum shot length for short videos and subtitle density). The server saves the updated model parameters and rule configurations as output for subsequent video generation tasks, gradually improving timeline accuracy, processing efficiency, and generation quality.

[0366] Application Example 2 The process flow corresponding to the specific processing in Use Case 2 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. In addition, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".

[0367] With the rapid growth in demand for multimedia content production and distribution, ordinary users hope to automatically obtain high-quality, emotion-matched personalized videos simply by uploading original materials and providing a brief description of their intent. However, existing technologies suffer from the following problems at the computer implementation level: (1) In computing devices, the processing of unedited multimodal information data (moving images, still images, audio) is usually fragmented. The lack of a unified computing process organization for steps such as material analysis, script generation, editing decisions, and rendering encoding makes it difficult for the server side to automatically complete the entire process from "user intent" to "playable video data" in an end-to-end pipeline, increasing processing complexity and resource consumption.

[0368] (2) Most existing solutions utilizing generative artificial intelligence models only use the model to generate text scripts or simple descriptions, without tightly coupling the "construction logic of natural language prompts" with the calculation results such as "underlying material features" and "playback constraints" at the system level. When generating prompts, the structured features and quality indicators obtained from image processing and audio processing are not fully utilized, making it difficult for generative artificial intelligence models to produce composition schemes that are highly consistent with the material and duration constraints, thereby reducing the efficiency and controllability of automatically generated videos.

[0369] (3) Traditional automatic video editing systems often treat emotion analysis as an additional function, simply adding filters or adjusting music after the video is finished. There is a lack of fine-grained computational coupling between emotion information and timeline editing decisions. The server cannot automatically allocate differentiated tone changes, visual effects, and background sound modes at the editing timeline level based on the content characteristics of different scene segments and the user's emotional state. This results in inconsistent overall emotional expression in the generated results and makes it difficult to efficiently personalize the video for different users.

[0370] (4) In resource-constrained server environments, there is a lack of a unified, programmable data flow and control flow structure to integrate material analysis, prompt generation, generative AI model invocation, emotion-driven expression parameter decision-making, and final encoding and storage. As a result, computing resources are unevenly allocated, there is a lot of redundant computation, and the system has limited throughput when facing massive requests, affecting scalability and response speed in actual deployment.

[0371] Therefore, it is necessary to provide a technical solution implemented on a server, which organically combines multimodal feature extraction of materials, automatic construction of prompts based on abstract information and constraints, generation of video composition schemes driven by generative artificial intelligence models, and automatic decision-making of timeline-level performance parameters driven by emotion state within the computer using a unified data structure and processing pipeline. This improves the controllability, efficiency, and personalized expression capabilities of the automatic video generation process at the computer technology level.

[0372] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 2 is achieved by the following means.

[0373] In this invention, the server includes: means for receiving and storing unedited information data; means for receiving and storing abstract information and constraints related to the generation of target video content; means for performing image and audio processing on the information data to extract scene segments and calculate content feature quantities and quality indicators for each scene segment; means for generating prompt statements in natural language form based on the abstract information and constraints corresponding to the generated target video content and inputting the prompt statements into a generative artificial intelligence model to obtain a video composition scheme; and means for determining the appropriate composition scheme based on the video composition scheme, the content feature quantities, and the quality indicators. The device comprises: a device for obtaining scene segments and their durations and generating editing timeline information; a device for acquiring emotion-related data related to the user and processing the emotion-related data to estimate the emotional state; a device for determining video performance parameters, including hue, brightness, visual effects, music type, and volume balance, based on the emotional state and the video composition scheme, and associating the video performance parameters with the editing timeline information to set visual and auditory effects; and a device for performing dynamic image processing and audio synthesis processing based on the editing timeline information and the video performance parameters to integrate multiple information data to generate final video data, and storing and assigning matching credits to the final video data. This allows for the formation of an end-to-end computational pipeline within the server. This pipeline encompasses receiving multimodal raw information data, calculating feature quantities and quality indicators, automatically generating prompts based on abstract information and constraints, driving a generative AI model to produce video composition schemes, and then dynamically determining video performance parameters at the timeline granularity based on emotional states and completing encoding and storage. This approach automatically generates personalized videos that meet duration constraints, content selection criteria, and individual emotional needs with relatively low computational resource consumption. Consequently, it substantially improves the processing efficiency, controllability, and scalability of computers in the fields of automatic video generation and multimodal emotion adaptive synthesis.

[0374] "Unedited information data" refers to multimedia digital data provided by users that has not been edited, synthesized, or processed with effects, including but not limited to moving image data, still image data, and audio data.

[0375] "Abstract information" refers to high-level information used to describe the overall intent and semantic requirements for generating target video content, including non-material-level descriptions such as theme, style, purpose, plot outline, and target audience.

[0376] "Constraints" refer to the structured restrictions imposed on the system when generating the target video, including parameters such as target playback duration, resolution, frame rate, content ratio, pacing requirements, and platform specifications.

[0377] Image processing refers to computational operations performed on dynamic or static image data, including frame extraction, scene segmentation, object detection, feature extraction, color analysis, and quality assessment.

[0378] "Audio processing" refers to computational operations performed on audio data, including feature extraction, rhythm analysis, spectrum analysis, volume statistics, calculation of emotion-related indicators, and quality assessment.

[0379] "Scene segment" refers to a time interval with relatively unified content or continuous visuals, defined by timeline analysis and shot boundary detection of dynamic image data.

[0380] "Content features" refer to numerical descriptions used to characterize the attributes of scene fragments or image content, including object categories, scene categories, action types, color distribution, composition features, and semantic tags.

[0381] "Quality metrics" refer to numerical parameters used to quantify the quality of scene segments or images in terms of sharpness, stability, noise level, exposure, and compositional integrity.

[0382] "Construction information" refers to intermediate representation data used to describe the internal structure of a target video, including scene order, approximate duration of each scene, key content types, and transition methods.

[0383] "Prompt statements" refer to instruction text expressed in natural language and used as input to generative artificial intelligence models. They contain abstract information, constraints, and guidance related to the characteristics of the material.

[0384] "Generative artificial intelligence models" refer to machine learning models that can automatically generate text, structured schemes, or other data outputs based on input prompts, including but not limited to large-scale language models and multimodal generative models.

[0385] "Video composition scheme" refers to the overall structural plan of the target video, which is output by the generative artificial intelligence model and analyzed by the system. It includes information such as the sequence of segments, the suggested duration of each segment, the key points of the content, the emotional direction, and the suggested background music.

[0386] "Edit timeline information" refers to structured data that arranges scene segments, transition effects, audio tracks, and related parameters on the timeline to guide subsequent compositing and encoding processes.

[0387] "Emotion-related data" refers to input data that is related to the estimation of a user's emotional state, including but not limited to text input, voice signals, facial images, historical interaction records, and physiological signals.

[0388] "Emotional state" refers to the type and intensity of a user's psychological emotions within a given time period, inferred from emotion-related data through a computational model. Examples include happiness, sadness, tension, and calmness.

[0389] "Video performance parameters" refer to a set of adjustable parameters used to control the final visual and auditory presentation of the video, including hue, brightness, contrast, saturation, filter type, visual effects, music type, and volume balance.

[0390] "Visual effects" refers to the visual presentation in video output through techniques such as color transformation, filter processing, transition effects, and overlay elements.

[0391] "Audio effects" refers to the overall sound presentation in video output through background music, ambient sound effects, narration audio tracks, and their mixing methods.

[0392] "Dynamic image processing" refers to processing operations performed on time-series image data, including cropping, splicing, scaling, stabilization, speed transformation, filter overlay, and transition effect compositing.

[0393] "Audio synthesis processing" refers to the process of performing various processing steps on multiple audio data, including alignment, trimming, mixing, volume adjustment, panning, and sound effect overlay, to generate a target audio track.

[0394] "Final video data" refers to the digital video file or video data stream that is ready for playback or distribution after timeline editing, visual effects settings, audio effects settings, and encoding processing have been completed.

[0395] "Assigned credit identifier" refers to the identification information used to uniquely identify the final video data so that it can be retrieved, accessed and managed in storage systems, transmission systems or distribution platforms, including but not limited to Uniform Resource Locator, Object Identifier or Internal Index Number.

[0396] In one embodiment of the present invention, the system includes a server, a terminal, and a user operating through the terminal. The server may be an information processing device having a multi-core general-purpose processor, a graphics processing unit, and persistent storage, and the terminal may be a mobile communication device or a computing device having a display unit, a camera unit, an audio acquisition unit, and a communication unit.

[0397] At the hardware level, a server may include: at least one general-purpose central processing unit, at least one graphics processing unit for deep learning inference, a semiconductor storage device for storing programs and data, interface circuitry for accessing external storage systems, and a network interface controller. At the software level, the server may run an operating system, network communication middleware, multimedia processing libraries, deep learning inference frameworks, and a database management system. Specific software examples include, but are not limited to: multimedia processing libraries for video and audio transcoding and encapsulation, image processing libraries for image and video frame processing, deep learning frameworks for building convolutional neural networks and recurrent neural networks, text processing libraries for natural language analysis, and statistical model libraries for sentiment analysis.

[0398] At the hardware level, the terminal may include: a display unit for showing a graphical user interface, a camera unit for capturing still and moving images, a microphone for capturing audio, a wireless communication module for local or wide area network communication, and a local storage device. At the software level, the terminal may execute applications for displaying a user interface, capturing materials and metadata, and communicating with a server.

[0399] Users perform the following operations through the terminal. After launching the multimedia generation application on the terminal, users select or capture unedited information data in the terminal's graphical user interface. Users can select at least one unedited video file, at least one image file, and at least one audio file as information data. Users further input abstract information on the terminal interface, such as theme descriptions like "adventure travel," "family memories," or "product introduction advertisement," or style descriptions like "bright and lively" or "nostalgic with a touch of sadness," or usage descriptions like "for sharing on social networks" or "for short video advertising." Users can also input their current mood via text (e.g., "I'm very happy lately," "I'm a little sad"), or provide a facial image through the terminal's camera, or provide voice through the microphone, to input emotion-related data into the server.

[0400] After receiving unedited data uploaded by the terminal, the server uses a multimedia processing library to perform format and integrity checks on the data. The server can call the multimedia processing library to decapsulate and transcode video, converting different video formats to a unified encoding format, and unifying the sampling rate and encoding format of audio, ensuring that subsequent access to video and audio frames uses a consistent data structure. The server stores the converted data in persistent storage and registers metadata attributes such as path, duration, frame rate, resolution, and recording time for each data item in the database.

[0401] During image and audio processing, the server uses an image processing library to extract frames from unedited video at fixed intervals. The server then uses a convolutional neural network (CNN)-based image recognition model to perform forward inference on each frame, obtaining feature vectors representing scene category, object category, and the presence or absence of faces. This CNN can employ a stacked structure of multiple convolutional layers, normalization layers, non-linear activation function layers, and pooling layers. In the offline phase, the server supervises the learning of this CNN using a large-scale labeled image dataset, updating the network weights through backpropagation and gradient descent by minimizing the cross-entropy loss function. In the online inference phase, the server uses frozen weights to perform forward computation on the input frames to output the category probability distribution and intermediate feature vectors.

[0402] At the video level, the server divides the unedited video into multiple scene segments by detecting visual differences between adjacent frames or based on a trained shot boundary detection model. For each scene segment, the server calculates content features, including but not limited to scene category histograms, object category histograms, average brightness, color saturation, and motion intensity. Simultaneously, the server calculates quality metrics for each scene segment, such as estimating sharpness based on image gradient statistics, assessing jitter based on optical flow stability, and estimating exposure based on brightness distribution. In the audio processing section, the server uses an audio analysis library to perform a short-time Fourier transform on the audio signal, extracting spectral features, rhythm features, and energy envelopes for subsequent matching of video rhythm and emotion.

[0403] In the abstract information processing and sentiment analysis sections, the server uses a text processing library to perform word segmentation and part-of-speech tagging on the abstract information input by the user, and extracts topic keywords, style keywords, and usage keywords. The server further uses a sentiment analysis model to classify the emotions in the text and speech-to-text. This model can be a text classification network based on a bidirectional recurrent neural network or an attention mechanism. During the training phase, the server uses text and speech-to-text data labeled with emotion tags and optimizes the network parameters based on the classification error. For facial expression analysis, the server uses a facial expression recognition model based on a convolutional network to extract and classify features from the user's facial region image, obtaining probability distributions for emotions such as joy, sadness, anger, and surprise. The server weighted and fused the emotion results from the text, speech, and image sources to generate a comprehensive emotion state vector, which serves as a control signal in subsequent performance parameter decisions.

[0404] When generating prompts, the server constructs natural language instructions based on abstract information, constraints, content features and quality metrics of scene segments, and overall emotional state. During this process, the server uses a specific data structure to organize information such as scene category distribution, quality ranking, segment duration statistics, and the user's target playback duration into an intermediate structure. This intermediate structure is then mapped to natural language descriptions to form the prompts. The prompts include not only high-level instructions such as "generate a video of how long" and "what theme and style," but also fine-grained control information such as "prioritize segments containing smiley faces and with high image quality" and "use strong music during climaxes."

[0405] For example, in a travel short film scenario, the server can generate the following prompts as input for a generative artificial intelligence model: "Based on the analysis of the following materials, design a 2-minute adventure travel video composition plan. The materials include beach, mountain road, and city night scene scenes. The high-quality clips are mainly concentrated on the beach at sunset and the city streets at night. Please prioritize clips that include smiling faces and stable footage. Start with a 10-second wide shot of the beach to introduce the video, focus on the mountain road walking and city night scene in the middle, and end with a close-up of the beach at sunset. The overall style should be bright and energetic, and suitable for a rhythmic background music." For example, in a family memory scenario, the server can generate the following prompt: "Based on photos and video footage of a family gathering, please design a 3-minute heartwarming family memory video script. Highlight moments of family members smiling, hugging, and eating together. Use soft, warm colors and upbeat music. Arrange the scenes chronologically from the beginning to the end of the gathering, and provide short text suitable for narration." After generating the prompt, the server inputs it into a generative AI model. This model employs a sequence generation network structure based on a multi-layered self-attention mechanism. During training, this network uses a large-scale text corpus to learn sequence distributions under natural language conditions. When calling the generative AI model, the server sets temperature parameters and sampling strategies to ensure that the output video composition schemes both conform to the prompt constraints and exhibit diversity. Upon receiving the model's output, the server parses it into a structured video composition scheme, including information such as scene segment order, suggested duration for each segment, plot logic, emotional fluctuations, and suggested background music.

[0406] When generating editing timeline information based on the video composition scheme, the server uses the content features and quality metrics of the aforementioned scene segments to map the abstract descriptions in the composition scheme to actual segments. At each logical location, the server selects the best segment from a set of candidate segments with corresponding scene tags and high quality metrics, and allocates the duration of each segment according to the target length. The server stores the source file identifier, start and end timestamps, and transition type of the selected segment into a timeline data structure, which can be a list or graph structure sorted by time. Through this feature-based and metric-based mapping, the server reduces the need for manual editors to review each piece of footage, improving the efficiency of automatic composition.

[0407] During the video performance parameter decision-making phase, the server uses the overall emotional state as control input, combined with the video composition scheme and timeline data, to assign specific color adjustment parameters, brightness adjustment parameters, filter type, special effects parameters, music type, and volume envelope to each scene segment. In the color adjustment section, the server can use lookup tables or parametric color transformation functions to adjust the overall saturation and contrast according to the emotional state; for example, increasing saturation and brightness in a "joyful" state, and decreasing saturation and introducing a color-cast filter in a "nostalgic" state. In the audio section, the server selects a background music track with a corresponding style based on the emotional state and plans the volume envelope according to the scene intensity, appropriately increasing the music volume at emotional climaxes and decreasing the volume in transitional segments to highlight the visual content.

[0408] During the actual video compositing process, the server utilizes a multimedia processing library to call the underlying encoder, performing operations such as cutting, scaling, stabilizing, speed transformation, and filter overlay on each segment of the timeline. The server then splices multiple processed segments in chronological order, inserting transition effects such as fade-in / fade-out and crossfade between segments. In the audio compositing section, the server synchronizes and mixes background music, ambient sound effects, and any narration tracks, using time alignment algorithms to ensure that visual and audio events correspond in time. Finally, the server inputs the complete timeline into the video encoder, encoding it into a target video file format that conforms to network distribution or storage specifications. The server assigns a credit identifier to this final video data and records the identifier and metadata information in the database.

[0409] In another implementation, the server can be specifically optimized for short-duration videos. In this implementation, the server uses the target playback duration as the core constraint and introduces a segment selection algorithm based on integer linear programming or a greedy strategy. The server calculates the "information density score" and "quality score" for each segment to maximize overall information content and quality within the constraint of total duration. Because the server explicitly incorporates the "target playback duration" and "segment selection strategy" into the generated prompts, the video composition scheme output by the generative AI model inherently tends to have a compact scene arrangement, thereby reducing subsequent adjustments and improving the overall system processing speed.

[0410] In another implementation, the server can apply rule constraints and consistency checks to the output of the generative AI model. After parsing the composition scheme, the server modifies or supplements the scheme based on predefined non-standard rules, such as: "two clips with resolution below the threshold cannot be used consecutively," "black screens cannot be inserted for more than a specified duration during climaxes," and "advertising videos must contain at least one product close-up clip." These rules do not simply reproduce manual editing; instead, they perform formal constraint checks within the computer based on content features and quality indicators, significantly reducing the proportion of substandard videos in large-scale batch generation scenarios.

[0411] Through its modular software and hardware architecture, the server integrates material analysis, prompt construction, generative AI model invocation, timeline composition decisions, and emotion-driven performance parameter determination into a unified processing pipeline. Because the server internally uses specific data structures to store content features and quality indicators, and encodes these features along with abstract information and constraints into prompts, the generative AI model receives richer and more structured contextual information during inference. Compared to the traditional method of inputting only brief text descriptions, the model's output video composition scheme shows higher consistency with the actual material, requires fewer corrections, and improves the overall computational efficiency of the generation process. Furthermore, by dynamically adjusting performance parameters at the timeline granularity and employing neural network-based emotion state inference, the server can perform automatic personalized rendering on a large number of user requests using a unified algorithm, reducing repeated rendering and unnecessary data transmission, thereby achieving higher throughput and lower communication load in real-world deployment environments.

[0412] In alternative implementations, the server can replace the generative AI model with other sequence generation networks or multimodal generation networks with similar functionality. The server can choose a network with a smaller parameter size to reduce inference latency, or a network with multimodal input capabilities to directly utilize visual and textual features, depending on the computational resources of the deployment environment. In all implementations, the server provides abstract information, constraints, and feature summaries to the generative AI model through a unified prompt interface, thus maintaining system architecture consistency. Similarly, in all implementations, the terminal presents material selection and intent input functions to the user in a graphical interface. Users do not need to understand the internal model details; they can obtain personalized final video data through natural language descriptions and simple operations.

[0413] Through the above-described embodiments, the server, terminal, and user collaborate to achieve automatic video generation based on unedited information data, abstract information, and constraints. By introducing specific data structures, model architectures, and rule constraints, this invention not only automates human editing work but also optimizes data flow and control flow within the computer, improving feature utilization efficiency, accuracy of scheme generation, and overall computing performance. This represents an improvement to multimedia processing and artificial intelligence computing technologies themselves.

[0414] use Figure 14 The processing flow is explained.

[0415] Step 1: Users select and collect materials and intent information using the terminal. Input includes video files, image files, audio files stored locally on the terminal, as well as text descriptions, emotion-related text, or voice input by the user in the interface. Users select or capture unedited videos, photos, and audio recordings in the graphical interface of the terminal application, and enter abstract information such as "adventure travel short film," "approximately 2 minutes long," and "bright and energetic style" in text boxes. The terminal performs preliminary validation on these files and texts (such as file size and format validity), and outputs a set of material data to be uploaded and an abstract information data structure.

[0416] Step 2: The terminal uploads materials and abstract information to the server. The input consists of the material data set and abstract information data structure generated in step 1. The terminal uses its network communication module and HTTP or HTTPS protocols to encapsulate the binary streams of video, images, and audio, along with the abstract information, into a request message and sends it to the server. The terminal includes metadata such as filename, duration, and resolution in the message. The data is processed into a network transmission format (e.g., chunked upload, content length marking), and the output is a network data stream that can be received and parsed on the server side.

[0417] Step 3: The server receives and stores unedited information data. Input consists of a multimedia binary stream and metadata received from the terminal. The server uses a network interface to parse HTTP / HTTPS messages, writes the binary content of each file to a temporary buffer, and then calls a multimedia processing library to detect the encoding format. If the encoding does not conform to a preset standard, the server calls a transcoding function to convert the video to a predetermined encoding format and the audio to a predetermined sampling rate and encoding. The server writes the transcoded files to a storage device and generates record entries in the database, recording the material identifier, storage path, duration, frame rate, resolution, etc. Output is a set of persistently stored unedited information data and corresponding database index information.

[0418] Step 4: The server performs content analysis on videos and images and extracts scene segments. The input consists of the unedited video and image files and their metadata stored in step 3. The server uses an image processing library to extract frames from each video at time intervals, inputs these frames into a pre-trained convolutional neural network classification model for forward inference, and obtains the scene category, object category, and intermediate feature vector for each frame. The server calculates similarity based on the feature differences between adjacent frames and applies a shot boundary detection algorithm to divide the video into several scene segments. The data processing includes: calculating feature vectors for frames, performing segmented clustering or thresholding along the time dimension, and outputting a list of scene segments for each video. The list includes the segment start and end times, representative frame feature vectors, scene labels, and initial values ​​of content features and quality indicators.

[0419] Step 5: The server performs feature extraction on the audio and generates audio feature data. The input is the audio file stored in step 3. The server uses an audio analysis library to perform frame segmentation and short-time Fourier transform on the audio signal, calculating the power spectrum, rhythm features, frequency band energy distribution, and pitch correlation features for each frame. The server also performs statistical analysis on the global rhythm (BPM), energy fluctuation patterns, etc., for the entire audio segment. The data processing involves transforming the time-domain waveform to the frequency domain and statistical features, and the output is an audio feature sequence with time stamps, as well as overall rhythm and energy feature parameters.

[0420] Step 6: The server parses abstract information and generates a theme and constraint structure. The input is the abstract information text uploaded in step 2 (including theme, style, purpose, target duration, etc.). The server uses a natural language processing library to perform word segmentation, part-of-speech tagging, and entity recognition on the text, extracting theme keywords (such as "travel," "family," "advertisement"), style tags (such as "bright," "nostalgic"), purpose tags (such as "social media sharing"), and numerical type constraints (such as converting "2 minutes" to 120 seconds). Data processing includes key phrase extraction and type identification, outputting a structured theme and constraint data structure, which includes target playback duration, target audience, style weights, etc.

[0421] Step 7: The server infers user emotional states based on multi-source data. Inputs include user text expressions, speech signals (or their transcribed text), and user facial images or video frames containing faces. The server first uses a speech recognition module to convert the speech signals into text, then inputs the text into an emotion classification network and the image into an expression recognition network; these networks are pre-trained neural network models. The server calculates probability distributions for text emotions (e.g., "happy," "sad"), speech tone emotions, and facial expression emotions, and uses weighted averaging or Bayesian fusion algorithms to obtain a comprehensive emotion vector. The data processing involves forward inference and probabilistic fusion of multi-channel features, and the output is an emotion state vector representing the emotion type and its intensity.

[0422] Step 8: The server constructs prompts and invokes a generative AI model. The inputs are the scene segment content features and quality metrics obtained in step 4, the audio features from step 5, the topic and constraint data structure from step 6, and the emotion state vector from step 7. Based on this data, the server generates an intermediate summary (e.g., a list of available segments, segment ratings, target duration, and expected emotional trajectory), and then converts this summary into prompts in natural language. The server sends the prompts as input to the generative AI model, setting the generation length and diversity parameters. The generative AI model performs sequence generation operations based on the prompts, outputting a video composition scheme text. The data processing involves a bidirectional mapping from structured features to natural language and then to structured script, with the output being a composition scheme text describing the scene sequence, the duration of each suggestion, the plot structure, and music suggestions.

[0423] Step 9: The server parses the output of the generative AI model and generates a video composition scheme structure. The input is the composition scheme text generated in step 8. The server uses a natural language processing module to segment the text into sentences and perform pattern matching, identifying the logical position (beginning, middle, end) of each scene segment, the required scene type, suggested duration, and emotional emphasis. The server organizes the recognition results into structured data, such as a list of time slots, with each time slot recording the target scene category, expected duration, and emotional marker. The data processing includes text-to-structure mapping and field filling, outputting a formalized video composition scheme data structure for subsequent segment selection.

[0424] Step 10: The server selects specific scene segments based on the composition scheme and segment features. The input is the scene segment list from step 4 (containing content feature quantities and quality indicators) and the video composition scheme data structure from step 9. For each time slot, the server selects one or more optimal segments from candidate segments that match the target scene category and sentiment markers, sorting them according to quality indicators and matching degree with the composition scheme. The server performs a trimming or splicing strategy on the selected segments based on the target duration to ensure the total duration meets the constraints. Data processing involves performing similarity sorting and selection in a multi-dimensional feature space, and arithmetic adjustments to the duration. The output is an edit timeline information containing actual segment identifiers, time truncation ranges, and sequence order.

[0425] Step 11: The server determines video performance parameters based on emotional state and composition scheme. The inputs are the emotional state vector generated in step 7 and the editing timeline information from step 10. The server records and assigns visual parameters such as hue, brightness, contrast, saturation, and filter type to each scene segment, and selects the corresponding music type and volume envelope shape. The server uses rule-based or learned mapping functions to map emotional intensity to parameter ranges (e.g., happy emotions correspond to higher saturation and faster tempo music). Data computation includes parameter mapping and normalization, and the output is a set of performance parameters that correspond one-to-one with the editing timeline information.

[0426] Step 12: The server performs video image processing and synthesizes a visual sequence. The inputs are the editing timeline information from step 10 and the visual performance parameters from step 11. The server calls image processing and video editing libraries, reads the original video clips within the corresponding time range, and performs cropping, scaling, stabilization, speed adjustment, and filter overlay on each clip. The server inserts specified transition effects (fade-in, fade-out, crossfade, etc.) between consecutive clips and splices all processed clips into a single video stream according to the timeline. Data processing involves resampling the frame sequence and performing pixel-level operations on the timeline, with the output being an intermediate video data stream that has not yet embedded the final audio track.

[0427] Step 13: The server performs audio synthesis and matches it to the video timeline. The inputs are the audio features from step 5, user-uploaded audio materials, the music type and volume envelope parameters obtained in step 11, and the timeline information from step 10. The server selects suitable background music from the audio material library based on the music type, and trims or loops the music signal according to the target duration. The server aligns the background music with any existing ambient sound effects and narration tracks, calculates the total volume and channel distribution at each moment according to the timeline, and uses a mixing algorithm to superimpose multiple audio tracks. The data processing involves weighted superposition and amplitude normalization of the multiple audio signals, and the output is the target audio track strictly aligned with the video timeline.

[0428] Step 14: The server encodes and stores the final video data. The inputs are the intermediate video data stream generated in step 12 and the target audio track generated in step 13. The server calls a multimedia encoding library to multiplex and encode the video and audio frame sequences into the target container format, setting the resolution, bitrate, encoder parameters, etc. During encoding, the server performs compression and encapsulation operations, converting the raw data into a compressed bitstream. The server writes the encoded final video file to a storage device and registers metadata such as file path, playback duration, and size in the database, while generating a unique matching identifier. The output is the final video data and its matching identifier that can be distributed over the network.

[0429] Step 15: The terminal obtains the final video identifier and displays the playback entry to the user. The input consists of the matching identifier returned by the server and metadata information related to the video. The terminal requests the playback or download address corresponding to this identifier via network request and generates a video thumbnail and a play button in the user interface. The terminal can request the video stream to start playback from the server and can also provide a share button for subsequent distribution. Data processing involves mapping the identifier and metadata to UI elements, and the output is a user-friendly playback and sharing interface.

[0430] Step 16: Users watch and provide feedback on the video on the terminal. The input consists of the final video footage and audio played on the terminal. After perceiving the video content visually and aurally, users can input their evaluations or modification requests on the terminal interface, such as "the music is too fast" or "I want more family photos." The terminal packages this textual feedback and possible rating values ​​and sends them to the server. The server receives this feedback as part of new abstract information and constraints, which can trigger a regeneration process to recalculate the editing timeline or performance parameters, thereby outputting a new version. Data calculus transforms human subjective opinions into structured constraint adjustment information, with the output being a feedback dataset for iterative optimization.

[0431] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires sound representing user input regarding the result of the specific processing. The control unit 46A sends the sound data representing user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0432] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0433] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.

[0434] For example, the collection unit is implemented by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart device 14 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the output device 40 of the smart device 14 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0435] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart device 14.

[0436] Second Implementation Method Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.

[0437] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. A server can be cited as an example of the data processing device 12.

[0438] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0439] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and communication I / F 44 are also connected to the bus 52.

[0440] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0441] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).

[0442] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0443] Figure 4 This illustrates an example of the main functions of the data processing device 12 and the smart glasses 214. For example... Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0444] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0445] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).

[0446] In the smart glasses 214, the processor 46 performs reception and output processing. The memory 50 stores the reception and output program 60. The processor 46 reads the reception and output program 60 from the memory 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is implemented by the processor 46 operating as a control unit 46A according to the reception and output program 60 executed on the RAM 48. Furthermore, the smart glasses 214 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290.

[0447] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart glasses 214. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0448] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0449] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0450] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0451] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0452] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A outputs the result of the specific processing to the speaker 240. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0453] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0454] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.

[0455] For example, the collection unit is implemented by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart glasses 214 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0456] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart glasses 214.

[0457] Third Implementation Method Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.

[0458] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. A server can be cited as an example of the data processing device 12.

[0459] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0460] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, display 343, and communication I / F 44 are also connected to the bus 52.

[0461] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0462] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).

[0463] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0464] Figure 6 This illustrates an example of the main functions of the data processing device 12 and the head-mounted terminal 314. For example... Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0465] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0466] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.

[0467] In the head-mounted terminal 314, the processor 46 performs the acceptance / output processing. The memory 50 stores the acceptance / output program 60. The processor 46 reads the acceptance / output program 60 from the memory 50 and executes the read acceptance / output program 60 on the RAM 48. The acceptance / output processing is implemented by the processor 46 operating as a control unit 46A according to the acceptance / output program 60 executed on the RAM 48.

[0468] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the head-mounted terminal 314. In the following description, the data processing device 12 will be referred to as the "server" and the head-mounted terminal 314 will be referred to as the "terminal".

[0469] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0470] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0471] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0472] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0473] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A outputs the result of the specific processing to the speaker 240 and the display 343. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0474] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 includes prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0475] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.

[0476] For example, the collection unit is implemented by the control unit 46A of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the head-mounted terminal 314 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 and display 343 of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0477] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the head-mounted terminal 314.

[0478] Fourth Implementation Method Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.

[0479] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. A server can be cited as an example of the data processing device 12.

[0480] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0481] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, controlled object 443, and communication I / F 44 are also connected to the bus 52.

[0482] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0483] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to photograph the area around robot 414 (e.g., the field of view defined by a perspective equivalent to the field of vision of an average healthy person).

[0484] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0485] The controlled object 443 includes a display device, LEDs (light-emitting diodes) for the eyes, and motors for driving the arms, hands, and feet. The posture or movement of the robot 414 is controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0486] Figure 8 This illustrates an example of the main functions of the data processing device 12 and the robot 414. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0487] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0488] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.

[0489] In robot 414, the processor 46 performs the acceptance and output processing. The memory 50 stores the acceptance and output program 60. The processor 46 reads the acceptance and output program 60 from the memory 50 and executes the read acceptance and output program 60 on RAM 48. The acceptance and output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance and output program 60 executed on RAM 48.

[0490] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the robot 414. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 will be referred to as the "terminal".

[0491] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0492] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0493] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0494] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0495] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A outputs the result of the specific processing to the speaker 240 and the controlled object 443. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0496] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0497] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.

[0498] For example, the collection unit is implemented by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the robot 414 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the robot 414 and the control object 443 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and can be varied.

[0499] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the robot 414.

[0500] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine a user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine a user's emotion based on an emotion graph that serves as a specific mapping (see [reference]). Figure 9 The emotion-specific model 59 can also determine the robot's emotion, and the specific processing unit 290 performs specific processing based on the robot's emotions.

[0501] Figure 9 This is a diagram representing an emotion map 400 that maps multiple emotions. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotion is. On the outer side of the concentric circles, emotions representing states or behaviors arising from mood are arranged. Emotions are concepts that include feelings and mental states. Emotions generated by reactions occurring in the brain are arranged roughly to the left of the concentric circles. Emotions derived from situational judgments are arranged roughly to the right of the concentric circles. Emotions generated by reactions occurring in the brain and derived from situational judgments are arranged roughly above and below the concentric circles. Furthermore, "pleasant" emotions are arranged above the concentric circles, and "unpleasant" emotions are arranged below them. Thus, in the emotion map 400, multiple emotions are mapped based on the structure that generates emotions, and emotions that are likely to occur simultaneously are mapped close to each other.

[0502] These emotions are distributed at the three o'clock position of the emotion map 400, typically fluctuating between peace and anxiety. In the right half of the emotion map 400, situational awareness dominates over internal sensation, thus resulting in an impression of calm.

[0503] The inner side of the emotion map 400 represents the inner state, while the outer side represents behavior. Therefore, the further outward you are from the emotion map 400, the more visible the emotion becomes (manifested in behavior).

[0504] Here, human emotions are based on various balances such as posture and blood sugar levels. When these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotions in robots, cars, motorcycles, etc., can also be created in the following way: based on various balances such as posture and remaining battery power, when these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a Brain Physiological Signal Analysis System for Voice Emotion Recognition and Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the sensory-dominated region, called "response," are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the situational cognition-dominated region, called "situation," are arranged.

[0505] In the emotion map, two types of emotions that promote learning are defined. One is a negative emotion on the situational side, in the middle or peripheral region of "repentance" or "reflection." This occurs when the robot experiences negative emotions such as "I don't want to experience this feeling again" or "I don't want to be blamed again." The other is a positive emotion on the response side, near the "desire" region. This occurs when there are positive feelings such as "wanting more" or "wanting to know more."

[0506] The emotion-specific model 59 inputs user input into a pre-trained neural network to obtain emotion values ​​representing each emotion shown in the emotion map 400, thereby determining the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network... Figure 10 As shown in the sentiment graph 900, it was trained in a way that sentiments that are configured close to each other have similar values. Figure 10 The text shows examples of emotions such as "peace of mind", "stability", and "reassurance" that have similar emotion values.

[0507] The above description focuses on the functions of the data processing device 12, but the system of this disclosure is not necessarily installed on a server. The system of this disclosure can also be installed as a general information processing system. This disclosure can also be installed, for example, as a software program running on a personal computer, an application running on a smartphone, etc. The method of this disclosure can also be provided to users in the form of SaaS (Software as a Service).

[0508] In the above embodiments, an example of a specific process being performed by a single computer 22 is given. However, the technology disclosed herein is not limited to this, and the specific process can also be distributed among multiple computers, including computer 22. For example, the data generation model 58 can be located on an external device of the data processing apparatus 12, where data is generated based on the input data.

[0509] In the above embodiments, examples of storing a specific processing program 56 in the memory 32 have been described, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed into the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0510] Alternatively, a specific processing program 56 may be pre-stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 according to the requirements of the data processing device 12.

[0511] In addition, it is not necessary to store all the specific processing program 56 in the storage device such as the server connected to the data processing device 12 via the network 54 or in the memory 32; a portion of the specific processing program 56 may be stored in advance.

[0512] As hardware resources for performing specific processes, various processors, as shown below, can be used. For example, a CPU can be listed as a processor, which functions as a general-purpose processor that performs specific processes by executing software, i.e., a program. Furthermore, processors can be listed as special-purpose circuits such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application-Specific Integrated Circuits), which are processors with circuitry specifically designed to perform specific processes. Each processor has built-in or connected memory, and each processor executes specific processes using that memory.

[0513] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resources for performing a specific process can be a single processor.

[0514] As an example of a single processor, there are two approaches: First, a processor is composed of a combination of one or more CPUs and software, which functions as a hardware resource to perform a specific process; second, as represented by a SoC (System-on-a-chip), a processor is used to implement the functionality of the entire system, which includes multiple hardware resources for performing a specific process, using a single IC (Integrated Circuit) chip. In this way, the specific process is implemented by using one or more of the aforementioned processors as hardware resources.

[0515] Furthermore, the hardware architecture of these various processors, more specifically, can utilize circuits that combine semiconductor elements and other circuit components. Moreover, the specific process described above is just one example. Therefore, without departing from the main point, unnecessary steps can certainly be deleted, new steps added, or the processing order changed.

[0516] The descriptions and illustrations above are detailed explanations of a portion of the technology disclosed herein, and are merely one example of the technology disclosed herein. For example, the above descriptions of the structure, function, effect, and results are just one example of the structure, function, effect, and results of a portion of the technology disclosed herein. Therefore, without departing from the spirit of the technology disclosed herein, unnecessary parts may be deleted, new elements added, or replacements may be made to the descriptions and illustrations above. Furthermore, to avoid confusion and facilitate understanding of a portion of the technology disclosed herein, explanations of common technical knowledge that do not require special explanation under the premise of being able to implement the technology disclosed herein have been omitted from the descriptions and illustrations above.

[0517] All documents, patent applications and technical specifications set forth in this specification are incorporated herein by reference to the same extent that each document, patent application and technical specification is specifically and individually described therein and referenced by reference.

[0518] In addition, the following notes are provided in response to the above explanation.

[0519] Example 1 (Note 1) An information processing system, characterized in that it comprises: A means of inputting information to obtain unedited video, image, and audio materials; Information conversion means for converting the unedited video material, image material and audio material into a standardized encoding method, screen resolution and time resolution, and for extracting visual feature quantities and structured accompanying information from the video material and the image material, and for extracting acoustic feature quantities and text information from the audio material; Information receiving means used to acquire abstract conditional information including the purpose, atmosphere, usage scenario, and duration of the target video; A means for generating prompt statements that provide a layered description of the composition, time allocation, scene selection, and effect additions of a target video based on the abstract conditional information, the visual feature quantities, the acoustic feature quantities, and the text information. A means of obtaining editing plans, which is used to issue video generation or video editing instructions to a generative artificial intelligence model based on the prompt statement and the feature quantity, and to obtain editing plan information including material identifier, start time, end time, transition type between scenes, visual effects, subtitle information and acoustic effects from the generative artificial intelligence model. A video generation method for extracting and connecting predetermined intervals from the unedited video, image, and audio materials based on the editing plan information, and performing scene transition effects, tone correction, text information overlay, and acoustic signal synthesis to generate complete video data; The method is used to provide the completed video data and obtain evaluation information or modification requests from users, reflect the evaluation information or modification requests in the abstract condition information to regenerate the prompt statement, and re-obtain the editing plan information and the iterative generation method of regenerating the completed video data through the generative artificial intelligence model. This is a sentiment analysis method used to modify the abstract condition information or the editing plan information based on information representing the user's emotional state, thereby adjusting the sentiment analysis of the completed video data content.

[0520] (Note 2) The information processing system according to Appendix 1 is characterized in that, When the abstract condition information includes conditions for providing short-duration videos, the prompt statement generation means, based on the visual feature quantity and the acoustic feature quantity, adds the indication content of the material selection criteria for meeting the expected playback duration and the quantity, length and order of the material intervals to be used to the prompt statement, and the editing plan acquisition means obtains the editing plan information corresponding to the indication content from the generative artificial intelligence model.

[0521] (Note 3) The information processing system according to Appendix 1 is characterized in that, The prompt statement generation method uses artificial intelligence technology to interpret the unedited video material, image material, and audio material, as well as the abstract conditional information, to generate prompt statements that reflect the interpretation results. The editing plan acquisition method uses the prompt statements and the feature quantities as conditions to drive the generative artificial intelligence model to run, thereby acquiring the editing plan information.

[0522] Application Example 1 (Note 1) An information processing system, characterized in that it comprises: An input means for receiving video information, still image information and audio information from a terminal device before editing in an information processing device, and for determining the format and size of the information to convert it into a predetermined format; An input method for obtaining string information from the user that represents the theme and purpose of the video content to be generated, and for performing structured processing on the string information to generate prompt statements that can be input into a generative artificial intelligence model. This is a material analysis method used to perform video parsing, image parsing, and audio parsing on the video information, the still image information, and the audio information, so as to extract the feature quantities corresponding to each information, and extract multiple candidate time intervals and candidate scenes based on the feature quantities. The method is used to perform natural language processing on the prompt statement to generate a numerical vector representing semantic information, and to use the numerical vector and the feature quantity as input to execute a generative artificial intelligence model to generate a generation instruction means for generating prompt statements containing control information for determining video composition that conforms to the topic and the purpose. Editing control means for determining timeline information, including the scene to be used, the order of the scenes, the duration of each scene, and the scene switching method between scenes, based on the generated prompt statement, the candidate time interval, and the candidate scene, and for editing the video information, the still image information, and the audio information according to the timeline information; A video generation means for performing video encoding and audio encoding processing based on the timeline information and screen switching method determined by the editing control means, so as to generate output video data that meets the predetermined compression method and image quality conditions; An emotion analysis method used to obtain operation history information or biometric information from a terminal device to infer the user's emotional state and adjust the timeline information or the screen switching method according to the emotional state. This is an output means for converting the output video data into a form that can be distributed to a terminal device and sent to the terminal device via a network for playback or storage on the terminal device.

[0523] (Note 2) According to the information processing system described in Appendix 1, the editing control means, when the prompt statement or the generated prompt statement contains conditions related to the public release of short-duration videos, analyzes the feature quantity and the numerical vector through the generative artificial intelligence model to calculate the target playback duration, and selects the time interval and number of shots that satisfy the target playback duration from the candidate time interval, thereby automatically setting the timeline information to a structure suitable for the public release of short-duration videos.

[0524] (Note 3) According to the information processing system described in Appendix 1, the generation instruction means utilizes a generative artificial intelligence model obtained by learning the correspondence between the video information, the still image information, the audio information, and the string information through artificial intelligence technology to generate a generation prompt statement containing multiple control parameters for indicating the video length, atmosphere, screen composition, and sound composition corresponding to the theme and the purpose, and the editing control means automatically determines the timeline information and the screen switching method based on the control parameters.

[0525] Example 2 (Note 1) An information processing system, characterized in that it comprises: An apparatus for receiving unedited video data, image data, and audio data via a communication network through a processing device, and for parsing and storing the data based on its type, recording duration, and image attributes; An apparatus for receiving natural language prompts, including target playback duration, aspect ratio, type of publishing media, and performance guidelines, from an information processing terminal via a processing device, and for converting and storing the prompts as structured video design information; An apparatus for performing natural language processing on the prompt statements and the video design information using a generative artificial intelligence model through a processing device, so as to generate video composition data containing time intervals, material types, text information content, image composition information and audio processing information; An apparatus for automatically adjusting the total duration of each time interval based on the video composition data to match the target playback duration, and for verifying and correcting the reference range of the referenced material based on the recording duration and attributes of the video data, image data, and audio data. A device for generating editing process definition data corresponding to the video composition data through a processing device, and controlling a video processing program or video editing device based on the editing process definition data to automatically perform material time extraction, image compositing, text information overlay and audio level adjustment to generate complete video data; An apparatus for storing the completed video data and its summary information in a storage area via a processing device, and distributing the completed video data to an information processing terminal upon receiving a request from the information processing terminal for viewing or access. An apparatus for recording the prompt statements, the video composition data, and historical information of the generation process through a processing device, and for updating the input format of the generative artificial intelligence model or the video composition data generation rules based on the historical information.

[0526] (Note 2) According to the information processing system described in Appendix 1, the processing device is configured to, when controlled according to a request to generate a short-duration video, extract the target playback duration and candidate material time intervals from the analysis results using a generative artificial intelligence model, automatically allocate the start and end times of each material within the target playback duration, and generate video composition data based on the allocation results to adjust the screen switching interval and information content of each time interval, thereby generating complete video data.

[0527] (Note 3) According to the information processing system described in Appendix 1, the processing device is configured to use artificial intelligence technology to calculate the correlation between the unedited video data, image data, and audio data and the prompt statement, determine the priority of each material based on the purpose, object, and performance conditions contained in the prompt statement, construct information input to the generative artificial intelligence model based on the priority to generate the video composition data, and generate completed video data based on the generated video composition data.

[0528] Application Example 2 (Note 1) An information processing system, characterized in that it comprises: A device for receiving unedited information data and storing the information data; A device for receiving and storing abstract information and constraints related to the generation of target video content; An apparatus for performing image and audio processing on the information data, extracting scene segments from the information data, and calculating content feature quantities and quality indicators for each scene segment; An apparatus for generating a prompt statement in natural language form from the composition information corresponding to the target video content based on the abstract information and the constraints, and inputting the prompt statement into a generative artificial intelligence model to obtain a video composition scheme; An apparatus for determining the scene segments to be used and the duration of each scene segment based on the video composition scheme, the content feature quantity, and the quality index, and for generating editing timeline information; A device for acquiring emotion-related data relating to a user and processing the emotion-related data to estimate an emotional state; An apparatus for determining video performance parameters, including hue, brightness, visual effects, music type, and volume balance, based on the emotional state and the video composition scheme, and for associating the video performance parameters with the editing timeline information to set visual and auditory effects. An apparatus for performing dynamic image processing and audio synthesis processing based on the editing timeline information and the video performance parameters, and integrating multiple information data to generate final video data; A means for storing the final video data and assigning a matching identifier to provide output information.

[0529] (Note 2) The information processing system according to Appendix 1 is characterized in that, The apparatus for generating the composition information is configured to, for the purpose of publishing a short-duration video, take the target playback duration specified as a constraint, the content feature quantity, and the quality index as a basis, include the selection criteria for scene segments and the time allocation strategy in the prompt statement, and obtain a video composition scheme adapted to the target playback duration from a generative artificial intelligence model.

[0530] (Note 3) The information processing system according to Appendix 1 is characterized in that, The device for determining the video performance parameters is configured to assign different tone transformation processing, filtering processing, scene switching processing, and background sound modes to each scene segment contained in the editing timeline information according to the emotional state, so as to dynamically adjust the overall emotional expression of the final video data.

Claims

1. An information processing system, characterized in that, include: processor; An input method, configured to input unedited video footage, image footage, and audio footage to the processor, so that the processor receives the footage; An information input means, the information input means being configured to input information about a video to be generated to the processor, so that the processor receives the information; A prompt generation unit is configured to enable the processor to parse the input material and the information, generate a prompt based on the parsing result to instruct the generative artificial intelligence model to automatically generate a video, and enable the processor to input the prompt to the generative artificial intelligence model to generate the video based on the prompt. A sentiment analysis method is configured to input a user's emotional state into the processor, so that the processor can analyze the user's emotions and adjust the content of the video based on the sentiment analysis results.

2. The information processing system according to claim 1, characterized in that, The processor is configured to, when the purpose is to publish a short video, parse the necessary playback duration and the shot segments that should be used in relation to the short video, generate a prompt based on the parsing results to instruct a generative artificial intelligence model to generate the short video, and input the prompt to the generative artificial intelligence model to generate the short video based on the prompt.

3. The information processing system according to claim 1, characterized in that, The processor is configured to use artificial intelligence technology to parse the input material and the information, generate a prompt based on the parsing result to instruct the generative artificial intelligence model to automatically generate a video, and input the prompt to the generative artificial intelligence model to generate the video based on the prompt.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A