Video generation method and device based on data portrait, and storage medium

By constructing data profiles and calling media generation engines in parallel to generate video content, the problem of insufficient scene coordination in existing technologies is solved, achieving stylistic coordination and narrative coherence in video generation, and improving generation efficiency.

CN122476247APending Publication Date: 2026-07-28SHENZHEN SHUZHIHUA INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN SHUZHIHUA INFORMATION TECH CO LTD
Filing Date
2026-06-18
Publication Date
2026-07-28

AI Technical Summary

Technical Problem

In existing technologies, video generation systems lack coordination between different scenes when generating video content, making it difficult to generate style-matched voiceovers, background music, animations, and visual effects based on emotional changes and content descriptions.

Method used

By constructing a data profile of the target object, a storyboard containing scene descriptions, narration text, and emotional tags is generated. In parallel, a media generation engine is called to generate voice-over, background music, animation materials, and stylized materials, which are then synthesized into a personalized video based on the timeline.

Benefits of technology

It achieves stylistic consistency and narrative coherence across different scenes, improves generation efficiency, and ensures that the specific emotional needs of each scene are met synchronously.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122476247A_ABST
    Figure CN122476247A_ABST
Patent Text Reader

Abstract

The application discloses a video generation method and device based on a data portrait and a storage medium, and belongs to the technical field of data processing. The method comprises the following steps: aggregating multi-source data of a target object from a data source, constructing a data portrait of the target object, inputting the data portrait into a story generation model, and generating a storyboard, wherein the storyboard comprises at least two scenes arranged based on a time axis, a scene is associated with a scene description, a side speech script and / or an emotional label, at least two media generation engines are called in parallel according to the emotional label and the scene description of the scene in the storyboard, voice dubbing, background music, animation materials and / or stylized materials are generated, and the generated voice dubbing, background music, animation materials and stylized materials are combined into a personalized video of the target object according to the time axis of the storyboard. The application combines the generated various materials into the personalized video of the target object based on the storyboard, and improves the overall coordination between various video generation scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a video generation method, device and storage medium based on data profiling. Background Technology

[0002] In school or home education settings, it is necessary to generate personalized growth stories, highlight moments, and other video content for each student or child.

[0003] In related technologies, users typically select a preset video template and input text or upload materials. The system directly uses the user's input text as narration, calls a speech synthesis engine to generate voiceover, and calls a music engine to generate background music according to the preset music style of the template. Then, the voiceover, background music, and user-uploaded materials are spliced ​​together in a fixed order to output a complete video.

[0004] In this process, the system uses a serial calling method, with each engine running independently. The processing of voice-over, background music, and materials is executed serially, meaning that the next engine is started only after the previous one is completed. The engines operate independently of each other and lack information exchange. Under these circumstances, the system struggles to generate style-matched voice-over, background music, animation, and visual effects for each scene in a targeted and parallel manner based on the emotional changes and content descriptions of different scenes in the video, resulting in insufficient coordination between scenes in the video.

[0005] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0006] The main purpose of this application is to provide a video generation method, device and storage medium based on data profiling, which aims to solve the technical problem of insufficient coordination between various scenes in the generated video content.

[0007] To achieve the above objectives, this application provides a video generation method based on data profiling, the method comprising the following steps: Aggregate multi-source data of the target object from data sources to construct a data profile of the target object; The data profile is input into the story generation model to generate a storyboard, which contains at least two scenes arranged according to a timeline, and the scenes are associated with scene descriptions, narration text and / or emotional tags; Based on the emotional tags and scene descriptions of the scene in the storyboard, at least two media generation engines are invoked in parallel to generate voice-over, background music, animation materials and / or stylized materials; Based on the timeline of the storyboard, the generated voice-over, background music, animation materials, and stylized materials are combined into a personalized video of the target object.

[0008] In one embodiment, the step of aggregating multi-source data of the target object from the data source to construct a data profile of the target object includes: The multi-source data is formed by collecting the sports rating records, growth record texts, campus activity materials and / or honor records of the target object from the data source. The multi-source data is stored in a structured manner based on a preset data model, which includes data type, timestamp, and content summary fields. The multi-source data, after being stored in a structured manner, is integrated into a data profile of the target object.

[0009] In one embodiment, the step of inputting the data profile into the story generation model to generate a storyboard includes: The data profile and the preset narrative template are input into the large language model. The narrative template defines the range of the number of scenes, the upper limit of the total duration, and the requirements for the emotional direction of the storyboard. Receive the target storyboard output by the large language model, wherein the target scene in the target storyboard includes scene identifier, scene description, narration text, sentiment tag and recommendation duration field; Perform timeline verification on the target storyboard, and select the target storyboard as the storyboard based on the verification results.

[0010] In one embodiment, the step of generating voice-over, background music, animation footage, and / or stylized footage by invoking at least two media generation engines in parallel based on the emotional tags and scene descriptions of the scene in the storyboard includes: The narration text for each scene in the storyboard is transmitted sequentially to the speech synthesis engine to generate dubbing audio clips corresponding to each scene; The emotional tag sequences of each scene in the storyboard are transmitted to the music generation engine to generate a background music track that matches the emotional tag sequences; The scene descriptions of each scene in the storyboard are transmitted to the style transfer engine, and a visual style matching the scene description is applied to the original photo or video material of the target object to generate stylized material.

[0011] In one embodiment, the step of synthesizing the generated voice-over, background music, animation footage, and stylized footage into a personalized video of the target object according to the timeline of the storyboard includes: Based on the arrangement order and recommended duration of each scene in the storyboard, the dubbing audio segments are segmented, and the segmented dubbing audio segments are aligned with the timeline; The background music track is overlaid onto the aligned dubbing audio track, and the background music volume is adjusted to avoid noise. The animation and stylized materials are spliced ​​together in chronological order; After inserting transition effects between adjacent materials, the audio track, the dubbing track, and the background music track are merged and output as the personalized video.

[0012] In one embodiment, after the step of inputting the data profile into the story generation model to generate a storyboard, the method further includes: The content fingerprint of the storyboard is calculated by performing a hash operation on the storyboard; Send an authorization request to the user terminal, the authorization request containing the content fingerprint, and receive an authorization instruction containing the content fingerprint returned by the guardian; Before invoking the media generation engine, verify whether the content fingerprint of the current storyboard matches the content fingerprint in the authorization instruction; If the verification results are consistent, the step of generating voice-over, background music, animation materials and / or stylized materials by calling at least two media generation engines in parallel based on the emotional tags and scene descriptions of the scene in the storyboard is executed.

[0013] In one embodiment, before generating voice-over, background music, animation footage, and / or stylized footage by simultaneously invoking at least two media generation engines based on the emotional tags and scene descriptions of the scene in the storyboard, the method further includes: The storyboard preview interface is displayed to the user terminal, and the target emotional intensity value is received from the user terminal for at least one scene. Based on the received target emotion intensity value, a target emotion curve is constructed, where the horizontal axis of the target emotion curve represents time and the vertical axis represents emotion intensity. The target emotion curve is used as a constraint and input into the story generation model so that the story generation model can adjust the duration of each scene in the storyboard or merge adjacent scenes.

[0014] In one embodiment, the step of concurrently calling at least two media generation engines to generate voice-over, background music, animation footage, and stylized footage further includes: Extract the emotional tags and scene descriptions of each scene in the storyboard, and calculate the global style vector of the storyboard; The global style vector is input into the music generation engine and the style transfer engine, so that the background music generated by the music generation engine maintains thematic consistency between adjacent scenes, and the style transfer engine applies a unified visual art style to adjacent scenes. Based on the difference in the emotional tag vectors between the two scenes, generate transitional audio clips or transitional visual effects.

[0015] In addition, to achieve the above objectives, this application also provides a video generation device based on data profiling, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the video generation method based on data profiling as described above.

[0016] In addition, to achieve the above objectives, this application also provides a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the video generation method based on data profiling as described above.

[0017] One or more technical solutions proposed in this application have at least the following technical effects: This application constructs a data profile of the target object by aggregating multi-source data from a data source. The data profile is then input into a story generation model to generate a storyboard containing at least two scenes arranged along a timeline, each scene being associated with a scene description, narration text, and / or emotional tags. Based on the emotional tags and scene descriptions in the storyboard, at least two media generation engines are invoked in parallel to generate voice-over, background music, animation materials, and / or stylized materials, respectively. According to the timeline of the storyboard, the generated materials are combined into a personalized video of the target object. This allows the voice-over, background music, animation materials, and stylized materials to be generated synchronously for the specific emotional needs of each scene, thereby reducing the impact of sequential generation on generation efficiency. At the same time, emotional and content constraints ensure the stylistic coordination and narrative coherence of media assets between different scenes. Attached Figure Description

[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a flowchart illustrating the first embodiment of the video generation method based on data profiling in this application; Figure 2 This is a flowchart illustrating the second embodiment of the video generation method based on data profiling in this application; Figure 3 This is a flowchart illustrating the third embodiment of the video generation method based on data profiling in this application; Figure 4 This is a schematic diagram of the structure of a data-based video generation device in the hardware operating environment involved in the embodiments of this application.

[0021] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0022] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.

[0023] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.

[0024] The main solution of this application embodiment is as follows: aggregate multi-source data of the target object from the data source to construct a data profile of the target object, input the data profile into the story generation model to generate a storyboard, wherein the storyboard contains at least two scenes arranged according to the timeline, and the scenes are associated with scene descriptions, narration text and / or emotional tags. According to the emotional tags and scene descriptions of the scenes in the storyboard, at least two media generation engines are called in parallel to generate voice-over, background music, animation materials and / or stylized materials. According to the timeline of the storyboard, the generated voice-over, background music, animation materials and stylized materials are combined into a personalized video of the target object.

[0025] In existing technologies, users typically select a preset video template and input text or upload materials. The system directly uses the user-input text as narration, calls a speech synthesis engine to generate voice-over, and simultaneously calls a music engine to generate background music based on the preset music style of the template. The voice-over, music, and user-uploaded materials are then concatenated in a fixed order to output a complete video. In this process, the system uses a serial calling method, with each engine running independently. The processing of voice-over, music, and materials is executed sequentially; that is, one engine is called before the next is started. The engines operate independently and lack information interaction. Under these circumstances, the system struggles to generate style-matched voice-over, music, animation, and visual effects for each scene in a targeted and parallel manner, based on the emotional changes and content descriptions of different scenes in the video. This results in insufficient coordination between scenes in the video.

[0026] This application constructs a data profile of the target object by aggregating multi-source data from a data source. The data profile is then input into a story generation model to generate a storyboard containing at least two scenes arranged along a timeline, each scene being associated with a scene description, narration text, and / or emotional tags. Based on the emotional tags and scene descriptions in the storyboard, at least two media generation engines are invoked in parallel to generate voice-over, background music, animation materials, and / or stylized materials, respectively. Finally, based on the timeline of the storyboard, the generated materials are combined into a personalized video of the target object. This allows voice-over, background music, animation materials, and stylized materials to be generated synchronously for the specific emotional needs of each scene, thereby eliminating the efficiency bottleneck caused by serial generation. At the same time, scene-level emotional and content constraints ensure the stylistic coordination and narrative coherence of media assets between different scenes.

[0027] To better understand the above technical solutions, exemplary embodiments of this application will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of this application are shown in the drawings, it should be understood that this application can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of this application and to fully convey the scope of this application to those skilled in the art.

[0028] It should be noted that the executing entity in this embodiment can be a video generation system, or a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device capable of performing the above functions, a video generation device based on data profiling, etc. This embodiment does not specifically limit it. The following uses a video generation system as an example to describe this embodiment and the following embodiments.

[0029] Based on this, embodiments of this application provide a video generation method based on data profiling, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the video generation method based on data profiling in this application.

[0030] In this embodiment, the video generation method based on data profiling includes steps S10 to S40: Step S10: Aggregate multi-source data of the target object from the data source to construct a data profile of the target object; In this embodiment, the video generation server deployed in the campus information management system is connected to various business systems within the campus, such as the sports scoring database, student growth record platform, campus activity media library, and / or honor and award registration system, via an internal network. The video generation system uses the unique identifier of the target object as the query key, sending data retrieval requests to each of the aforementioned data sources. After receiving the raw data returned by each data source, it performs field mapping, type conversion, and structured storage according to a preset data profile model, integrating the heterogeneous data scattered across various business systems into a complete data profile record indexed by the target object's identifier. This data profile covers multiple dimensions of information about the target object, including sports performance, teacher evaluation records, campus activity participation videos, and honor and achievement tags, providing a structured input foundation for subsequent processing.

[0031] It should be noted that data profiling refers to a personalized digital description formed by the structured integration of multi-dimensional data generated by a target individual in a campus setting. This description reflects the target individual's growth trajectory, behavioral characteristics, and achievements. Sports scoring records originate from structured data stored in the database after the sports analysis system performs posture recognition and comprehensive scoring on student sports videos. Growth record texts come from written comments entered by teachers through their terminals during daily teaching and from semester summaries written by students themselves. Campus activity materials are video footage collected by acquisition devices during campus activities and uploaded to the media library after review. Honor records are from award information entered by the school's honor registration system after each commendation activity. These various types of data use different data models and field naming conventions during their original storage. They need to be unified through field mapping and format conversion at the data access layer of the video generation system before being aggregated into a single profiling model.

[0032] Specifically, a structured query is sent to the sports scoring database through the data access layer. The query condition is set so that the student ID number of the target object equals the identity identifier of the target object being processed, and the query scope is all sports scoring records of the target object within a complete semester. The records returned by the sports scoring database include the sports name, comprehensive score value, detailed scores for each dimension, and score generation timestamp. After receiving these records, the score value and score timestamp are extracted and mapped according to the ability dimension fields in the preset data profile model. The score value is stored under the corresponding sports subfield in the ability dimension field, and the score timestamp is stored in the timestamp field of the ability dimension record. At the same time, a query request is sent to the student growth record platform, which stores teachers' daily comments and self-summary texts for students. The text records returned by the growth record platform include a record type identifier and a record generation time. The video generation system maps the text content to the behavior description field and comment summary field in the data profile model according to the record type identifier, and uses the record generation time as the timestamp of the field.

[0033] Furthermore, the video generation system sends a query request to the campus activity media library, which stores video footage of students during campus activities, along with the recording time, activity name, and a list of participants. Upon receiving the returned list of media file storage paths and corresponding metadata, the video generation system extracts the storage path, recording timestamp, and activity tags of the media files, combining them into media material entries and mapping them to the media material field of the data profile model. The video generation system then sends a query request to the honors and awards registration system, which records various competition awards, honorary titles, and award dates received by students. The records returned by the honors and awards registration system include the honor name, award level, award date, and honor description text. The video generation system combines the honor name and award date into achievement tag entries and maps them to the achievement tag field of the data profile model. After completing all the above mapping operations, the video generation system aggregates all mapped field values ​​into a data profile record using the target object's identity as the key and persistently stores this record in the local profile database.

[0034] As an optional implementation, the video generation system achieves data aggregation through timed batch extraction. The system triggers a full extraction task at a preset time each day. This task first reads the list of target object identifiers stored in the local profile database. Then, for each identifier in the list, it sends incremental query requests to the sports scoring database, the student growth record platform, the campus activity media library, and the honor and award registration system. The condition for the incremental query is that the update timestamp of the record is between the last extraction time and the current time.

[0035] After each data source returns incremental data records that meet the criteria, the video generation system performs field mapping and type conversion on each incremental record, and then merges the converted data with the existing target object profile in the local profile database. Specifically, for rating records in the capability dimension fields, the video generation system deduplicates the data based on the rating timestamp, retaining the record with the most recent timestamp. For comment text in the behavior description field, the video generation system appends the comment text based on the record's creation time, without overwriting existing records. For media material fields, the video generation system deduplicates the data based on the inclusion timestamp, retaining the material entry with the most recent timestamp. For achievement tag fields, the video generation system deduplicates the data based on the award date, retaining the record with the most recent date. After the merging is complete, the video generation system rewrites the updated complete profile record into the local profile database, overwriting the old profile record for the target object.

[0036] As an alternative implementation, the video generation system employs an event-driven approach to achieve data aggregation. The system deploys data change listeners at the write interfaces of each data source. These listeners capture write events by subscribing to the data change log streams of each data source. When the sports scoring database receives a new scoring record and completes the write, a change event containing the new record's primary key and write timestamp is generated in the database's change log. The listener captures this event, parses its content, extracts the target object's identity and score value from the new scoring record, encapsulates the extracted data into a data change message, and pushes it to the video generation system's data access message queue. The video generation system, acting as a consumer of the message queue, consumes each data change message in real time, parses the target object's identity and changed data content from the message body, and integrates the changed data into the target object's existing profile according to a pre-defined data profiling model. After integration, the video generation system synchronously writes the updated profile record to its local profile database.

[0037] For example, after the video generation system sends a structured query to the sports scoring database, the database returns multiple records corresponding to scores for multiple sports events, each containing a corresponding score generation timestamp. The video generation system stores the score values ​​of these records into the respective sports event subfields under the ability dimension field, and stores each score timestamp into the timestamp attribute of the corresponding subfield. Simultaneously, after the video generation system sends a query request to the student growth record platform, the platform returns multiple text records, including teacher comment type records and self-summary type records, each containing a generation timestamp. The video generation system maps teacher comment type text to the behavior description field and self-summary type text to the comment summary field, using their respective generation times as the timestamps of the corresponding fields. After the video generation system sends a query request to the campus activity media library, the media library returns multiple video materials of the target object included in different campus activities, along with their storage paths, inclusion timestamps, and activity tags. The video generation system combines this information into multiple media material entries and maps them to the media material field. For example, after the video generation system sends a query request to the honor and recognition registration system, the system returns records containing the honor name and the date of award. The video generation system combines the honor name and the date of award into achievement tag entries and maps them to the achievement tag field. After item-by-item mapping, the video generation system uses the target object's identity identifier as an index to aggregate all sports event rating records in the ability dimension field, teacher comments and self-evaluation texts in the behavior description field, video data storage paths in the media material field, and honor title records in the achievement tag field into a complete data profile record, which is then written to the local profile database.

[0038] Step S20: Input the data profile into the story generation model to generate a storyboard; In this embodiment, the video generation system reads the constructed data profile from the profile database, converts it from a structured field format into a natural language text sequence that meets the input requirements of the large language model, and uses this text sequence, along with a preset narrative template, as input prompts. The system then calls the large language model service via an application programming interface (API). Upon receiving the prompts, the large language model generates a structured storyboard containing multiple scenes, each equipped with scene description text, narration, and emotional tags, according to the number of scenes, total duration, and emotional direction requirements defined in the narrative template. After receiving the output from the large language model, the video generation system performs format parsing and data cleaning, extracts the field values ​​of each scene, assembles them into a scene sequence, and performs timeline verification and adjustment on the scene sequence to ensure the total duration meets the preset duration constraint. The verified storyboard is then used as the official storyboard for the current task.

[0039] The story generation model is a computational model capable of converting structured input data into a narrative content sequence. Built on a large language model, it receives structured data containing field names and values ​​as contextual input and outputs structured text including scene divisions and descriptions according to predefined narrative constraints. A storyboard is a narrative structure composed of multiple scenes arranged chronologically. Each scene includes scene description text to describe the visual content, narration text for voice-over, and emotion tags to identify the emotional tone. Scene descriptions guide the generation of subsequent visual materials, narration text for subsequent audio content synthesis, and emotion tags for subsequent background music generation.

[0040] Specifically, the video generation system reads the target object's data profile record from the profile database. This record contains multiple field names and corresponding field values. The video generation system then initiates a text converter, which converts the key-value pair structured data into a sequence of natural language sentences. The converter iterates through each field in the data profile. For the sports item rating value in the ability dimension field, the converter generates a sentence structure corresponding to that sports item. For the teacher's comments text in the behavior description field, the converter directly quotes the original text and adds the corresponding prefix. For the activity tag list in the media material field, the converter generates a sentence structure for participating in the corresponding activity. For the honor name list in the achievement tag field, the converter generates a sentence structure for obtaining the corresponding honor.

[0041] The converter groups and concatenates all the above sentence structures according to field type into a coherent structured input text, separating each sentence with a period, and appends the target object's name and grade information to the end of the text. The video generation system then reads a preset narrative template, which is a structured cue word framework containing lower and upper limits for the number of scenes, lower and upper limits for the total duration, a list of preset patterns for emotional trajectory, and an example of a storyboard output format. The video generation system combines the structured input text with the narrative template to form a complete cue word, sends the cue word to the application programming interface of the large language model via Hypertext Transfer Protocol, and sets the response format to structured data format for subsequent parsing.

[0042] After receiving the prompt words, the large language model, based on the target object data described in the structured input text, autonomously determines the number of scenes, the recommended duration of each scene, and the order of scenes within the constraints defined by the narrative template, outputting an initial storyboard containing a list of scenes. The video generation system, upon receiving the response from the large language model, verifies whether the data format of the response conforms to the preset field structure requirements, i.e., whether it contains a scene list field and whether each scene contains four sub-fields: scene description, narration text, sentiment tag, and recommended duration. If the verification passes, the video generation system extracts the recommended duration field value of each scene and sums them to obtain the total duration, comparing the total duration with the upper and lower limits defined in the narrative template. If the total duration falls within the upper or lower limit, the initial storyboard is directly selected as the official storyboard. If the total duration exceeds the upper limit, the video generation system calculates the excess ratio and performs reverse proportional compression according to the numerical value of the recommended duration of each scene. If the total duration is below the lower limit, the video generation system prioritizes expanding the recommended duration of adjacent scenes with specific sentiment tags until the total duration reaches the lower limit requirement. After the timeline verification and adjustment are completed, the video generation system assigns scene numbers to each scene according to the order of arrangement. The scene description, narration text, emotional tags and adjusted recommendation duration of each scene are stored using the scene number as the key to form the formal storyboard.

[0043] Optionally, the video generation system also performs an authorization verification process on the generated storyboard. The video generation system performs a hash operation on the storyboard to calculate its hash value as a content fingerprint, which uniquely identifies the current content state of the storyboard. The video generation system sends an authorization request to the guardian's terminal, which includes the storyboard's preview information and the calculated content fingerprint. Upon receiving the authorization request, the guardian's terminal displays a storyboard preview interface. After confirming the storyboard content in the preview interface, the guardian performs a signature operation to generate an authorization instruction. This authorization instruction includes the storyboard content fingerprint at the time of the guardian's confirmation and the guardian's digital signature. After receiving the authorization instruction returned by the guardian's terminal, the video generation system extracts the content fingerprint carried in the authorization instruction and recalculates the hash value of the current storyboard to obtain the current content fingerprint. The content fingerprint in the authorization instruction is compared and verified with the current content fingerprint. If the verification result is consistent, the subsequent process continues; if the verification result is inconsistent, it indicates that the storyboard content has been modified after authorization, the subsequent process is paused, and a prompt message is output.

[0044] As an optional implementation, before invoking the large language model, the number of honorary records and the extreme value distribution in sports rating records of the target object are statistically analyzed from the data profile. The video generation system dynamically adjusts the lower and upper limits of the number of scenes defined in the narrative template based on the comparison between the number of honors and a preset honor abundance threshold. When the number of honors exceeds the first threshold, the upper limit of the number of scenes is increased, ensuring that the generated storyboard fully covers all the highlight moments of the target object. Simultaneously, the video generation system determines the fluctuation range of the emotional trajectory based on the changes in sports ratings. When the fluctuation range exceeds the second threshold, the preset emotional trajectory mode of the narrative template tends to include obvious high and low emotional contrasts; when the fluctuation range is below the third threshold, the preset emotional trajectory mode of the narrative template tends to be smooth and gradual.

[0045] As an alternative implementation, when the timeline verification fails, the video generation system does not adjust the duration of all scenes by a uniform ratio. Instead, it first identifies the emotional tag type of each scene. For key scenes with specific emotional tag types, the video generation system keeps their recommended duration unchanged; for other types of scenes, the video generation system compresses or expands the duration, thereby preserving the core emotional expression effect of the storyboard to the greatest extent while meeting the total duration constraint.

[0046] Specifically, the video generation system first calculates the sum of the recommended durations of all key scenes. If the sum exceeds the total duration limit, the video generation system compresses the duration of all scenes proportionally. If the sum does not exceed the total duration limit, the video generation system only compresses other types of scenes. The compression amount is the average compression amount distributed across all other types of scenes, which is the difference between the total duration limit and the sum of the key scene durations.

[0047] Step S30: Based on the emotional tags and scene descriptions in the storyboard, call at least two media generation engines in parallel to generate voice-over, background music, animation materials and / or stylized materials; In this embodiment, the video generation system uses the generated storyboard as input, extracts different types of data fields from it, and routes them to the corresponding media generation engines. The video generation system extracts the narration text for each scene and concatenates it into a complete narration text sequence, which is then sent to the speech synthesis engine to generate voice-over audio. It also extracts the emotional tags for each scene to form an emotional tag sequence, which is sent to the music generation engine to generate background music tracks. Furthermore, it combines the scene descriptions of each scene with the original video data in the constructed data profile and sends this data to the style transfer engine to generate stylized visual materials. Finally, based on the character action requirements in the scene descriptions, it calls the skeletal-driven animation engine to generate dynamic effects.

[0048] Optionally, the video generation system uses a concurrent thread pool to initiate multiple generation requests simultaneously. After all engines return their generation results, the media materials output by each engine are indexed and stored according to scene identifiers for use in subsequent compositing steps. A media generation engine refers to a computational service unit specifically designed to generate specific types of media content. These include speech synthesis engines for converting text to speech, music generation engines for generating instrumental music based on emotional tags, style transfer engines for applying specific visual styles to images or videos, and skeletal-driven animation engines for converting static images into dynamic visuals.

[0049] Specifically, the video generation system iterates through each scene in the storyboard, extracts the narration text field values ​​for each scene, and concatenates all the narration texts into a complete narration text sequence in ascending order of scene number. A preset-duration silence separator is inserted between the narration texts of adjacent scenes; this silence separator is used to create natural speech pauses between scenes during subsequent synthesis. The video generation system sends the concatenated complete narration text sequence to the speech synthesis engine via Hypertext Transfer Protocol (HTTP). This request also includes timbre identification parameters, instructing the speech synthesis engine to use a speaker model that matches the target audience's age and gender. After receiving the request, the speech synthesis engine sequentially inputs each sentence in the narration text sequence into a text-to-acoustic feature model. This model converts the text into an acoustic feature sequence, which is then input into a vocoder model to reconstruct waveform audio. The output waveform audio file and its corresponding word timestamp list are then generated. The word timestamp list records the start and end times of each word in the audio file.

[0050] Optionally, the video generation system simultaneously extracts the emotion tag field values ​​for each scene from the storyboard, arranges them according to scene number to form an emotion tag sequence, and sends this sequence to the music generation engine. Upon receiving the emotion tag sequence, the music generation engine maps each emotion tag to a corresponding combination of musical parameters, including mode, tempo, orchestration timbre, and dynamic range. The music generation engine generates music clips corresponding to each scene sequentially according to the order of the emotion tags, with the duration of each music clip matching the recommended duration for the corresponding scene. The music clips are then sequentially spliced ​​into a complete background music track, while simultaneously recording the start and end times of each music clip within the track. The video generation system extracts the storage path of the original image data of the target object from the media material field of the data profile, and extracts the composition requirements and visual style keywords for each scene from the scene description field of the storyboard. The video generation system combines the storage path of the original material, the visual style keywords, and the scene identifier of the corresponding scene into a style transfer request, which is then sent to the style transfer engine. Upon receiving the request, the style transfer engine reads the original image data from the storage path and fuses the content features of the original image with the style features corresponding to the visual style keywords.

[0051] Optionally, a text fragment containing a description of a person's actions is extracted from the scene description, and a person image of the target object is extracted from the data profile. The text describing the person's actions and the person image are then sent to the skeleton-driven animation engine. Upon receiving the request, the skeleton-driven animation engine performs human pose estimation on the person image, extracts the coordinates of two-dimensional skeletal keypoints of the person in the photo, determines the target action type and action amplitude parameters based on the text describing the person's actions, converts the static skeletal keypoint coordinates into a dynamic skeletal keypoint sequence through a motion generation network, and uses the dynamic skeletal keypoint sequence to drive image deformation of the original image to generate a dynamic video clip.

[0052] As an optional implementation, before invoking the media generation engines in parallel, the video generation system first sends a health check request to the status monitoring endpoint of each engine to obtain the current load queue length and average response latency of each engine. For engines whose load queue length exceeds a preset capacity threshold, the video generation system switches their corresponding generation tasks to the same type of engine deployed on a standby computing node, achieving task-level load balancing. The video generation system is configured with two sets of media generation engine instances, primary and backup. When the video generation system detects that the load queue length of the primary instance exceeds a preset percentage of the capacity threshold, the video generation system routes the current batch of generation tasks to the backup instance while continuing to monitor the load changes of the primary instance. When the load queue length of the primary instance falls back below the preset percentage of the capacity threshold, the video generation system switches subsequent tasks back to the primary instance for processing.

[0053] As an alternative implementation, when the video generation system calls the music generation engine and style transfer engine, it additionally calculates the global style vector of the storyboard and passes this global style vector as additional input to both engines. This ensures that the background music generated by the music generation engine maintains consistency in thematic melody across scene segments, while the style transfer engine applies a unified visual art style to adjacent scenes. The video generation system obtains the global style vector by weighted averaging the vectorized representations of the emotional tags for each scene in the storyboard, where the weight of each scene is proportional to its recommended duration. After receiving the global style vector, the music generation engine uses it as a constraint when generating music segments for each scene, ensuring coherence in modal tonic and rhythmic structure. After receiving the global style vector, the style transfer engine uses the mean hue and texture roughness from the global style vector as style benchmarks when generating stylized images for each scene, ensuring visual continuity in color saturation and brushstroke texture for stylized images of adjacent scenes.

[0054] Step S40: Based on the storyboard's timeline, combine the generated voiceover, background music, animation materials, and stylized materials into a personalized video for the target object.

[0055] In this embodiment, the video generation system takes voice-over audio, background music tracks, stylized materials, and animation materials returned by multiple media generation engines as input, and uses a timeline composed of the scene arrangement order and recommended duration in the storyboard as the synthesis benchmark. The video generation system segments the voice-over audio into independent segments corresponding to each scene, placing each segment at its corresponding starting position on the timeline according to the scene order. Then, the background music track is overlaid on the voice-over track, and volume avoidance processing is applied to the background music during the periods when the voice-over segments are present. Stylized materials and animation materials are arranged on the video track according to scene order, and transition effects are inserted between adjacent materials. The video track and audio track are multiplexed and encoded to output a complete personalized video file.

[0056] Specifically, the video generation system constructs a linear timeline with preset precision based on the arrangement order and recommended duration of each scene in the storyboard. The starting point of this timeline is the initial time, and the ending point is the sum of the recommended durations for each scene. The video generation system reads the complete waveform audio file returned by the speech synthesis engine and its corresponding word timestamp list. Each word timestamp in this list is based on the start time of the complete waveform audio file. Based on the narration text for each scene in the storyboard, the video generation system searches the word timestamp list for the timestamps corresponding to the first and last words of the narration text for each scene. Using the start timestamp of the first word as the start position of the dubbing segment for that scene and the end timestamp of the last word as the end position, the complete waveform audio is divided into multiple dubbing segments, each carrying a scene identifier for the corresponding scene.

[0057] Each voice-over segment is placed on the audio track according to the starting position of the timeline corresponding to the scene identifier. A crossfade transition is inserted between adjacent voice-over segments. Specifically, the volume of the preceding segment linearly decreases to zero for a preset duration at the end, while the volume of the following segment linearly increases to a preset volume value for a preset duration at the beginning, thus eliminating perceptible interruptions at the segment splicing points. The video generation system then overlays the complete background music track audio file returned by the music generation engine as a second audio track onto the audio track where the voice-over segments have been placed. The video generation system iterates through each time point on the timeline, determining whether that time point falls within the coverage area of ​​any voice-over segment. If so, the volume gain of the background music at that time point is reduced to a preset percentage of the voice-over volume; otherwise, the volume gain of the background music is restored to 100%, thus achieving volume avoidance processing that prioritizes voice-over.

[0058] Furthermore, after completing all processing of the audio track, the video generation system enters the video track processing flow. The video generation system reads the stylized image files corresponding to each scene returned by the style transfer engine and the animation clip files returned by the skeletal-driven animation engine. According to the scene identifier of each scene, it places each material file at the starting position on the video track corresponding to the recommended duration of that scene, and stretches or compresses the duration of each material to the recommended duration of that scene. Stretching or compression is achieved through interpolation sampling. Based on the transition type field in the scene description of each scene in the storyboard, corresponding transition effects are inserted at the junction of materials from two adjacent scenes. The duration of the transition effect is a preset percentage of the smaller of the recommended durations of the two adjacent scenes. After completing all processing of the video track, the video generation system synchronously multiplexes the video track containing transition effects with the audio track after volume avoidance processing using an audio-video multiplexing encoder. Video frames and audio frames are encapsulated into the same container format media unit according to the same timestamp on the timeline, and the encoded output is a personalized video file with a container format and resolution supported by the target playback terminal.

[0059] As an optional implementation, before compositing the video, the video generation system obtains the capability parameters of the target playback terminal via an application programming interface (API). If the resolution of the target playback terminal is lower than the native resolution of the current video footage, the video generation system activates a video scaler before encoding, proportionally scaling the original resolution of the video track to a size matching the resolution of the target playback terminal, and dynamically adjusts the target bitrate of the video encoder according to the bitrate capability parameters of the target playback terminal. If the target playback terminal reports that its maximum supported bitrate is lower than the original bitrate of the current video footage, the video generation system adjusts the target bitrate of the video encoder to the maximum bitrate value supported by the terminal, while maintaining the image quality within an acceptable range by adjusting the quantization parameters of the encoder.

[0060] As an alternative implementation, after completing the video encoding output, the video generation system performs a content security verification process on the generated personalized video file. The system inputs the audio track from the video file into a speech recognition engine to convert it into text. The converted text is then input into a sensitive word detection filter for scanning for prohibited words. Simultaneously, keyframe images from the video file are input into a visual content review model for detecting inappropriate visual elements. When neither audio / text detection nor visual detection detects any anomalies, the video generation system marks the video file as distributable. If an anomaly is detected at either stage, the system marks the video file as awaiting manual review, suspends the automatic distribution process, and generates a review alarm notification that is pushed to the management terminal. The review alarm notification includes the time and location of the detected anomaly and an anomaly type identifier, allowing administrators to quickly locate the problematic content.

[0061] This application embodiment aggregates multi-source data of the target object from a data source to construct a data profile of the target object. The data profile is then input into a story generation model to generate a storyboard containing at least two scenes arranged according to a timeline, with each scene associated with a scene description, narration text, and / or emotional tags. Then, based on the emotional tags and scene descriptions in the storyboard, at least two media generation engines are invoked in parallel to generate voice-over, background music, animation materials, and / or stylized materials, respectively. Finally, based on the timeline of the storyboard, the generated materials are combined into a personalized video of the target object. This allows the voice-over, background music, animation materials, and stylized materials to be generated synchronously for the specific emotional needs of each scene, thereby eliminating the efficiency bottleneck caused by serial generation. At the same time, scene-level emotional and content constraints ensure the stylistic coordination and narrative coherence of media assets between different scenes.

[0062] Based on the same inventive concept, this application also provides a second embodiment, referring to... Figure 2 , Figure 2 This is a flowchart illustrating the second embodiment of the video generation method based on data profiling in this application.

[0063] In this embodiment, the video generation method based on data profiling further includes steps S21 to S24: Step S21: Calculate the content fingerprint of the storyboard by performing a hash operation on the storyboard; Step S22: Send an authorization request to the user terminal; Step S23: Before invoking the media generation engine, verify whether the content fingerprint of the current storyboard is consistent with the content fingerprint in the authorization instruction; Step S24: If the verification results are consistent, based on the emotional tags and scene descriptions in the storyboard, call at least two media generation engines in parallel to generate voice-over, background music, animation materials and / or stylized materials.

[0064] In this embodiment, after the storyboard is generated, the video generation system also needs to perform content fingerprint calculation and authorization verification on the storyboard. The content fingerprint refers to a fixed-length hash value obtained by hashing the complete content data of the storyboard. This hash value is unique and collision-resistant; that is, the probability of two different storyboard contents generating different hash values ​​is extremely low, and the original content cannot be deduced from the hash value. The content fingerprint generates an unforgeable digital digest of the storyboard's content state, serving as the basis for integrity verification.

[0065] Specifically, after completing the storyboard's timeline verification and scene sequence assembly, the video generation system uses the storyboard's complete data structure as input data for hash operations. The system extracts the scene description field, narration text field, emotion tag field, and recommended duration field values ​​for all scenes from the storyboard data structure. These field values ​​are then concatenated into a continuous byte sequence according to the scene number. A secure hash algorithm is applied to this byte sequence to generate a fixed-length hash value as the content fingerprint. The video generation system then encapsulates the storyboard's preview content and the calculated content fingerprint into an authorization request message, which is sent to the user's terminal, such as the terminal of the target object's guardian, via push notification service.

[0066] Optionally, after receiving the authorization request message, the user terminal displays a storyboard preview interface on its screen. The preview interface shows scene descriptions, narration text, and emotional tags for each scene in sequence, and simultaneously displays a brief representation of the content fingerprint for the guardian to verify. After confirming the storyboard content is correct, the user performs a signature operation on the terminal, triggering the user terminal to return an authorization instruction message containing the content fingerprint to the video generation system. Upon receiving the authorization instruction message, the video generation system extracts the content fingerprint field value from the message, and simultaneously reads the complete storyboard data stored locally and recalculates its hash value to obtain the current content fingerprint.

[0067] Furthermore, the video generation system compares the content fingerprint in the authorization command with the current content fingerprint byte by byte. If the two fingerprints match exactly, it indicates that the storyboard content has not been modified since authorization, and the video generation system marks the authorization verification status as passed, subsequently triggering the media generation steps. If the two fingerprints do not match, it indicates that the storyboard content has been modified after authorization, and the video generation system refuses to execute the subsequent media generation steps, marks the task status as authorization invalid, and generates an authorization invalidation notification to be pushed to the user terminal and management terminal.

[0068] Optionally, a preview interface of the storyboard can be displayed to the target audience or their guardian via a user terminal, and the target emotional intensity value for at least one scene can be received. Based on the received target emotional intensity value, a target emotional curve is constructed, with time on the horizontal axis and emotional intensity on the vertical axis. The target emotional curve is then used as a constraint and input into the story generation model to allow the model to adjust the duration of each scene in the storyboard or merge adjacent scenes.

[0069] Specifically, after the initial generation of the storyboard and preliminary verification of the timeline are completed, the preview data of the storyboard is pushed to the user terminal. The preview data includes scene description text, narration text, and initial emotion tags for each scene, and an emotion intensity adjustment control is set below the preview area of ​​each scene. The target audience or their guardian can use this control to specify a desired emotion intensity value for each scene. The value ranges from a continuous interval, with a larger value indicating a stronger emotion that the user wants to present in the final video. After setting the emotion intensity for all scenes, the user terminal encapsulates the mapping relationship between all scene identifiers and their corresponding emotion intensity values ​​into a feedback message and returns it to the video generation system. After receiving the feedback message, the video generation system extracts each scene identifier and its corresponding target emotion intensity value. The video generation system divides the timeline into multiple continuous time intervals, with the start time of the storyboard timeline as zero, according to the recommended duration of each scene. Each time interval corresponds to one scene. The video generation system assigns the target sentiment intensity value for each scene to the corresponding time interval, plotting a discrete point sequence on a time axis coordinate system with the midpoint of the time interval as the x-axis and the sentiment intensity value as the y-axis. The system then interpolates and fits this discrete point sequence to generate a continuously differentiable target sentiment curve on the time axis. This target sentiment curve is then input into the story generation model.

[0070] The story generation model uses the target sentiment curve as a constraint for duration allocation and scene merging. It calculates the baseline sentiment intensity value corresponding to the initial sentiment tag of each scene in the current storyboard. This baseline value is compared with the expected value at the corresponding time position on the target sentiment curve. If the baseline sentiment intensity value of a scene is lower than the expected value at the corresponding position on the target curve, and the difference exceeds a preset deviation tolerance, the model increases the recommended duration of that scene. If the baseline sentiment intensity value of a scene is higher than the expected value at the corresponding position on the target curve, and the difference exceeds a preset deviation tolerance, the model compresses the recommended duration of that scene. When the difference between the expected values ​​of two adjacent scenes at corresponding positions on the target sentiment curve is less than a preset merging threshold, the model merges these two scenes into one. The description and narration of the merged scene are then regenerated by the model.

[0071] Since the system described in Embodiment 2 of this application is a system used to implement the method of Embodiment 1 of this application, those skilled in the art can understand the specific structure and variations of the system based on the method described in Embodiment 1 of this application, and therefore will not be described again here. All systems used in the method of Embodiment 1 of this application fall within the scope of protection of this application.

[0072] Based on the same inventive concept, this application also provides a third embodiment, referring to... Figure 3 , Figure 3This is a flowchart illustrating the third embodiment of the video generation method based on data profiling in this application.

[0073] In this embodiment, the video generation method based on data profiling further includes steps S31 to S33: Step S31: Extract the emotional tags and scene descriptions of each scene in the storyboard, and calculate the global style vector of the storyboard; Step S32: Input the global style vector into the music generation engine and the style transfer engine to ensure that the background music generated by the music generation engine maintains thematic consistency between adjacent scenes, and to make the style transfer engine apply a unified visual art style to adjacent scenes. Step S33: Generate transition audio clips or transition visual effects based on the difference in emotional tag vectors between the two scenes.

[0074] In this embodiment, while sending the emotion tag sequence and scene description to the music generation engine and style transfer engine respectively, the video generation system additionally calculates a global style vector and synchronously inputs it into both engines as supplementary input. The global style vector provides a unified style benchmark for the music generation engine and style transfer engine, allowing each engine to refer to this benchmark for conditional constraints when generating media materials for different scenes, thereby avoiding style jumps caused by independent generation of each scene. It is a multi-dimensional vector obtained by aggregating the vectorized representations of emotion tags for all scenes in the storyboard. This vector numerically expresses the central position and distribution range of the entire storyboard in the emotional space. A transitional audio segment refers to a short musical fragment that acts as a bridge between musical passages in two adjacent scenes, its function being to smoothly connect the music of the two scenes in terms of mode, tempo, and orchestration. A transitional visual effect refers to a visual effect with a gradual change in nature located between visual materials in two adjacent scenes.

[0075] Specifically, the video generation system extracts the sentiment tag field value for each scene from the storyboard and inputs each sentiment tag into a preset tag embedding model. This model maps discrete sentiment tag names to continuous vectors of fixed dimensions. The video generation system weights the sentiment tag vectors of all scenes according to the recommended duration of each scene; the longer the recommended duration, the greater the weight of the sentiment tag vector corresponding to the scene during aggregation. The video generation system sums and averages all the weighted sentiment tag vectors to obtain a global style vector of fixed dimensions. The video generation system sends this global style vector as an additional input, along with the sentiment tag sequence, to the music generation engine. When generating music clips for each scene, the music generation engine concatenates the sentiment tag vector of that scene with the global style vector to form a conditional vector. This conditional vector serves as a constraint for generating music clips, ensuring that the music clips for each scene maintain consistency with the global style vector in terms of modal tonic, rhythmic skeleton, and orchestration base. The video generation system also sends this global style vector as an additional input, along with the scene descriptions for each scene, to the style transfer engine.

[0076] The style transfer engine uses the mean hue and texture roughness in the global style vector as baseline style parameters, and visual style keywords in the scene description as local adjustment parameters, ensuring visual continuity in color saturation and brushstroke texture between stylized images of adjacent scenes. After generating the music clips and stylized materials for each scene, the video generation system further extracts the emotional tag vectors of two adjacent scenes and calculates the Euclidean distance between the two vectors as the difference. When the difference exceeds a preset transition threshold, the video generation system triggers the transition generation process. For the music track, the video generation system generates a transition audio clip based on the difference in emotional tag vectors between the two scenes. The key of this transition audio clip is taken as the middle key between the key of the previous clip and the key of the next clip, the tempo is taken as the middle tempo, and the orchestration is taken as the mixture ratio of the orchestration of the previous clip and the orchestration of the next clip, thus forming a natural auditory transition. For the video track, the video generation system generates a transition visual effect based on the difference in emotional tag vectors between the two scenes. The hue of this transition visual effect gradually changes from the hue of the previous clip to the hue of the next clip.

[0077] This application provides a video generation device based on data profiling. The device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the video generation method based on data profiling in Embodiment 1 described above.

[0078] The following is for reference. Figure 4The diagram illustrates a structural schematic of a video generation device based on data profiling suitable for implementing embodiments of this application. The video generation device based on data profiling in these embodiments may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 4 The video generation device based on data profiling shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0079] like Figure 4 As shown, the video generation device based on data profiling may include a processing unit 1001 (e.g., a core processor, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 1002 or a program loaded from storage device 1003 into random access memory (RAM) 1004. The random access memory 1004 also stores various programs and data required for the operation of the video generation device based on data profiling. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the data-based video generation device to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows a data-based video generation device with various systems, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems can be implemented alternatively.

[0080] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0081] The video generation device based on data profiling provided in this application, employing the video generation method based on data profiling in the above embodiments, can solve the technical problem of insufficient coordination between scenes in the generated video content. Compared with the prior art, the beneficial effects of the video generation device based on data profiling provided in this application are the same as those of the video generation method based on data profiling provided in the above embodiments, and other technical features in this video generation device based on data profiling are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0082] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0083] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0084] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the video generation method based on data profiling in the above embodiments.

[0085] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination thereof.

[0086] The aforementioned computer-readable storage medium may be included in a video generation device based on data profiling; or it may exist independently and not be assembled into a video generation device based on data profiling.

[0087] The aforementioned computer-readable storage medium carries one or more programs that, when executed by a video generation device based on a data profile, cause the video generation device based on a data profile to: aggregate multi-source data of a target object from a data source, construct a data profile of the target object, input the data profile into a story generation model, and generate a storyboard, wherein the storyboard contains at least two scenes arranged according to a timeline, the scenes being associated with scene descriptions, narration text, and / or emotional tags; based on the emotional tags and scene descriptions of the scenes in the storyboard, at least two media generation engines are invoked in parallel to generate voice-over, background music, animation materials, and / or stylized materials; and based on the timeline of the storyboard, the generated voice-over, background music, animation materials, and stylized materials are synthesized into a personalized video of the target object.

[0088] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0089] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0090] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0091] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described video generation method based on data profiling. This addresses the technical problem of insufficient coordination between scenes in the generated video content. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the video generation method based on data profiling provided in the above embodiments, and will not be elaborated upon here.

[0092] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A video generation method based on data profiling, characterized in that, The method includes the following steps: Aggregate multi-source data of the target object from data sources to construct a data profile of the target object; The data profile is input into the story generation model to generate a storyboard, which contains at least two scenes arranged according to a timeline, and the scenes are associated with scene descriptions, narration text and / or emotional tags; Based on the emotional tags and scene descriptions of the scene in the storyboard, at least two media generation engines are invoked in parallel to generate voice-over, background music, animation materials and / or stylized materials; Based on the timeline of the storyboard, the generated voice-over, background music, animation materials, and stylized materials are combined into a personalized video of the target object.

2. The video generation method based on data profiling as described in claim 1, characterized in that, The step of aggregating multi-source data of the target object from the data source to construct a data profile of the target object includes: The multi-source data is formed by collecting the sports rating records, growth record texts, campus activity materials and / or honor records of the target object from the data source. The multi-source data is stored in a structured manner based on a preset data model, which includes data type, timestamp, and content summary fields. The multi-source data, after being stored in a structured manner, is integrated into a data profile of the target object.

3. The video generation method based on data profiling as described in claim 1, characterized in that, The step of inputting the data profile into the story generation model to generate a storyboard includes: The data profile and the preset narrative template are input into the large language model. The narrative template defines the range of the number of scenes, the upper limit of the total duration, and the requirements for the emotional direction of the storyboard. Receive the target storyboard output by the large language model, wherein the target scene in the target storyboard includes scene identifier, scene description, narration text, sentiment tag and recommendation duration field; Perform timeline verification on the target storyboard, and select the target storyboard as the storyboard based on the verification results.

4. The video generation method based on data profiling as described in claim 1, characterized in that, The step of generating voice-over, background music, animation materials, and / or stylized materials by invoking at least two media generation engines in parallel based on the emotional tags and scene descriptions of the scene in the storyboard includes: The narration text for each scene in the storyboard is transmitted sequentially to the speech synthesis engine to generate dubbing audio clips corresponding to each scene; The emotional tag sequences of each scene in the storyboard are transmitted to the music generation engine to generate a background music track that matches the emotional tag sequences; The scene descriptions of each scene in the storyboard are transmitted to the style transfer engine, and a visual style matching the scene description is applied to the original photo or video material of the target object to generate stylized material.

5. The video generation method based on data profiling as described in claim 1, characterized in that, The step of combining the generated voice-over, background music, animation materials, and stylized materials into a personalized video of the target object according to the timeline of the storyboard includes: Based on the arrangement order and recommended duration of each scene in the storyboard, the dubbing audio segments are segmented, and the segmented dubbing audio segments are aligned with the timeline; The background music track is overlaid onto the aligned dubbing audio track, and the background music volume is adjusted to avoid noise. The animation and stylized materials are spliced ​​together in chronological order; After inserting transition effects between adjacent materials, the audio track, the dubbing track, and the background music track are merged and output as the personalized video.

6. The video generation method based on data profiling as described in claim 1, characterized in that, After the step of inputting the data profile into the story generation model to generate a storyboard, the method further includes: The content fingerprint of the storyboard is calculated by performing a hash operation on the storyboard; Send an authorization request to the user terminal, the authorization request containing the content fingerprint, and receive an authorization instruction containing the content fingerprint returned by the guardian; Before invoking the media generation engine, verify whether the content fingerprint of the current storyboard matches the content fingerprint in the authorization instruction; If the verification results are consistent, the step of generating voice-over, background music, animation materials and / or stylized materials by calling at least two media generation engines in parallel based on the emotional tags and scene descriptions of the scene in the storyboard is executed.

7. The video generation method based on data profiling as described in claim 1, characterized in that, Before generating voice-over, background music, animation assets, and / or stylized assets by simultaneously invoking at least two media generation engines based on the emotional tags and scene descriptions of the scene in the storyboard, the process further includes: The storyboard preview interface is displayed to the user terminal, and the target emotional intensity value is received from the user terminal for at least one scene. Based on the received target emotion intensity value, a target emotion curve is constructed, where the horizontal axis of the target emotion curve represents time and the vertical axis represents emotion intensity. The target emotion curve is used as a constraint and input into the story generation model so that the story generation model can adjust the duration of each scene in the storyboard or merge adjacent scenes.

8. The video generation method based on data profiling as described in claim 1, characterized in that, The steps of generating voiceovers, background music, animation assets, and stylized assets by calling at least two media generation engines in parallel also include: Extract the emotional tags and scene descriptions of each scene in the storyboard, and calculate the global style vector of the storyboard; The global style vector is input into the music generation engine and the style transfer engine, so that the background music generated by the music generation engine maintains thematic consistency between adjacent scenes, and the style transfer engine applies a unified visual art style to adjacent scenes. Based on the difference in the emotional tag vectors between the two scenes, generate transitional audio clips or transitional visual effects.

9. A video generation device based on data profiling, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the video generation method based on data profiling as described in any one of claims 1 to 8.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the video generation method based on data profiling as described in any one of claims 1 to 8.