Data video generation method and apparatus, device, and medium
By automatically extracting target data sets and generating visualizations, narratives, and animations through a data analysis intelligent agent, the problem of complex and time-consuming data video production is solved, enabling non-professional users to generate efficient and high-quality data videos.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HONG KONG UNIV OF SCI & TECH (GUANGZHOU)
- Filing Date
- 2026-01-09
- Publication Date
- 2026-04-21
AI Technical Summary
The production process of data videos in existing technologies is complex and time-consuming, requiring collaboration among multiple parties. Non-professional users find it difficult to understand user intent and target audience, resulting in a high barrier to entry and an inability to automatically discover valuable story points.
By acquiring user description data and raw data tables through a data analysis intelligent agent, the target data set is automatically extracted, and visualization, narrative, and animation configurations are generated. Multi-dimensional quality assessment and iterative optimization are then performed to achieve end-to-end automatic generation of data videos.
It significantly lowers the barrier to entry and time cost of creating data videos, enabling non-professional users to easily create high-quality data videos and ensuring semantic consistency and temporal coherence of visualization, narrative, and animation.
Smart Images

Figure CN121482220B_ABST
Abstract
Description
Technical Field
[0001] This application relates to, but is not limited to, the field of data video generation technology, and in particular to a data video generation method, apparatus, device, and medium. Background Technology
[0002] Data video, as an emerging data storytelling medium, is widely used in business, education, social media, and other fields. By integrating multimodal elements such as visualization, animation, narrative, and audio, data video can convey data insights in a vivid and intuitive way, attracting viewers' attention and enhancing information recall. However, the production process of data video is extremely complex and time-consuming, requiring data analysts to design visualizations, animators to create motion effects, writers to write narratives, audio engineers to handle voice-overs, and multimedia experts to coordinate all components. The entire process typically involves iterative iterations and multi-party collaboration. Related technologies require users to prepare intermediate products such as static visualizations, narrative text, or annotated narration in advance, rather than starting directly from raw data and natural language requirements. This makes it difficult to help users automatically discover valuable story points from data and understand the user's true intentions and target audience, thus creating a high barrier to entry for non-professional users. Summary of the Invention
[0003] This application provides a data video generation method, apparatus, device, and medium that can achieve end-to-end automatic generation of data videos, significantly reducing the creation threshold and time cost.
[0004] In a first aspect, embodiments of this application provide a data video generation method, the data video generation method comprising:
[0005] Obtain user description data and raw data tables;
[0006] Based on the user description data, data is extracted from the original data table to obtain a target data set, and visualization configuration, narrative configuration, and animation configuration are generated according to the target data set. The target data set includes multiple data groups, and each data group includes data type, data content, data evidence, and importance score.
[0007] Render the first target video corresponding to the original data table according to the visualization configuration, the narrative configuration, and the animation configuration;
[0008] The quality of the first target video is assessed, and the assessment results are obtained.
[0009] The first target video is optimized based on the evaluation results to obtain the second target video.
[0010] The data video generation method according to the first aspect of this application has at least the following beneficial effects: After acquiring user description data and a raw data table, a data analysis agent jointly understands the user's intent and data content, automatically extracting a target data set related to the user's needs. Visualization configurations, narrative configurations, and animation configurations are generated based on the target dataset, and a first target video is generated according to the configurations. Finally, the first target video undergoes multi-dimensional quality evaluation and iterative optimization until a high-quality second target video is generated. Through an end-to-end automated process generation strategy, the threshold and time cost of data video creation are significantly reduced, enabling non-professional users to easily create high-quality data videos. A multi-configuration-driven intermediate representation mechanism ensures semantic consistency and temporal coordination among visualization, narrative, and animation. A multi-dimensional feedback loop and quality control mechanism guarantee the high quality and reliability of the final generated video.
[0011] According to some embodiments of the first aspect of this application, the step of extracting data from the original data table based on the user description data to obtain a target data set includes:
[0012] Based on the user description data, the user's target intent is determined;
[0013] Data is extracted from the original data table based on the user description data to obtain trend data, abnormal data, comparison data, and correlation data.
[0014] The trend data, the abnormal data, the comparison data, and the relevance data are assigned importance scores based on the user's target intent, and the trend data, the abnormal data, the comparison data, the relevance data, and the importance scores of each data point are used as a target data set.
[0015] According to some embodiments of the first aspect of this application, the step of generating the visualization configuration, narrative configuration, and animation configuration based on the target data set includes:
[0016] Based on the data types in all the data groups, determine the chart type; based on the chart type and the data groups, generate the visualization configuration;
[0017] Scene planning is performed based on the importance scores within all data groups and the logical relationships between each data group to determine the scene type, playback order, and video duration. Based on the scene content constructed by the visualization configuration, narration text is generated for all scene content, and the link relationship between the narration text and visual elements is established to obtain the narrative configuration.
[0018] Based on the narrative configuration and visualization configuration, entrance animation, progressive animation of data elements, emphasis animation, annotation animation, and scene transition animation are generated for the scene content to obtain the animation configuration.
[0019] According to some embodiments of the first aspect of this application, rendering the first target video corresponding to the original data table according to the visualization configuration, the narrative configuration, and the animation configuration includes:
[0020] The visualization configuration is parsed to obtain the chart rendering code;
[0021] Generate a visual chart based on the chart rendering code;
[0022] Based on the visual elements in the visualization chart, an element registry is established; wherein, the element registry records the element identifier, element type, data filter, raw data index and bounding box information of each visual element;
[0023] The data filters in the narrative configuration and the animation configuration are parsed based on the element registry;
[0024] Based on the parsed narrative configuration and animation configuration, a first target video is generated.
[0025] According to some embodiments of the first aspect of this application, the step of performing a quality assessment on the first target video to obtain an assessment result includes:
[0026] The data accuracy of the first target video is evaluated to obtain a data accuracy score;
[0027] A visual quality assessment is performed on the first target video to obtain a visual quality score;
[0028] The narrative coherence of the first target video is evaluated to obtain a narrative coherence score;
[0029] The component coordination is evaluated on the first target video to obtain a component coordination score;
[0030] The temporal rationality of the first target video is evaluated to obtain a temporal rationality score;
[0031] The evaluation result is obtained by integrating the data accuracy score, the visualization quality score, the narrative coherence score, the component coordination score, and the temporal rationality score.
[0032] According to some embodiments of the first aspect of this application, optimizing the first target video based on the evaluation result to obtain a second target video includes:
[0033] Based on the evaluation results, determine whether the first target video needs to be optimized;
[0034] If the first target video needs to be optimized, feedback information is generated and the visualization configuration, narrative configuration, and animation configuration are updated according to the feedback information. Then, the second target video is generated according to the updated visualization configuration, narrative configuration, and animation configuration.
[0035] If no optimization is required for the first target video, then the first target video will be output as the second target video.
[0036] According to some embodiments of the first aspect of this application, optimizing the first target video based on the evaluation result to obtain a second target video includes:
[0037] If the data accuracy score or the visualization quality score is lower than a preset first threshold, then first feedback information for the visualization design is generated, and the visualization configuration is updated.
[0038] If the narrative coherence score is lower than a preset second threshold, then a second feedback message for the narrative design is generated, and the narrative configuration is updated.
[0039] If the component coordination score or the timing rationality score is lower than a preset third threshold, then a third feedback message for animation coordination is generated, and the animation configuration is updated.
[0040] A second target video is generated based on the updated visualization configuration, narrative configuration, and animation configuration, and a quality assessment is performed again.
[0041] Secondly, embodiments of the present invention provide an operation control device for implementing the data video generation method provided in the second aspect embodiments.
[0042] Thirdly, embodiments of the present invention provide an electronic device including the operation control device provided in the second aspect embodiments above.
[0043] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing computer-executable instructions for causing a computer to perform the data video generation method described in the first aspect of the embodiments above. Attached Figure Description
[0044] The accompanying drawings are provided to further understand the technical solutions of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the technical solutions of the present invention, and do not constitute a limitation on the technical solutions of the present invention.
[0045] Figure 1 A detailed flowchart of the data video generation method provided in the embodiments of this application;
[0046] Figure 2 for Figure 1 The detailed flowchart of step S200;
[0047] Figure 3 for Figure 1 Another specific flowchart of step S200;
[0048] Figure 4 for Figure 1 The detailed flowchart of step S300;
[0049] Figure 5 for Figure 1 The detailed flowchart of step S400;
[0050] Figure 6 A schematic diagram of the operation control device is provided for the embodiments of this application. Detailed Implementation
[0051] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0052] It is understandable that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, or the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0053] Data video, as an emerging data storytelling medium, is widely used in business, education, social media, and other fields. By integrating multimodal elements such as visualization, animation, narrative, and audio, data video can convey data insights in a vivid and intuitive way, attracting viewers' attention and enhancing information recall. However, the production process of data video is extremely complex and time-consuming, requiring data analysts to design visualizations, animators to create motion effects, writers to write narratives, audio engineers to handle voice-overs, and multimedia experts to coordinate all components. The entire process typically involves iterative iterations and multi-party collaboration. Related technologies require users to prepare intermediate products such as static visualizations, narrative text, or annotated narration in advance, rather than starting directly from raw data and natural language requirements. This makes it difficult to help users automatically discover valuable story points from data and understand the user's true intentions and target audience, thus creating a high barrier to entry for non-professional users.
[0054] Based on this, embodiments of this application provide a data video generation method, apparatus, device, and medium. After acquiring user description data and a raw data table, a data analysis agent jointly understands the user's intent and data content, automatically extracting a target data set related to the user's needs. Visualization configurations, narrative configurations, and animation configurations are generated based on the target dataset, and a first target video is generated according to these configurations. Finally, the first target video undergoes multi-dimensional quality evaluation and iterative optimization until a high-quality second target video is generated. This end-to-end automated process generation strategy significantly reduces the threshold and time cost of data video creation, enabling non-professional users to easily create high-quality data videos. A multi-configuration-driven intermediate representation mechanism ensures semantic consistency and temporal coordination among visualization, narrative, and animation. A multi-dimensional feedback loop and quality control mechanism guarantee the high quality and reliability of the final generated video.
[0055] Firstly, referring to Figure 1 , Figure 1 A detailed flowchart of the data video generation method provided in this application embodiment includes, but is not limited to, the following steps:
[0056] Step S100: Obtain user description data and raw data table;
[0057] Step S200: Extract data from the original data table based on user description data to obtain the target data set, and generate visualization configuration, narrative configuration and animation configuration based on the target data set;
[0058] Step S300: Render the first target video corresponding to the original data table according to the visualization configuration, narrative configuration, and animation configuration;
[0059] Step S400: Perform a quality assessment on the first target video to obtain the assessment result;
[0060] Step S500: Optimize the first target video based on the evaluation results to obtain the second target video.
[0061] It's understandable that user description data consists of natural language descriptions of the initial icon data, while the original data table is a structured data table. User description data helps determine the user's data analysis needs, allowing data extraction from the original data table to ensure the data sets in the target data set align with the user's intent. Videos are composed of multiple frames, each containing both visual and textual content. The continuity of the video requires connecting these frames through animation. Therefore, visualization, narrative, and animation configurations can be generated based on the target data set. This separation avoids the complexity of traditional methods that separate the storage of visualization, narrative, and animation, allowing for scene-by-scene generation and correction during subsequent optimization, reducing cognitive load and improving optimization efficiency. Furthermore, since the visualization, narrative, and animation configurations are generated based on the target data set, which is derived from user description data extracted from the original data table, the accuracy and reliability of these configurations are effectively improved. After rendering the first target video corresponding to the original data table based on the visualization configuration, narrative configuration, and animation configuration, the quality of the first target video can be evaluated, and then optimized based on the evaluation results to make the second target video better meet the user's needs.
[0062] Reference Figure 2 , Figure 2 for Figure 1 The detailed flowchart of step S200 includes, but is not limited to, the following steps:
[0063] Step S210: Determine the user's target intent based on the user description data;
[0064] Step S220: Extract data from the original data table based on the user description data to obtain trend data, abnormal data, comparison data, and correlation data;
[0065] Step S230: Based on the user's target intent, the trend data, abnormal data, comparative data, and related data are scored for importance, and the trend data, abnormal data, comparative data, and related data, along with the importance scores of each data, are used as the target data set.
[0066] Understandably, after obtaining user description data, since this data is presented in the user's natural language, we can first determine the user's target intent based on it. For example, if the user description is "Analyze sales trends in various regions from 2020 to 2023, focusing on the fastest-growing regions," then we can determine the user's target intent as sales trends, the time range as 2020-2023, the analysis dimension as region, and the specific focus as the fastest-growing regions. Based on the user description data, we extract data from the original data table to obtain trend data (identifying rising, falling, or fluctuating patterns), outlier data (finding data points outside the normal range), comparative data (comparing differences between different categories or time periods), and correlation data (the relationships between variables). Extracting data from the original data table yields multiple data points, but the importance of each data point varies. Therefore, we can assign importance scores to trend data, outlier data, comparative data, and correlation data based on the user's target intent. The target data set consists of multiple data groups. Each data group includes a data identifier, data type, data content, data evidence, and importance score. The data identifier is a unique identifier used to mark the data group to ensure its uniqueness. Data types include trends, outliers, comparisons, and correlations. Data content is a natural language description of the data. Data evidence consists of the specific data in the original data table supporting the data group. The importance score is the degree of relevance of the data group to the user's target intent. The higher the relevance to the user's target intent, the more important the user values the data, and the higher the importance score. In addition, all data groups can be sorted and filtered based on their importance scores.
[0067] Specifically, the user's description data is "Analyze sales trends in various regions from 2020 to 2023, focusing on the fastest-growing regions." Therefore, the user's target intent is sales trends, the time range is 2020-2023, the analysis dimension is region, and the specific focus is the fastest-growing regions. The original data table contains the following fields: year, region, sales revenue, and growth rate. Based on the user's description data, the following data was extracted from the original data table: Trend Data 1: The Southern region maintained the highest growth rate for three consecutive years (12%-18%); Comparison Data: Sales revenue in the Southern region increased from 8 million to 15 million, showing the largest increase; Trend Data 2: All regions showed an overall growth trend, but the growth rates differed significantly. Since the user's specific focus is on the fastest-growing regions, the importance scores of Trend Data 1 and Comparison Data are higher than those of Trend Data 2.
[0068] Reference Figure 3 , Figure 3 for Figure 1 Another specific flowchart for step S200 includes, but is not limited to, the following steps:
[0069] Step S240: Determine the chart type based on the data types in all data groups;
[0070] Step S250: Generate visualization configuration based on chart type and data group;
[0071] Step S260: Based on the importance scores within all data groups and the logical relationships between each data group, scenario planning is performed to determine the scenario type, playback order, and video duration. Based on the scenario content constructed by the visual configuration, narration text is generated for all scenario content, and the link relationship between the narration text and visual elements is established to obtain the narrative configuration.
[0072] Step S270: Based on the narrative configuration and visualization configuration, generate entrance animation, progressive animation of data elements, emphasis animation, annotation animation and scene transition animation for the scene content to obtain the animation configuration.
[0073] Understandingly, visualization configuration includes visualization labels, chart types, titles, data binding, data transformation, and style settings. Visual labels represent the visualization configuration; chart types are determined based on the data types in the data groups, for example, line charts for trends and bar charts for comparisons; data binding maps fields in the initial chart to visual channels, for example, the horizontal axis represents years and the vertical axis represents sales revenue; data transformation includes grouping, aggregation, and filtering; and style settings include color schemes and layout parameters. Since different data groups have different importance scores and different logical relationships, directly planning scenarios based on data groups would lead to scenario confusion. Therefore, scenario planning can be based on the importance scores of all data groups and their logical relationships, determining the scenario type, playback order, and video duration. Then, based on the scenario content constructed by the visualization configuration, explanatory text is generated for all scenario content, and links are established between the explanatory text and visual elements to obtain the narrative configuration; finally, based on the narrative configuration and visualization configuration, entrance animations, progressive animations of data elements, emphasis animations, and annotation animations are generated for the scenario content to obtain the animation configuration.
[0074] It should be noted that when establishing the link between explanatory text and visual elements, data filters can be used to abstractly reference the visual elements. These data filters are expressions describing data conditions, such as "Region = South District". During the rendering phase, the abstract references are resolved into specific visual elements based on the data filters. This abstract referencing mechanism eliminates the need to know the specific visual elements in advance during the configuration generation phase, thus improving the system's flexibility.
[0075] It should be noted that the entrance animation is the overall entrance design animation for the chart (such as fade-in, zoom); the data element progressive animation can be designed with targeted effects according to the chart type (such as drawing the path of a line chart, increasing or sequentially appearing the bars of a bar chart, and showing the data points one by one in a scatter plot), and the delay time and animation duration of each element are calculated based on the amount of data; the emphasis animation is designed to emphasize narrative segments of specific data (such as highlighting, glowing, zooming); the annotation animation is designed to provide supplementary explanations (such as displaying labels, drawing auxiliary lines); and the scene transition animation is a multi-chart video, designing transition animations between charts.
[0076] Specifically, trend data 1: The Southern region has maintained the highest growth rate (12%-18%) for three consecutive years; comparative data: Sales in the Southern region increased from 8 million to 15 million, showing the largest increase; trend data 2: All regions show an overall growth trend, but the growth rates differ significantly. The visualization configuration includes: chart type: line chart (suitable for displaying trends), horizontal axis: year (time type), vertical axis: sales (numerical type), color: region (category type, creating multiple lines), data transformation: grouping by year and region. The narrative configuration includes:
[0077] Scene 1 (Opening scene, 0-3.0 seconds):
[0078] Scene type: Opening scene;
[0079] Content: Title: "Analysis of Sales Trends in Various Regions from 2020 to 2023", Subtitle: "Data-Driven Insights";
[0080] Narrative: Segment 1 (0-3.0 seconds) "Let's look at the sales performance in each region from 2020 to 2023."
[0081] Scenario 2 (Chart scenario, 3.0-15.0 seconds):
[0082] Scenario type: Chart scenario;
[0083] Content: Referencing visual configuration;
[0084] Narrative:
[0085] Segment 1 (3.0-7.5 seconds): "The South Region performed exceptionally well, with sales increasing from 8 million to 15 million."
[0086] Segment 2 (7.5-12.0 seconds): "Its growth rate has remained above 12% for three consecutive years, far exceeding other regions" (Because the importance scores of trend data 1 and comparison data are high, trend data 1 and comparison data can be selected for display in scenario 2, and trend data 1 and comparison data are related).
[0087] Scenario 3 (Statistical card scenario, 15.0-18.5 seconds):
[0088] Scenario type: Statistics card scenario;
[0089] Content: Three statistical cards display key indicators;
[0090] Card 1: The number "87.5%", labeled "Southern Region Growth Rate";
[0091] Card 2: The number "15 million", labeled "Southern Region Sales in 2023";
[0092] Card 3: Number "TOP 1", label "Southern Region Ranking";
[0093] Narrative: Segment 1 (15.0-18.5 seconds) "The South District ranked first with a total growth rate of 87.5%".
[0094] Scene 4 (Final scene, 18.5-21.0 seconds):
[0095] Scene type: Ending scene;
[0096] Content: Summary text: "Data Speaks Louder Than Words, Southern Region Leads the Way";
[0097] Narrative: Segment 1 (18.5-21.0 seconds) "This is the story behind the data."
[0098] The animation configuration can be as follows: (0-1.0 seconds): title fades in; (3.0-4.0 seconds): chart fades in; (4.0-6.5 seconds): each line is drawn in sequence; (7.5-12.0 seconds): the south line is highlighted; (7.5-12.0 seconds): growth rate label is displayed; (15.0-17.0 seconds): three cards pop up in sequence with a scrolling number effect; (18.5-20.0 seconds): summary text fades in and enlarges.
[0099] It's important to note that visual configuration, narrative configuration, and animation configuration are considered software design patterns. These patterns describe system behavior through declarative configuration files, rather than directly writing implementation code. The generation of visual, narrative, and animation configurations can employ multi-agent collaboration. Multi-agent collaboration is a system architecture that decomposes complex tasks into multiple specialized sub-tasks, with multiple agents each performing their specific duties and collaborating to complete the task. Multi-agent collaboration improves the system's modularity and task completion quality. Specifically, it involves using different agents to generate visual, narrative, and animation configurations.
[0100] Reference Figure 4 , Figure 4 for Figure 1 The detailed flowchart of step S300 includes, but is not limited to, the following steps:
[0101] Step S310: Analyze the visualization configuration to obtain the chart rendering code;
[0102] Step S320: Generate a visual chart based on the chart rendering code;
[0103] Step S330: Based on the visual elements in the visualization chart, establish an element registry;
[0104] Step S340: Parse the data filters in the narrative configuration and animation configuration based on the element registry;
[0105] Step S350: Generate the first target video based on the parsed narrative configuration and animation configuration.
[0106] Understandably, for chart scenarios, the system parses the referenced visualization configuration, then uses the chart type, data binding, data transformation, and style configuration within the visualization configuration to generate the corresponding chart rendering code. Based on the chart rendering code, visual elements are generated from a pre-defined visualization library, and an element registry for that scenario is established. The element registry records detailed information for each visual element, including: element identifier (unique identifier, such as `line_south`), element type (line, bar, region, etc.), data filter (data conditions corresponding to the element, such as "region = South"), raw data index (an array of raw data row indices corresponding to the element), and bounding box information (the element's position and size). This scenario-by-scenario approach to building the element registry allows the system to independently track and manage the visual elements for each scenario. The element registry employs a dual-indexing mechanism to record the mapping relationship between visual elements and data. This dual-indexing mechanism includes: a visual position index, used to locate specific visual elements in the rendered visualization chart; and a raw data row index, used to trace the raw data rows corresponding to the visual elements. This dual-indexing mechanism ensures data traceability even after data aggregation, ensuring that narratives and animations accurately point to the correct data and visual elements. The renderer parses the data filters in the narrative and animation configurations of the scene, resolving abstract references into concrete visual elements based on the scene's element registry. For example, the data filter "Region = South" in the narrative configuration is resolved to the `line_south` element in the scene's element registry. This scene-by-scene parsing method ensures that the narrative and animation of each scene are precisely linked to the scene's visual elements. Finally, a complete video timeline is generated based on the parsed narrative and animation configurations. The video timeline is a chronological sequence of events, including the start and end times of each scene, the playback time of the narration, and the trigger times of the animation effects. The narration (via text-to-speech or pre-recorded audio) and animation effects of each scene are triggered sequentially and chronologically to generate the first target video.
[0107] It should be noted that the renderer is responsible for converting declarative configurations into data videos. The renderer can enhance and encapsulate existing visualization libraries (such as D3.js and AntV) and animation libraries (such as GSAP and Remotion) to achieve functions such as configuration parsing, element registry construction, abstract reference resolution, generation of chart type-specific animation effects, and generation of video timelines.
[0108] Reference Figure 5 , Figure 5 for Figure 1 The detailed flowchart of step S400 includes, but is not limited to, the following steps:
[0109] Step S410: Evaluate the data accuracy of the first target video to obtain a data accuracy score;
[0110] Step S420: Perform a visual quality assessment on the first target video to obtain a visual quality score;
[0111] Step S430: Perform a narrative coherence assessment on the first target video to obtain a narrative coherence score;
[0112] Step S440: Perform component coordination evaluation on the first target video to obtain a component coordination score;
[0113] Step S450: Evaluate the temporal rationality of the first target video and obtain a temporal rationality score;
[0114] Step S460: The evaluation results are obtained by integrating the data accuracy score, visualization quality score, narrative coherence score, component coordination score, and temporal rationality score.
[0115] Understandably, evaluating the primary target video can accurately reflect the data. Data accuracy assessment verifies the correctness of data transformation operations and confirms that the numbers and facts in the narrative text are consistent with the data. For example, checking whether the sales figures displayed in the charts are consistent with the original data, and checking whether the growth rate mentioned in the narrative text is accurate. Visualization quality assessment checks whether visual coding conforms to best practices and verifies whether the data-to-ink ratio is reasonable. For example, checking whether trend insights use line charts and whether color coding is easily distinguishable. Narrative coherence assessment judges logical coherence and story structure, and verifies whether the narrative pacing is reasonable. For example, checking whether the narration text is fluent and whether the logical relationships between insights are clear. Component coordination assessment verifies the synchronization between animation and narrative and assesses the semantic consistency between components. Timing rationality assessment checks whether the animation timeline and total video duration are reasonable. For example, checking whether the data mentioned in the narrative text is consistent with the highlighted visual elements, checking whether the animation trigger time is synchronized with the narration, and checking whether the animation speed and video duration are within a reasonable range. Then, the scores from each dimension are integrated and calculated to obtain a comprehensive evaluation result. A weighted average approach can be used, with weights set according to the importance of different dimensions. For example, the evaluation result = 0.3 * data accuracy score + 0.2 * visualization quality score + 0.2 * narrative coherence score + 0.3 * component coordination score. By conducting a multi-dimensional quality assessment of the primary target video, evaluating it from multiple dimensions such as data accuracy, visualization quality, narrative coherence, component coordination, and temporal rationality, the final generated video is ensured to be of high quality.
[0116] It should be noted that if the data accuracy score or visualization quality score is lower than the preset first threshold, first feedback information for visualization design is generated and the visualization configuration is updated; if the narrative coherence score is lower than the preset second threshold, second feedback information for narrative design is generated and the narrative configuration is updated; if the component coordination score or temporal rationality score is lower than the preset third threshold, third feedback information for animation coordination is generated and the animation configuration is updated; a second target video is generated based on the updated visualization configuration, narrative configuration, and animation configuration, and the quality is reassessed.
[0117] Specifically, if the data accuracy score or visualization quality score is lower than a preset first threshold, first feedback information for the visualization design is generated, and the visualization configuration is updated. The first feedback information includes specific problem descriptions and improvement suggestions, such as "The data points in the line chart are inconsistent with the original data; please check the data transformation operation" or "The color contrast of the bar chart is insufficient; it is recommended to adjust the color scheme." If the narrative coherence score is lower than a preset second threshold, second feedback information for the narrative design is generated, and the narrative configuration is updated. The feedback information includes specific problem descriptions and improvement suggestions, such as "There are grammatical errors in the narration text; please correct them" or "The logical relationship between insights is unclear; it is recommended to adjust the narrative order." If the component coordination score or timing rationality score is lower than a preset third threshold, third feedback information for animation coordination is generated, and the animation configuration is updated. The feedback information includes specific problem descriptions and improvement suggestions, such as "The animation trigger time is not synchronized with the narration; please adjust the animation timing" or "The animation speed is too fast; it is recommended to extend the animation duration." Based on the updated visualization configuration, narrative configuration, and animation configuration, the data video is re-rendered, and a quality assessment is performed again until all dimension scores reach the preset thresholds or the maximum number of iterations is reached. The maximum number of iterations can be set to 3 or 5 rounds to prevent infinite loops.
[0118] Understandably, if all dimensions of the score reach the preset threshold, it means that there is no need to optimize the first target video, and the first target video can be output as the second target video.
[0119] Secondly, referring to Figure 6 This application provides an operation control device 600, including a memory 610, a processor 620, and a computer program stored in the memory 610 and executable on the processor 620. The processor 620 executes the program to implement the data video generation method of the first aspect embodiment above, for example, by executing... Figure 1 Method steps S100 to S500 Figure 2 Method steps S210 to S230, Figure 3 Method steps S240 to S270, Figure 4Method steps S310 to S350 and Figure 5 Method steps S410 to S460.
[0120] The memory 610, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs, including the data video generation method in the above embodiments of this application. The processor 620 implements the data video generation method in the above embodiments of this application by running the non-transitory software program and instructions stored in the memory 610.
[0121] The memory 610 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data required for executing the data video generation method described in the above embodiments. Furthermore, the memory 610 may include a high-speed random access memory 610, and may also include non-transitory memory 610, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. It should be noted that the memory 610 may include remotely located memories 610 relative to the processor 620, and these remote memories 610 can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0122] Thirdly, embodiments of this application provide an electronic device including an operation control device 600 as described in the second aspect embodiment. After acquiring user description data and a raw data table, the device uses a data analysis agent to jointly understand user intent and data content, automatically extracting a target data set related to user needs. Based on the target dataset, it generates visualization configurations, narrative configurations, and animation configurations, and generates a first target video based on these configurations. Finally, it performs multi-dimensional quality evaluation and iterative optimization on the first target video until a high-quality second target video is generated. This end-to-end automated process generation strategy significantly reduces the threshold and time cost of data video creation, enabling non-professional users to easily create high-quality data videos. A multi-configuration-driven intermediate representation mechanism ensures semantic consistency and temporal coordination among visualization, narrative, and animation. A multi-dimensional feedback loop and quality control mechanism guarantee the high quality and reliability of the final generated video.
[0123] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions for causing a computer to perform the data video generation method of the first aspect embodiment above, for example, executing... Figure 1 Method steps S100 to S500 Figure 2 Method steps S210 to S230, Figure 3 Method steps S240 to S270, Figure 4 Method steps S310 to S350 and Figure 5 The method steps S410 to S460 are described above. Those skilled in the art will understand that all or some of the steps and systems disclosed in the above methods can be implemented as software, firmware, hardware, and suitable combinations thereof. Some or all physical components can be implemented as processors, such as central processing units, digital signal processors, or microprocessors executing software, or as hardware, or as integrated circuits, such as application-specific integrated circuits (ASICs). Such software can be distributed on a computer-readable medium, which can include computer-readable storage media or non-transitory media and communication media or transient media. As is known to those skilled in the art, a computer-readable storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc DVD or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible by a computer. Furthermore, as is known to those skilled in the art, communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information transmission medium.
[0124] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0125] The above is a detailed description of the preferred embodiments of this application. However, this application is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this application. All such equivalent modifications or substitutions are included within the scope defined by the claims of this application.
Claims
1. A method of generating a data video, characterized by, The data video generation method includes: Obtain user description data and raw data tables; Based on the user description data, data is extracted from the original data table to obtain a target data set, and visualization configuration, narrative configuration, and animation configuration are generated according to the target data set. The target data set includes multiple data groups, and each data group includes data type, data content, data evidence, and importance score. Render the first target video corresponding to the original data table according to the visualization configuration, the narrative configuration, and the animation configuration; The quality of the first target video is assessed, and the assessment results are obtained. The first target video is optimized based on the evaluation results to obtain the second target video; The step of extracting data from the original data table based on the user description data to obtain the target data set includes: Based on the user description data, the user's target intent is determined; Data is extracted from the original data table based on the user description data to obtain trend data, abnormal data, comparison data, and correlation data. The trend data, abnormal data, comparative data, and relevance data are assigned importance scores based on the user's target intent. The trend data, abnormal data, comparative data, and relevance data, along with the importance scores of each data point, are then used as a target data set. The importance scores characterize the degree of correlation between each data point and the user's target intent.
2. The data video generation method of claim 1, wherein, The step of generating visualization configuration, narrative configuration, and animation configuration based on the target data set includes: The chart type is determined based on the data types in all the data groups. Generate a visualization configuration based on the chart type and the data group; Scene planning is performed based on the importance scores within all data groups and the logical relationships between each data group to determine the scene type, playback order, and video duration. Based on the scene content constructed by the visualization configuration, narration text is generated for all scene content, and the link relationship between the narration text and visual elements is established to obtain the narrative configuration. Based on the narrative configuration and visualization configuration, entrance animation, progressive animation of data elements, emphasis animation, annotation animation, and scene transition animation are generated for the scene content to obtain the animation configuration.
3. The data video generation method of claim 1, wherein, The step of rendering the first target video corresponding to the original data table according to the visualization configuration, the narrative configuration, and the animation configuration includes: The visualization configuration is parsed to obtain the chart rendering code; Generate a visual chart based on the chart rendering code; Based on the visual elements in the visualization chart, an element registry is established; wherein, the element registry records the element identifier, element type, data filter, raw data index and bounding box information of each visual element; The data filters in the narrative configuration and the animation configuration are parsed based on the element registry; Based on the parsed narrative configuration and animation configuration, a first target video is generated.
4. The data video generation method of claim 1, wherein, The quality assessment of the first target video, to obtain the assessment result, includes: The data accuracy of the first target video is evaluated to obtain a data accuracy score; A visual quality assessment is performed on the first target video to obtain a visual quality score; The narrative coherence of the first target video is evaluated to obtain a narrative coherence score; The component coordination is evaluated on the first target video to obtain a component coordination score; The temporal rationality of the first target video is evaluated to obtain a temporal rationality score; The evaluation result is obtained by integrating the data accuracy score, the visualization quality score, the narrative coherence score, the component coordination score, and the temporal rationality score.
5. The data video generation method of claim 1, wherein, The step of optimizing the first target video based on the evaluation result to obtain the second target video includes: Based on the evaluation results, determine whether the first target video needs to be optimized; If the first target video needs to be optimized, feedback information is generated and the visualization configuration, narrative configuration, and animation configuration are updated according to the feedback information. Then, the second target video is generated according to the updated visualization configuration, narrative configuration, and animation configuration. If no optimization is required for the first target video, then the first target video will be output as the second target video.
6. The data video generation method of claim 4, wherein, The step of optimizing the first target video based on the evaluation result to obtain the second target video includes: If the data accuracy score or the visualization quality score is lower than a preset first threshold, then first feedback information for the visualization design is generated, and the visualization configuration is updated. If the narrative coherence score is lower than a preset second threshold, then a second feedback message for the narrative design is generated, and the narrative configuration is updated. If the component coordination score or the timing rationality score is lower than a preset third threshold, then a third feedback message for animation coordination is generated, and the animation configuration is updated. A second target video is generated based on the updated visualization configuration, narrative configuration, and animation configuration, and a quality assessment is performed again.
7. A running control device characterized by comprising: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to implement the data video generation method as described in any one of claims 1 to 6.
8. An electronic device, comprising: Includes the operation control device as described in claim 7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to perform the data video generation method as described in any one of claims 1 to 6.