A short video production method and system

CN122601945APending Publication Date: 2026-08-18闽南科技学院
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610701151.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-20
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0005]因此,本发明所要解决的问题在于现有短视频制作中剧情逻辑生硬、情感表达单一、版权与风格割裂、缺乏真三维沉浸体验及无法实时自进化等问题

Benefits of technology

[0022]The beneficial effects of this invention are as follows: by constructing a "multi-dimensional causal narrative graph" and a "neuro-emotional resonance" system, it achieves deep intelligent transformation of short videos from logical deduction to emotional audiovisual generation; by utilizing "neural style fingerprints" to implicitly integrate copyright protection into personalized styles, it pioneers a "style is copyright" security mechanism; by combining "holographic light field" rendering technology to break through two-dimensional limitations and provide a naked-eye three-dimensional immersive experience; by relying on the "fluid video" protocol to achieve end-side adaptive and smooth delivery, and by leveraging the "collective intelligence feedback loop" to establish a real-time self-evolving closed loop, it ultimately achieves a comprehensive technological breakthrough in high content resonance, high copyright security, high immersive experience, high transmission reliability, and model self-iteration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122601945A_ABST
    Figure CN122601945A_ABST
Patent Text Reader

Abstract

The application discloses a short video production method and system, relates to the cross technical field of artificial intelligence and multimedia technology, and comprises a logic deduction engine based on a multi-dimensional causal narrative graph, introduces a cross-modal "neural emotional resonance" scheduling system, carries out dynamic copyright embedding and personalized generation based on a "neural style fingerprint", uses "holographic light field" fusion rendering technology, and realizes self-evolution creation based on a "group wisdom feedback loop" through an end-side adaptive "fluid video" delivery protocol. The application realizes the deep intelligentization of short videos from logic deduction to emotional audio-visual generation; the "neural style fingerprint" is used to invisibly fuse copyright protection into personalized style, the "holographic light field" rendering technology is combined to break the two-dimensional limitation and provide a naked-eye three-dimensional immersive experience; the "fluid video" protocol is used to realize end-side adaptive smooth delivery, and the "group wisdom feedback loop" is used to establish a real-time self-evolution closed loop.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the intersection of artificial intelligence and multimedia technology, and in particular to a method and system for producing short videos. Background Technology

[0002] With the acceleration of social digitalization and the popularization of mobile internet, short videos have become a core carrier for information dissemination, cultural entertainment, and commercial marketing. Their content production is undergoing a profound transformation from "human creation" to "AI-assisted generation." Currently, automated script generation based on large language models, deep learning-driven video editing, and personalized recommendation algorithms are widely used on major platforms, significantly lowering the creation threshold and improving distribution efficiency. At the same time, users' expectations for video content have upgraded from a simple two-dimensional audiovisual experience to a pursuit of immersive, interactive, and true 3D visual effects. Cutting-edge technologies such as holographic display, light field technology, and adaptive streaming media transmission protocols are beginning to emerge in high-end application scenarios, attempting to build a more intelligent, three-dimensional, and smooth digital content ecosystem.

[0003] Currently, existing short video production and distribution technologies still have significant limitations, making it difficult to meet the needs of next-generation immersive content. Firstly, at the content generation level, mainstream technologies largely rely on fixed linear script templates, lacking the ability to dynamically extrapolate the deep causal logic and complex emotional threads of a story. This results in videos with stiff emotional expression, simplistic plot logic, and an inability to adjust cross-modal audiovisual interactions based on real-time user emotions. Secondly, regarding copyright protection and personalization, traditional watermarking technologies are easily removed and can damage the aesthetic appeal of the image, making it difficult to achieve a deep integration of copyright information and artistic style, let alone dynamically generate personalized content based on user status. Thirdly, in the presentation and delivery stage, existing videos are still limited to two-dimensional planar display, lacking true three-dimensional depth of field, and traditional fixed bitrate transmission protocols are prone to stuttering when facing network fluctuations, unable to dynamically adjust rendering precision based on terminal computing power. Finally, existing systems generally lack closed-loop evolution mechanisms; model iteration relies on offline training and cannot utilize real-time feedback from global users for online self-optimization, causing content creation capabilities to lag behind the rapid evolution of user aesthetics. Summary of the Invention

[0004] In view of the problems existing in current short video production methods and systems, this invention is proposed.

[0005] Therefore, the problems that this invention aims to solve are the issues in existing short video production, such as rigid plot logic, simplistic emotional expression, disconnect between copyright and style, lack of true 3D immersive experience, and inability to self-evolve in real time.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0007] In a first aspect, embodiments of the present invention provide a method for producing short videos, which includes a logical deduction engine based on a "multidimensional causal narrative graph" and, based on the emotional nodes generated in the deduction, a cross-modal "neuro-emotional resonance" scheduling system is introduced.

[0008] Under the scheduling of the cross-modal "neuro-emotional resonance" scheduling system, dynamic copyright embedding and personalized generation based on "neural style fingerprint" are performed, and "holographic light field" fusion rendering technology is used.

[0009] Through an edge-adaptive "fluid video" delivery protocol, it adapts to the environment and achieves self-iteration based on a self-evolving creation model of "collective intelligence feedback loop".

[0010] As a preferred embodiment of the short video production method described in this invention, the logical deduction engine based on the "multidimensional causal narrative graph" includes a language model semantic understanding and psychological script theory, used to explore the causal chain hidden behind the story. By constructing a graph containing characters' psychological states, potential conflict climaxes, emotional turning points, and ambiguous subtexts, and transforming the script into a narrative network, it provides a solid logical foundation and data source for subsequent emotional resonance scheduling, stylized rendering, and branch plot triggering.

[0011] As a preferred embodiment of the short video production method described in this invention, the introduction of the cross-modal "neuro-emotional resonance" scheduling system includes, during the logical deduction process based on the "multi-dimensional causal narrative graph", once the engine identifies and generates an emotional node, the device system will activate the cross-modal "neuro-emotional resonance" synchronous scheduling system, convert the emotional weight contained in the node into multi-modal control instructions, and drive the audiovisual rendering engine to adjust the light and shadow tone of the screen, the movement speed of the camera, the frequency fluctuation of the background music, and the particle density of the special effects;

[0012] The cross-modal "neuro-emotional resonance" scheduling system refers to an intelligent hub that integrates a biosignal decoder, a multimodal emotion alignment engine, and a rendering controller.

[0013] As a preferred embodiment of the short video production method described in this invention, the dynamic copyright embedding and personalized generation based on "neural style fingerprint" includes converting emotional vectors into visual and auditory features. The cross-modal "neural emotional resonance" scheduling system extracts the user's interaction habits, device environment, and emotional state as "personalization seeds" and adjusts the stylization parameters of the rendering pipeline to make each frame a unique customized content. At the same time, the cross-modal "neural emotional resonance" scheduling system encodes invisible copyright watermarks into tiny style perturbations that resonate with the emotional rhythm, ensuring that these copyright identifiers not only have robustness against attacks but are also integrated into the personalized artistic style, achieving a unity of content presentation and copyright protection.

[0014] As a preferred embodiment of the short video production method described in this invention, the "holographic light field" fusion rendering technology includes four core modules: three-dimensional scene geometric reconstruction, multi-view light field sampling and encoding, light wave propagation physical simulation, and optical display driving. It uses multi-camera arrays or computational generation technology to collect and encode millions of light data with directional, phase, and amplitude information to construct a four-dimensional light field database. Based on the wave optics principle, it calculates the interference and diffraction effects of light in space to synthesize a continuous parallax image that conforms to the visual characteristics of the human eye. The calculated light field signal is then projected into physical space through hardware terminals such as spatial light modulators or microlens arrays.

[0015] As a preferred embodiment of the short video production method described in this invention, the edge-adaptive "fluid video" delivery protocol includes four core mechanisms: dynamic semantic segmentation encoding, network-aware routing, edge intelligent prefetching caching, and progressive rendering decoding based on device computing power. It utilizes AI to semantically segment video content into different micro-data streams and compress them. By monitoring bandwidth fluctuations, latency jitter, and battery status at the user's end, it adjusts the transmission path and bitrate strategy to ensure lossless delivery of semantic frames. Simultaneously, it combines edge computing nodes to predict the user's viewing trajectory for millisecond-level preloading to eliminate stuttering. Finally, based on the GPU computing power limit of the terminal device, it selects the decoding accuracy and rendering resolution, realizing a shift from "fixed bitrate transmission" to a "flowing on demand, with image quality varying according to the context" delivery mode.

[0016] As a preferred embodiment of the short video production method described in this invention, the self-evolving creation model based on a "collective intelligence feedback loop" includes four core components: multimodal behavior acquisition, distributed emotional value assessment, adversarial strategy gradient update, and knowledge graph reconstruction. The device system captures feedback data such as micro-expressions, dwell time, modification trajectory, and social sharing from global users during the interaction process, constructs a high-dimensional group preference vector, and performs emotional weighting and value quantification on these feedback data through a consensus mechanism to identify creation patterns and style trends. Then, the learning algorithm is used to transform collective intelligence into reward signals, automatically adjusting the potential spatial distribution and parameter weights of the generation model. While preserving the diversity of individual creativity, low-quality noise is eliminated, forming a closed-loop ecosystem of "user feedback driving model iteration and model evolution guiding creation".

[0017] Secondly, embodiments of the present invention provide a short video production system, which includes: an intelligent narrative and emotion scheduling module, which is based on a logic deduction engine of "multi-dimensional causal narrative graph" and introduces a cross-modal "neuro-emotional resonance" scheduling system according to the emotional nodes generated in the deduction.

[0018] The holographic rendering and copyright generation module, under the scheduling of the cross-modal "neuro-emotional resonance" scheduling system, uses dynamic copyright embedding and personalized generation based on "neural style fingerprint" and applies "holographic light field" fusion rendering technology.

[0019] The adaptive delivery and self-evolving ecosystem module adapts to the environment through an edge-adaptive "fluid video" protocol and achieves self-iteration based on a self-evolving creation model of "collective intelligence feedback loop".

[0020] Thirdly, embodiments of the present invention provide a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement any step of the above-described short video production method.

[0021] Fourthly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the above-described short video production method.

[0022] The beneficial effects of this invention are as follows: by constructing a "multi-dimensional causal narrative graph" and a "neuro-emotional resonance" system, it achieves deep intelligent transformation of short videos from logical deduction to emotional audiovisual generation; by utilizing "neural style fingerprints" to implicitly integrate copyright protection into personalized styles, it pioneers a "style is copyright" security mechanism; by combining "holographic light field" rendering technology to break through two-dimensional limitations and provide a naked-eye three-dimensional immersive experience; by relying on the "fluid video" protocol to achieve end-side adaptive and smooth delivery, and by leveraging the "collective intelligence feedback loop" to establish a real-time self-evolving closed loop, it ultimately achieves a comprehensive technological breakthrough in high content resonance, high copyright security, high immersive experience, high transmission reliability, and model self-iteration. Attached Figure Description

[0023] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:

[0024] Figure 1 This is a flowchart of a method and system for producing short videos.

[0025] Figure 2 This is a schematic diagram of a method and system for producing short videos. Detailed Implementation

[0026] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.

[0027] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0028] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0029] This invention is described in detail with reference to the schematic diagrams. When detailing the embodiments of this invention, for ease of explanation, the cross-sectional views illustrating the device structure may be partially enlarged, not adhering to the usual scale. Furthermore, the schematic diagrams are merely examples and should not be construed as limiting the scope of protection of this invention. In actual fabrication, the three-dimensional spatial dimensions of length, width, and depth should be included.

[0030] Furthermore, in the description of this invention, it should be noted that the terms "upper," "lower," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. These terms are used solely for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. In addition, the terms "first," "second," or "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0031] Unless otherwise explicitly specified and limited, the terms "installation," "connection," and "joining" in this invention should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; similarly, they can refer to mechanical connections, electrical connections, or direct connections, or indirect connections through an intermediate medium, or internal connections between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0032] Example 1

[0033] Reference Figure 1 and Figure 2 This is the first embodiment of the present invention, which provides a method for producing short videos, including:

[0034] S1: A logic deduction engine based on a "multidimensional causal narrative graph" and a cross-modal "neuro-emotional resonance" scheduling system based on the emotional nodes generated during the deduction.

[0035] Among them, the logic deduction engine based on the "multidimensional causal narrative graph" integrates language model semantic understanding and psychological script theory to explore the causal chain hidden behind the story. By constructing a graph that includes characters' psychological states, potential conflict climaxes, emotional turning points and ambiguous subtexts, and transforming the script into a narrative network, it provides a solid logical foundation and data source for subsequent emotional resonance scheduling, stylized rendering and branch plot triggering.

[0036] S1.1: The introduction of a cross-modal “neuro-emotional resonance” scheduling system includes the following: In the logical deduction process based on the “multi-dimensional causal narrative graph”, once the engine identifies and generates an emotional node, the device system will activate the cross-modal “neuro-emotional resonance” synchronous scheduling system, which will convert the emotional weight contained in the node into multi-modal control instructions and drive the audiovisual rendering engine to adjust the light and shadow tones of the screen, the speed of camera movement, the frequency fluctuations of the background music, and the particle density of the special effects.

[0037] The cross-modal "neuro-emotional resonance" scheduling system refers to an intelligent hub that integrates a biosignal decoder, a multimodal emotion alignment engine, and a rendering controller.

[0038] S2: Under the scheduling of the cross-modal "neuro-emotional resonance" scheduling system, dynamic copyright embedding and personalized generation based on "neural style fingerprint" are used, and "holographic light field" fusion rendering technology is applied.

[0039] Among them, the dynamic copyright embedding and personalized generation based on "neural style fingerprint" includes transforming emotional vectors into visual and auditory features. The cross-modal "neuro-emotional resonance" scheduling system extracts the user's interaction habits, device environment, and emotional state as "personalization seeds" and adjusts the stylization parameters of the rendering pipeline to make each frame a unique customized content. At the same time, the cross-modal "neuro-emotional resonance" scheduling system encodes invisible copyright watermarks into tiny style perturbations that resonate with the emotional rhythm, ensuring that these copyright marks not only have robustness against attacks but are also integrated into the personalized artistic style, achieving a unity of content presentation and copyright protection.

[0040] S2.1: The “holographic light field” fusion rendering technology includes four core modules: three-dimensional scene geometric reconstruction, multi-view light field sampling and encoding, light wave propagation physical simulation, and optical display driving. It uses multi-camera arrays or computational generation technology to collect and encode millions of light data with directional, phase, and amplitude information to build a four-dimensional light field database. Based on the wave optics principle, it calculates the interference and diffraction effects of light in space to synthesize a continuous parallax image that conforms to the visual characteristics of the human eye. The calculated light field signal is then projected into physical space through hardware terminals such as spatial light modulators or microlens arrays.

[0041] S3: Adapts to the environment through an edge-adaptive "fluid video" delivery protocol and achieves self-iteration based on a self-evolving creation model of "collective intelligence feedback loop".

[0042] The edge-adaptive "fluid video" delivery protocol includes four core mechanisms: dynamic semantic segmentation encoding, network-aware routing, edge intelligent prefetching caching, and progressive rendering decoding based on device computing power. It uses AI to semantically segment video content into different micro-data streams and compress them. By monitoring the bandwidth fluctuations, latency jitter, and battery status of the user's terminal, it adjusts the transmission path and bitrate strategy to ensure lossless delivery of semantic frames. At the same time, it combines edge computing nodes to predict the user's viewing trajectory for millisecond-level preloading to eliminate stuttering. Finally, based on the GPU computing power limit of the terminal device, it selects the decoding accuracy and rendering resolution, realizing the transformation from "fixed bitrate transmission" to "on-demand flow, with image quality changing according to the context" delivery mode.

[0043] S3.1: The self-evolving creation model based on the "collective intelligence feedback loop" includes four core components: multimodal behavior acquisition, distributed emotional value assessment, adversarial strategy gradient update, and knowledge graph reconstruction. The device system captures feedback data such as micro-expressions, dwell time, modification trajectory, and social sharing from global users during the interaction process, constructs a high-dimensional group preference vector, and performs emotional weighting and value quantification on this feedback data through a consensus mechanism to identify creation patterns and style trends. Then, the learning algorithm is used to transform collective intelligence into reward signals, automatically adjusting the potential spatial distribution and parameter weights of the generated model. While preserving the diversity of individual creativity, low-quality noise is eliminated, forming a closed-loop ecosystem of "user feedback driving model iteration and model evolution guiding creation".

[0044] Furthermore, this embodiment also provides a short video production system, including:

[0045] S310: Intelligent narrative and emotion scheduling module, which is based on the logical deduction engine of "multi-dimensional causal narrative graph" and introduces a cross-modal "neuro-emotional resonance" scheduling system according to the emotional nodes generated in the deduction.

[0046] S320: Holographic rendering and copyright generation module. Under the scheduling of the cross-modal "neuro-emotional resonance" scheduling system, it uses dynamic copyright embedding and personalized generation based on "neural style fingerprint" and applies "holographic light field" fusion rendering technology.

[0047] S330: Adaptive Delivery and Self-Evolving Ecosystem Module. It adapts to the environment through the edge-adaptive "fluid video" protocol and achieves self-iteration based on the self-evolving creation model of "collective intelligence feedback loop".

[0048] The above-mentioned unit modules can be embedded in the processor of the computer device in hardware form or independent of it, or they can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of the above modules.

[0049] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0050] In summary, this invention discloses a method and system for producing short videos. By integrating a "multi-dimensional causal narrative graph" with a "neuro-emotional resonance" scheduling system, it achieves intelligent generation from deep logical deduction to cross-modal emotional audiovisual generation. It utilizes "neural style fingerprints" to implicitly embed copyright protection into personalized styles, pioneering a "style as copyright" mechanism. It combines "holographic light field" rendering technology to construct a naked-eye 3D immersive experience. It relies on the edge-adaptive "fluid video" protocol to ensure smooth delivery in complex environments, and establishes a real-time self-evolving closed loop with the help of a "collective intelligence feedback loop." Ultimately, it achieves a comprehensive technological breakthrough in high content resonance, high copyright security, high immersive experience, high transmission reliability, and continuous model iteration.

[0051] Example 2

[0052] Reference Figures 1-2 This is the second embodiment of the present invention, which provides a method for producing short videos. In order to verify the beneficial effects of the present invention, a simulation experiment is conducted for scientific demonstration.

[0053] On the short video platform A, the system first receives the user's input theme "A Lonely Detective in a Cyberpunk Style," and uses the "Multi-Dimensional Causal Narrative Graph" engine to deduce the psychological turning point and emotional climax of the protagonist when he discovers a key clue on a rainy night. Then, the "Neuro-Emotional Resonance" scheduling system transforms this emotional node into specific audiovisual instructions, driving the rendering engine to adjust the screen tone to a strong contrast between cool blue and neon purple, simulate breathing-like tremors in camera movement, and synchronize the background music with low-frequency pulse rhythms. At the same time, the system extracts the current user's device model and preferences as "personalized seeds," generating unique screen textures and implicitly encoding the copyright watermark as particle effects that fluctuate with the music rhythm. Next, the "Holographic Light Field" module calculates light interference based on the principle of wave optics, generating a four-dimensional light field video stream that can present realistic depth of field on a naked-eye 3D screen. Finally, the "Fluid Video" protocol monitors the user's weak network environment in real time, automatically segments the video semantics and reduces the rendering precision of non-critical areas to ensure continuous smooth playback, and after the user finishes watching, converts the duration of their viewing and micro-expression feedback into reward signals, updating the creation model parameters online to optimize the generation effect of subsequent similar themes. The comparison between the present invention and the prior art is shown in Table 1 below:

[0054] Table 1. Comparison of the present invention with the prior art

[0055] Comparison Dimensions Existing technology This invention Narrative Logic and Generation Mechanism It relies on fixed linear script templates or simple keyword matching, lacks in-depth causal reasoning, has a simplistic plot logic, and stiff emotional expression. Based on the "multidimensional causal narrative map" and integrating psychological theories, this paper deeply explores the characters' psychology, the climax of the conflict and the subtext, and constructs a dynamic narrative network to achieve a logically rigorous and multi-branched plot development. Emotional expression and audiovisual interaction Audiovisual elements (images, music, special effects) are usually generated independently or pieced together based on simple rules, and cannot be dynamically adjusted in real time according to the plot and emotions, thus lacking a sense of resonance. The "cross-modal neural emotional resonance" scheduling system is introduced, which transforms the weight of emotional nodes into multimodal control commands for light and shadow, camera movement, sound effects and particle density in real time, so as to achieve deep synchronous resonance between audiovisual and emotional aspects. Copyright protection and personalization Traditional watermarks are easily removed and can damage the aesthetics of the image; the content is mostly generated in a standardized manner, making it difficult to achieve true "personalized" customization. It pioneered the "neural style fingerprint" technology, which encodes copyright watermarks as tiny style perturbations that resonate with emotional rhythm (style is copyright), and combines them with user state seeds to achieve unique customization of each frame. Visual presentation experience Limited to two-dimensional displays, lacking true depth of field, immersive experiences rely on expensive VR / AR headsets. By using "holographic light field" fusion rendering technology, a four-dimensional light field database is constructed and light interference and diffraction are simulated. Through a specific terminal, an immersive experience with naked-eye 3D, continuous parallax, and real depth of field is achieved. Transmission delivery and adaptation Using a fixed bitrate transmission protocol, it is prone to lag when the network fluctuates, and cannot dynamically adjust the image quality according to the terminal's computing power, resulting in a fragmented experience. Employing an edge-adaptive "fluid video" delivery protocol, based on semantic segmentation, network-aware routing, and edge prefetching, it achieves seamless and smooth delivery with "on-demand flow and image quality that changes with the environment". Model evolution mechanism Relying on offline training and manually labeled data, the model has a long iteration cycle, cannot respond to changes in user aesthetics in real time, and lacks closed-loop feedback. By constructing a "collective intelligence feedback loop," we can collect feedback data such as micro-expressions and dwell time from users around the world in real time. Through adversarial strategy gradient updates, we can automatically adjust model parameters and form a self-evolving ecosystem of "feedback-driven iteration."

[0056] Table 1 describes the disruptive advantages of this invention compared to existing technologies: Existing technologies are limited by linear script templates and independently generated audiovisual elements, resulting in simple plot logic, lack of emotional resonance, rigid copyright protection, and visual experience limited to a two-dimensional plane. They also face the dilemma of transmission lag and model iteration delays. In contrast, this invention achieves deep psychological logic deduction through a "multi-dimensional causal narrative graph," achieves real-time dynamic synchronization of audiovisual and emotional responses through "neuro-emotional resonance," pioneers an invisible watermark and personalized customization mechanism of "style as copyright," provides a naked-eye 3D immersive experience with the help of "holographic light field" technology, and achieves adaptive smooth delivery and model self-evolution based on real-time feedback by relying on the "fluid video" protocol. Thus, it constructs a full-link intelligent closed-loop system from content creation, emotional rendering, security rights confirmation to immersive delivery and continuous optimization.

[0057] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for producing short videos, characterized in that: include, Based on the logic deduction engine of "multidimensional causal narrative graph", and based on the emotional nodes generated in the deduction, a cross-modal "neuro-emotional resonance" scheduling system is introduced; Under the scheduling of the cross-modal "neuro-emotional resonance" scheduling system, dynamic copyright embedding and personalized generation based on "neural style fingerprint" are performed, and "holographic light field" fusion rendering technology is used. Through an edge-adaptive "fluid video" delivery protocol, it adapts to the environment and achieves self-iteration based on a self-evolving creation model of "collective intelligence feedback loop".

2. The method for producing short videos as described in claim 1, characterized in that: The logic deduction engine based on the "multidimensional causal narrative graph" integrates language model semantic understanding and psychological script theory to uncover the causal chain hidden behind the story. It constructs a graph that includes characters' psychological states, potential conflict climaxes, emotional turning points, and ambiguous subtexts, and transforms the script into a narrative network, providing a logical foundation and data source for subsequent emotional resonance scheduling, stylized rendering, and branch plot triggering.

3. The method for producing short videos as described in claim 2, characterized in that: The introduction of the cross-modal "neuro-emotional resonance" scheduling system includes the following: In the logical deduction process based on the "multi-dimensional causal narrative graph", once the engine identifies and generates an emotional node, the device system will activate the cross-modal "neuro-emotional resonance" synchronous scheduling system, which will convert the emotional weight contained in the node into multi-modal control instructions and drive the audiovisual rendering engine to adjust the light and shadow tones of the screen, the speed of camera movement, the frequency fluctuations of the background music, and the particle density of the special effects. The cross-modal "neuro-emotional resonance" scheduling system refers to an intelligent hub that integrates a biosignal decoder, a multimodal emotion alignment engine, and a rendering controller.

4. The method for producing short videos as described in claim 3, characterized in that: The dynamic copyright embedding and personalized generation based on "neural style fingerprint" includes converting emotional vectors into visual and auditory features. The cross-modal "neuro-emotional resonance" scheduling system can extract users' interaction habits, device environment, and emotional state as "personalization seeds" and adjust the stylization parameters of the rendering pipeline to make each frame a unique customized content. At the same time, the cross-modal "neuro-emotional resonance" scheduling system encodes the invisible copyright watermark into a tiny style perturbation that is in sync with the emotional rhythm, ensuring that the copyright mark has robustness against attacks and is integrated into the personalized artistic style.

5. The method for producing short videos as described in claim 4, characterized in that: The "holographic light field" fusion rendering technology includes four core modules: three-dimensional scene geometric reconstruction, multi-view light field sampling and encoding, light wave propagation physical simulation, and optical display driving. It uses multi-camera arrays or computational generation technology to collect and encode millions of light data with directional, phase, and amplitude information to build a four-dimensional light field database. Based on the wave optics principle, it calculates the interference and diffraction effects of light in space to synthesize a continuous parallax image that conforms to the visual characteristics of the human eye. The calculated light field signal is then projected into physical space through a hardware terminal such as a spatial light modulator or microlens array.

6. The method for producing short videos as described in claim 5, characterized in that: The edge-adaptive "fluid video" delivery protocol includes four core mechanisms: dynamic semantic segmentation encoding, network-aware routing, edge intelligent prefetching caching, and progressive rendering decoding based on device computing power. It uses AI to semantically segment video content into different micro-data streams and compress them. By monitoring the bandwidth fluctuations, latency jitter, and battery status of the user terminal, it adjusts the transmission path and bitrate strategy to ensure lossless delivery of semantic frames. At the same time, it combines edge computing nodes to predict the user's viewing trajectory for millisecond-level preloading to eliminate stuttering. Finally, based on the GPU computing power limit of the terminal device, it selects the decoding accuracy and rendering resolution, realizing the transformation from "fixed bitrate transmission" to "on-demand flow, image quality changes with the context" delivery mode.

7. The method for producing short videos as described in claim 6, characterized in that: The self-evolving creation model based on the "collective intelligence feedback loop" comprises four core components: multimodal behavior acquisition, distributed emotional value assessment, adversarial strategy gradient update, and knowledge graph reconstruction. The device system captures micro-expressions, dwell time, modification trajectory, and social sharing feedback data of global users during the interaction process, constructs a high-dimensional group preference vector, and performs emotional weighting and value quantification on the above feedback data through a consensus mechanism to identify creation patterns and style trends. Then, the learning algorithm transforms collective intelligence into reward signals, automatically adjusting the potential spatial distribution and parameter weights of the generated model. While preserving the diversity of individual creativity, it eliminates low-quality noise, forming a closed-loop ecosystem of "user feedback driving model iteration and model evolution guiding creation." 8. A short video production system, based on the short video production method according to any one of claims 1 to 7, characterized in that: include, The intelligent narrative and emotion scheduling module is based on the logical deduction engine of "multi-dimensional causal narrative graph" and introduces a cross-modal "neuro-emotional resonance" scheduling system based on the emotional nodes generated in the deduction. The holographic rendering and copyright generation module, under the scheduling of the cross-modal "neuro-emotional resonance" scheduling system, uses dynamic copyright embedding and personalized generation based on "neural style fingerprint" and applies "holographic light field" fusion rendering technology. The adaptive delivery and self-evolving ecosystem module adapts to the environment through the edge-adaptive "fluid video" protocol and achieves self-iteration based on the self-evolving creation model of "collective intelligence feedback loop".

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the short video production method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the short video production method according to any one of claims 1 to 7.