Image-based video generation method and device, equipment and storage medium
By generating multi-scale features and utilizing the differentiable search of the spatiotemporal architecture generator, combined with a temporal recurrent neural network, the optimal spatiotemporal processing path is automatically determined, solving the problem of existing video generation systems relying on manual processing, and achieving efficient and coherent video generation, which is suitable for the financial and medical fields.
Patent Information
- Application Number
- CN202511066121.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-10-14
AI Technical Summary
Existing image-based video generation systems rely on manual processing of time and space, resulting in unstable temporal coherence of the generated videos, low efficiency, and inability to adapt to diverse task requirements.
By receiving the input static image for preprocessing, generating multi-scale features, using the differentiable search of the spatiotemporal architecture generator to determine the optimal spatiotemporal processing path, combining the temporal recurrent neural network to generate video frames, automatically determining the optimal spatiotemporal processing path, eliminating manual intervention, and improving adaptability and efficiency.
It achieves adaptability to diverse needs and improves timing processing efficiency, ensures the temporal continuity and detail authenticity of the generated video, balances generation quality and computational efficiency, and enhances the application value of video generation technology in the financial and medical fields.
Smart Images

Figure CN120786149A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence, financial technology and digital medicine, and in particular to image-based video generation methods, devices, equipment and storage media. Background Art
[0002] Currently, deep learning technology has made significant progress in the field of image-based video generation. The application of generative adversarial networks (GANs) and diffusion models has effectively improved the visual quality and temporal coherence of generated videos. Therefore, it is widely used in various fields, especially in video synthesis, content creation, and virtual reality application scenarios in the financial and medical fields.
[0003] However, existing video generation systems rely heavily on manual intervention for spatiotemporal processing. This typically involves manually presetting spatiotemporal processing module combinations or manually adjusting hierarchical parameters. These rigid, manual processing methods severely restrict the system's flexibility and adaptability. When faced with diverse input content, particularly static scenes and dynamic objects, simple motions and complex actions, such as dynamic market visualization in the financial sector and simulation of the dynamic evolution of lesions in the medical field, these scenarios place stringent demands on video timing accuracy and detailed authenticity, and the requirements for spatiotemporal interaction vary greatly. These rigid, manual processing methods are unable to adapt to diverse task requirements, resulting in unstable temporal coherence in the generated videos. Furthermore, the time-consuming manual processing process leads to low video generation efficiency. Summary of the Invention
[0004] The present invention provides an image-based video generation method, apparatus, device and storage medium to solve the technical problem that existing image-based video generation systems rely on manual processing of time and space problems, resulting in low quality and efficiency.
[0005] In a first aspect, a method for generating an image-based video is provided, comprising:
[0006] Receive the input static image and perform preprocessing to generate multi-scale features;
[0007] Based on the multi-scale features, the optimal spatiotemporal processing path is determined through a differentiable search of the spatiotemporal architecture generator to generate output features that incorporate temporal dynamic information;
[0008] Generate video frames based on the temporal recurrent neural network and the output features;
[0009] The video frame is input into a video frame sequence, and a target video is generated through the video frame sequence.
[0010] In a second aspect, an image-based video generation apparatus is provided, comprising:
[0011] The receiving module is configured to receive an input static image and perform preprocessing to generate multi-scale features;
[0012] The fusion module is configured to determine an optimal spatio-temporal processing path based on the multi-scale features through a differentiable search of a spatio-temporal architecture generator, and generate output features that fuse temporal dynamic information;
[0013] The generation module is configured to generate video frames based on a temporal recurrent neural network and the output features;
[0014] The output module is configured to input the video frames into a video frame sequence, and generate a target video through the video frame sequence.
[0015] In a third aspect, a computer device is provided, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the image-based video generation method when executing the computer program.
[0016] In a fourth aspect, a computer readable storage medium is provided, which stores a computer program, and the computer program implements the steps of the image-based video generation method when executed by a processor.
[0017] The present application has the following beneficial effects compared with the prior art: the present application automatically determines the optimal spatio-temporal processing path through differentiable search, eliminates manual intervention, improves the adaptability to diversified requirements and the efficiency of temporal processing, and guarantees the temporal coherence and detail authenticity of the generated video by dynamically generating output features that fuse temporal information and generating video, thereby balancing the generation quality and computational efficiency and improving the application value of the video generation technology in high-precision scenarios such as video synthesis, content creation, and virtual reality in the fields of finance and medicine.
[0018] The above description is only a summary of the technical solutions of the present application, in order to more clearly understand the technical means of the present application, the application can be implemented according to the content of the specification, and in order to make the above and other purposes, characteristics and advantages of the present application more obvious and easy to understand, the following preferred embodiments are described in detail as follows. BRIEF DESCRIPTION OF DRAWINGS
[0019] Figure 1 is an application environment schematic diagram of the image-based video generation method in an embodiment of the present application;
[0020] Figure 2 is a flowchart of the image-based video generation method in an embodiment of the present application;
[0021] Figure 3 is Figure 2 is a specific implementation flowchart of step S10 in the embodiment;
[0022] Figure 4 is Figure 2 a specific implementation flowchart of step S20 in the method;
[0023] Figure 5 is Figure 4 another specific implementation flowchart of step S23 in the method;
[0024] Figure 6 is Figure 2 a specific implementation flowchart of step S30 in the method;
[0025] Figure 7 is a structural schematic diagram of an image-based video generation apparatus in an embodiment of the present application;
[0026] Figure 8 is a structural schematic diagram of a computer device in an embodiment of the present application;
[0027] Figure 9 is another structural schematic diagram of a computer device in an embodiment of the present application. DETAILED DESCRIPTION
[0028] In order to make the objectives, technical solutions, and advantages of the present application clearer, further detailed descriptions will be given below with reference to the drawings and specific embodiments. The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings of the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of the present application.
[0029] It should be understood that, when used in the specification and the appended claims, the terms “comprise” and “include” indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not exclude one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0030] It should also be understood that the terms used in the present application specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the present application specification and the appended claims, the singular forms “a”, “an” and “the” are intended to include the plural forms, unless the context clearly indicates otherwise.
[0031] It should be further understood that the term “and / or” used in the present application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes these combinations.
[0032] Referring to Figure 1 and Figure 2 , Figure 1 The application scenario of the image-based video generation method provided by the embodiment of the present application is shown in the figure. The client communicates with the server through the network. The user can upload a static image through the client, and the client sends the static image to the server through the network; the server receives the input static image and performs preprocessing to generate multi-scale features; based on the multi-scale features, the optimal spatiotemporal processing path is determined through the differentiable search of the spatiotemporal architecture generator to generate output features that fuse temporal dynamic information; video frames are generated based on the temporal recurrent neural network and the output features; the video frames are input into a video frame sequence, and the target video is generated through the video frame sequence. The client can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers and portable wearable devices; the server can be implemented by an independent server or a server cluster composed of multiple servers. The present application will be described in detail through specific embodiments.
[0033] Referring to Figure 2 as shown, Figure 2 The schematic flowchart of the image-based video generation method provided by the embodiment of the present application is shown in the figure. The image-based video generation method comprises the following steps:
[0034] S10: receiving an input static image and performing preprocessing to generate multi-scale features.
[0035] In this embodiment, the multi-scale features refer to image features extracted at different spatial resolutions and abstraction levels, such as edges, textures, semantic information, etc., which ensure that the video generation model can capture multi-level information from local details to global content. The static image is the original input data of video generation and is the basis for spatiotemporal expansion of the model, and all subsequent temporal dynamic information is generated based on the content of the static image. Step S10 unifies the input data format through preprocessing and extracts multi-scale features, solves the problem of inconsistent original static image data formats, and provides a diversified feature basis for subsequent spatiotemporal processing, improves the adaptability of the model to different scenes, and lays a data foundation for generating high-quality videos.
[0036] It can be understood that the static image can be selected according to actual needs, and the static image includes but is not limited to historical K-line chart, fund flow atlas in the financial field, medical static image, pathological section, operation planning chart and the like in the medical field. In specific implementation, the user can upload the static image through the client, and the static image is processed by the server to generate a video based on the static image. For example, the user inputs a 30-day K-line change screenshot of a stock K-line chart through the client, and generates a stock trend video through the image-based video generation method of the embodiment. The user inputs several medical images such as CT section static images through the client, and generates a lesion change simulation video through the image-based video generation method of the embodiment.
[0037] In some embodiments of the application, as shown in Figure 3 A specific multi-scale feature generation scheme is provided, S10, that is, receiving an input static image and preprocessing to generate multi-scale features, specifically including the following steps S11-S13.
[0038] S11: receiving an input static image.
[0039] S12: performing standardization processing on the static image to obtain standardized image data.
[0040] S13: performing feature extraction at different abstraction levels and spatial resolutions from the standardized image data to obtain multi-scale features.
[0041] The standardization processing refers to scaling and normalizing operations on the pixel values of the static image, which is used to eliminate the brightness and contrast differences between different static images, make the input data distribution consistent, and facilitate stable model training and feature extraction. The abstraction level is the progressive relationship of image features from low-level such as edge, color, to high-level such as semantic category; the spatial resolution refers to the size of the feature map, such as 128x128, 64x64, and the like. In the embodiment, by extracting multi-scale features, the multi-abstraction level and spatial resolution information of the image is captured, which not only retains local details such as texture and edge, but also includes global semantics such as object category and scene structure, providing comprehensive feature support for subsequent spatio-temporal architecture search, and improving the detail authenticity and semantic consistency of the generated video.
[0042] S20: determining an optimal spatio-temporal processing path based on the multi-scale features through a differentiable search of a spatio-temporal architecture generator to generate output features fused with time-series dynamic information.
[0043] The spatio-temporal architecture generator (STAG) is a hardware-aware system that dynamically generates an optimal video processing path through differentiable neural architecture search (DNAS). It can construct and optimize a spatio-temporal hybrid computing architecture in real time according to the content complexity and motion requirements of the input static image, significantly improving the computing efficiency while ensuring the generated quality. In this embodiment, the spatio-temporal architecture generator is responsible for automatically selecting appropriate spatio-temporal processing modules and computing resource allocation strategies based on the input multi-scale features, achieving dynamic optimization of the architecture and solving the problem of insufficient flexibility of traditional fixed architectures. Differentiable search is a neural architecture search method that optimizes architecture parameters through gradient descent. In this embodiment, it is used to automatically find the optimal spatio-temporal processing path within a predefined search space, without the need for manual architecture design, improving the efficiency and accuracy of architecture search. The optimal spatio-temporal processing path refers to the combination of spatio-temporal modules that balances the generated quality and computing efficiency found during the search process. In this embodiment, it serves as the specific process of processing feature map fusion of spatio-temporal information, ensuring the rationality and efficiency of feature fusion. The temporal dynamic information refers to the dynamic content-related information such as motion and state that changes over time in a video. In this embodiment, the temporal dynamic information is generated by expanding the multi-scale features of the static image, providing key information for the temporal coherence of the video.
[0044] For step S20, the optimal spatio-temporal processing path is determined based on the multi-scale features through the differentiable search of the spatio-temporal architecture generator, and the output features that fuse the temporal dynamic information are generated. The content and motion complexity of the static image are reflected through the multi-scale features, and the optimal spatio-temporal processing path is automatically selected by the spatio-temporal architecture generator, enabling the model to adaptively process spatio-temporal information according to the input content. This ensures the temporal coherence and semantic accuracy of the generated video, optimizes the computing efficiency, avoids redundant computation, and achieves a balance between generation quality and computing cost.
[0045] In specific implementations, such as in the application scenario of lesion change simulation video generation in the medical field, the multi-scale features demonstrate that the complexity of the static image is high, for example, the texture of the junction between cancerous tissue and healthy tissue in a CT slice is mixed, and blood vessels are densely distributed. The spatio-temporal architecture generator enables multi-layer axial attention to model the cancer cell migration path along the blood vessel direction, uses multi-channel number enhancement to capture the details of microvascular permeation, and performs multi-scale fusion to integrate features at the cell level with high resolution and the organ level with low resolution, and then completes resource allocation and spatio-temporal processing of complex static images.
[0046] In a specific implementation, for example, in a stock trend prediction video generation application scenario in the financial field, the multi-scale feature shows that the complexity of the static image is low, for example, the lines of the K-line chart are obvious and in narrow box oscillation. The spatio-temporal architecture generator uses factorized 3D convolution to decompose the spatio-temporal calculation, suppresses overfitting to the slight fluctuations of the lines by using a low number of channels, and maintains the basic trend memory by using a lightweight time series recurrent neural network variant, such as the 20-day moving average direction, and then completes the resource allocation and spatio-temporal processing of the simple static image.
[0047] It can be understood that the core component of the spatio-temporal architecture generator is the spatio-temporal module, which is the smallest hardware-aware unit that performs basic spatio-temporal calculation and fuses spatial and temporal information through specific mathematical operations. The selection and combination of the spatio-temporal module directly reflects the depth of the system's understanding of the static image and the intelligent scheduling ability of the computing resources. The essence is to adapt to the spatio-temporal modeling needs of different scenarios through dynamic architecture, and in this embodiment, the selection and combination of the spatio-temporal module are determined by the optimal spatio-temporal processing path, thereby realizing the dynamic adjustment of the computing resources.
[0048] More specifically, the working principle of the spatio-temporal architecture generator is based on the following mathematical framework:
[0049]
[0050] where α represents the architecture parameter, which defines the selection and combination of the spatio-temporal module in the spatio-temporal architecture generator, θ represents the model weight parameter, and are the loss functions on the validation set and the training set, respectively, FLOPs(α) represents the computational complexity of architecture α, and λ is a weighting coefficient for balancing the generation quality and the computational efficiency. In a specific implementation, in a predefined search space , the optimal architecture parameter α is found, such that the weighted sum of the validation set loss and the computational complexity FLOPs is minimized, thereby achieving optimal resources.
[0051] In some embodiments of the present application, as shown in Figure 4 , a specific output feature generation scheme is provided, in S20, that is, based on the multi-scale feature, the optimal spatio-temporal processing path is determined by the differentiable search of the spatio-temporal architecture generator to generate output features that fuse time series dynamic information, specifically including the following steps S21-S24.
[0052] S21: using the spatio-temporal architecture generator to encode the multi-scale feature to obtain a feature map of multi-level abstract information from global to local.
[0053] In this embodiment, feature encoding is a process of converting input multi-scale features into a feature representation more suitable for subsequent processing, more specifically, for converting multi-scale features into a feature map containing multi-level abstract information from global semantics to local details, making the features more conducive to spatio-temporal information fusion and processing. Multi-level abstract information refers to information of different levels of abstraction from low-level local details such as edges, textures, to high-level global semantics such as object categories, scene structures, so as to ensure that subsequent spatio-temporal processing can take into account both details and overall semantics. The latent feature representation of the encoded multi-scale features after feature map encoding contains the matrix or tensor of the multi-scale features, which carries various types of information encoded from the multi-scale features.
[0054] S22: Calculate the information entropy of the feature map layer by layer.
[0055] It can be understood that information entropy is a measure of information uncertainty or complexity, and the greater the value, the more complex and uncertain the information, which is used to quantify the complexity of information in the feature map, providing a basis for subsequent channel adjustment and architecture search, so that the model can adaptively allocate computing resources according to the complexity of the features.
[0056] For step S22, the information entropy is calculated layer by layer according to different levels of the feature map, such as different resolutions or abstraction levels, which can more finely capture the differences in information complexity of different levels of features, facilitating the model to allocate more computing resources to feature maps with high information complexity and reduce computing resources for feature maps with simple information, thereby improving computing efficiency while ensuring processing effect.
[0057] S23: Using a double-layer optimization framework, perform differentiable search based on the information entropy in a predefined search space to obtain an optimal spatio-temporal processing path.
[0058] The double-layer optimization framework includes a nested optimization structure of upper-layer optimization such as architecture parameter optimization and lower-layer optimization such as model weight parameter optimization, which is used to optimize the architecture parameters and model weights simultaneously, realize joint optimization of generation quality and computing efficiency, and find an optimal architecture that takes both into account. In this embodiment, the architecture search is limited to a range, i.e., a predefined search space, which contains a set of various possible spatio-temporal aggregation operators (such as factorized 3D CNN convolution, axial attention mechanism), channel configurations, etc., making the search process more targeted and efficient, and avoiding meaningless search.
[0059] In some embodiments of the present application, as shown in Figure 5 An optimal spatio-temporal processing path generation scheme is provided, S23, i.e., using a double-layer optimization framework, performing differentiable search based on the information entropy in a predefined search space to obtain an optimal spatio-temporal processing path, which specifically includes the following steps S231-S236.
[0060] S231: predicting a channel adjustment amount through a learnable mapping function according to the information entropy.
[0061] By mapping the information entropy to the channel adjustment amount, the learnable mapping function can adaptively adjust the mapping relationship according to the actual data distribution, thereby improving the accuracy of the channel adjustment amount prediction. The learnable mapping function is a function that can continuously optimize parameters through training, has a learnable characteristic, and is beneficial to dynamically adjusting the channel number. The channel adjustment amount refers to a numerical value for adjusting the number of convolution layer channels in the spatio-temporal aggregation process, which is used to dynamically change the channel number so that the channel number can be adapted according to the information entropy of the feature map. When the information entropy is high, the channel is increased for more detailed processing. When the information entropy is low, the channel is reduced to save resources, avoid waste of computing resources caused by too many channels, or insufficient feature processing caused by too few channels, and improve the resource utilization efficiency and feature processing capability of the model.
[0062] S232: determining a target channel number based on the channel adjustment amount and a preset basic channel number.
[0063] Based on the channel adjustment amount and the preset basic channel number, the target channel number is determined, which is beneficial to converting the predicted channel adjustment amount into an actual usable channel number, retaining the basic processing capability ensured by the basic channel number, and realizing flexible adaptation of the channel number through the adjustment amount, thereby ensuring the sufficiency of feature processing and the efficiency of calculation. The preset basic channel number is the initial channel number of the convolution layer, which is set in advance as a reference value for dynamic adjustment of the channel number, thereby ensuring the stability and rationality of the channel adjustment. The target channel number is the final channel number of the convolution layer after the channel adjustment, which is the channel number of the convolution layer during actual work in the case, directly determines the computational complexity and feature processing capability of the layer, and is a key parameter for balancing the computational efficiency and processing effect.
[0064] S233: obtaining an architecture parameter through differentiable search in a double-layer optimization framework according to the calculated information entropy.
[0065] The architecture parameter is obtained through differentiable search in a double-layer optimization framework according to the information entropy, which is used to determine the specific composition of the network architecture of the spatio-temporal architecture generator, so that the obtained architecture parameter can adapt to the information complexity of the feature map. It can be understood that the architecture parameter is a parameter that defines the network architecture structure, including the selection of the spatio-temporal module, the connection mode, the channel configuration, etc. In the embodiment, the architecture parameter determines the specific composition of the optimal spatio-temporal processing path, and directly affects the performance and computational efficiency of the network.
[0066] S234: constructing a set of spatio-temporal aggregation operators based on the predefined search space, the set of spatio-temporal aggregation operators being configured to generate a set of candidate operators with channel constraints according to the target number of channels.
[0067] The step S234 generates the set of candidate operators that meet the computing resource constraints, provides specific operator selection for subsequent architecture parameter instantiation, and ensures that the candidate operators both cover various spatio-temporal processing capabilities and meet the resource limitations of dynamic channel allocation. The set of spatio-temporal aggregation operators is a set composed of various operators for fusing spatial and temporal information, and is the basis for constructing candidate operators in this embodiment. Each operator can implement different spatio-temporal information processing functions. The set of candidate operators with channel constraints is a set of operators obtained by configuring the set of spatio-temporal aggregation operators according to the target number of channels. In this case, the number of channels of each operator is limited to meet the requirements of dynamic resource allocation, and provides candidates that meet the resource constraints for subsequent selection of the optimal operator combination.
[0068] S235: optimizing the architecture parameters through a loss function, and selecting an architecture parameter with the minimum loss from the optimized architecture parameters.
[0069] In this embodiment, the architecture parameters are continuously optimized through the guidance of the loss function, and the architecture parameter with the optimal performance is selected to ensure that the corresponding network architecture can generate high-quality and time-sequential videos. The loss function is a function for measuring the difference between the predicted results of the model and the true results. In this embodiment, the loss function includes adversarial loss, CLIP semantic alignment loss, and temporal consistency loss, which are used to quantify the network performance corresponding to the architecture parameters and provide a direction for architecture parameter optimization. The architecture parameter with the minimum loss refers to the architecture parameter that minimizes the value of the loss function after optimization. It can be understood that the architecture parameter with the minimum loss is the optimal architecture parameter, which corresponds to the network architecture with the optimal performance, can maximize the quality and time-sequential consistency of the generated video, and also takes into account the computing efficiency.
[0070] S236: instantiating the set of candidate operators using the architecture parameter with the minimum loss to obtain an optimal spatio-temporal processing path.
[0071] The instantiation refers to the process of specificizing the abstract architecture parameters into executable operator combinations and connection methods. In this embodiment, the instantiation operation causes the set of candidate operators to form a specific processing flow according to the definition of the optimal architecture parameter, and converts the theoretical architecture parameters into an actual executable processing path.
[0072] It can be understood that the space-time module is a specific functional carrier for the space-time information processing of the STAG, which integrates multiple space-time aggregation operators inside, and completes the basic space-time calculation through a specific combination method. The space-time module is the smallest hardware-aware unit for performing basic space-time calculation, that is, the smallest hardware-aware unit of the space-time architecture generator. The space-time module integrates multiple space-time aggregation operators inside, and the function of the space-time module is determined by the space-time aggregation operators contained therein. For example, a space-time module focusing on capturing short-term dynamics may mainly use axial attention mechanism, while a module focusing on long-term dependency may mainly use recurrent skip connection. The space-time aggregation operator includes factorized 3D CNN convolution, axial attention mechanism, recurrent skip connection and the like, and is a specific mathematical operation for realizing space-time information fusion. The core function of the space-time module is to fuse space and time information, and this process is mainly realized through convolution operation. The factorized 3D CNN convolution, that is, the 3D convolution is decomposed into time convolution and space convolution, respectively capturing the features of time dimension and space dimension, is the basic operation for realizing space-time information fusion of the space-time module. The dynamic channel allocation mechanism of the embodiment directly acts on the convolution layer, adjusts the number of channels of the convolution layer, matches the calculation resources of the space-time module with the complexity of the input features, and ensures that the space-time module retains key information while processing efficiently.
[0073] The following are specific implementation examples of steps S21-S23:
[0074] Predefined search space under the space-time architecture generator The space-time aggregation operator includes:
[0075] The factorized 3D CNN convolution has the expression: Wherein, W t is a time convolution kernel, t is a time, W h and W w are space convolution kernels, h is height, w is width, represents a tensor product operation.
[0076] The axial attention mechanism can calculate attention along the time, height and width dimensions respectively.
[0077] The recurrent skip connection realizes long-term dependency modeling of time sequence information through the time sequence recurrent neural network variant LSTM.
[0078] In specific implementation, the features of the feature map at different spatial resolutions are weighted by the axial attention mechanism, and then a spatial resolution LSTM is used to obtain multi-scale features f i from small to large, where i represents the feature of the i-th scale.
[0079] When performing dynamic channel allocation, the number of channels c of each 3DCNN convolutional layer l layer (t*h*w*c) is dynamically adjusted according to the input complexity, and the adjustment basis is:
[0080]
[0081] Wherein, w l is the channel adjustment amount of the lth layer, round is the rounding operation, is the preset basic channel number, Δw l is the maximum channel adjustment amount, artificially set, z l is the input feature of the lth layer, and Entropy(·) is the information entropy.
[0082] S24: The optimal space-time processing path is used to fuse the space-time information of the feature map, and the output feature fused with the time sequence dynamic information is generated.
[0083] For step S24, the static feature map is converted into a feature containing dynamic time information, realizing the key conversion from image feature to video feature, so that the output feature has spatial details and time sequence dynamics, providing a high-quality feature basis for subsequent generation of coherent and real video frames. Among them, the fusion of space-time information refers to the process of integrating the spatial information such as the position and shape of the object in the feature map, and the time information such as the motion trend and state change of the object. The output feature fused with the time sequence dynamic information is the feature after the fusion of space-time information, which not only retains the spatial details of the original multi-scale feature, but also contains dynamic information changing with time, and is the direct input of generating video frames, which determines the quality and time sequence consistency of the video frames.
[0084] For example, when generating a fund net value change video, the space-time information of the fund net value feature map is fused through the optimal space-time processing path, such as the spatial distribution of net value in different time periods and the rising and falling dynamics with time, and the generated output feature contains the time sequence change rule of the net value, and the video generated accordingly can clearly show the fluctuation trend of the fund net value, helping investors understand the fund performance. When generating a simulation video of cell division process, the space-time information of the cell feature map is fused using the optimal space-time processing path, such as the spatial form of the cell and the time change in the division process, and the output feature contains the time sequence dynamics of cell division, and the generated video can intuitively present the whole process of cell division, providing intelligent, convenient and efficient assistance for biological research and medical education.
[0085] In some embodiments of the present application, S24, i.e., the optimal space-time processing path is used to fuse the space-time information of the feature map, and the output feature fused with the time sequence dynamic information is generated, specifically including the following steps:
[0086] Extracting hierarchical spatiotemporal information from the feature map according to the optimal spatiotemporal processing path;
[0087] Through the gated fusion mechanism, the spatiotemporal information is weightedly fused according to scale to generate output features that incorporate temporal dynamic information.
[0088] Hierarchical spatiotemporal information extraction refers to extracting spatiotemporal information separately according to different levels of feature maps, such as different resolutions or levels of abstraction. In this embodiment, it can more meticulously capture the spatiotemporal characteristics of features at each level, avoid information omission, and improve the comprehensiveness of feature extraction. The gated fusion mechanism is a mechanism that uses learnable gating parameters to perform weighted fusion of spatiotemporal information of different scales. In this embodiment, the sigmoid function is used to activate the gating parameters, controlling the fusion weights of features at different scales, giving more weight to important features and improving the effectiveness of the fused features.
[0089] In specific implementation, the output features are fused using a gated fusion mechanism. More specifically, the following expression is used: t =∑ i σ(g i )·f i ;
[0090] Among them, f i is the feature of the i-th scale of the feature map; g i is the learnable gating parameter in the spatial resolution LSTM; σ is the sigmoid function. Each scale f of the current frame t i All need to be convolved through the convolution layer of layer l, and finally output the potential state vector, that is, the output feature z t The potential state vector is used as the input of the time series LSTM to obtain the hidden layer state vector h t , in order to obtain the potential and more comprehensive information of the feature map.
[0091] S30: Generate video frames based on the temporal recurrent neural network and the output features.
[0092] The time sequence recurrent neural network (LSTM) is a neural network capable of processing sequence data and having a memory function, and can predict a current state by using historical information. In the embodiment, the time sequence recurrent neural network is used to generate continuous video frames based on the output features and historical memory states, to ensure the time sequence continuity between the video frames, and is a core model for realizing conversion from the features to the video frames. The video frames are generated based on the time sequence recurrent neural network and the output features, to convert the abstract output features into specific visual frame pictures. By means of the memory function of the time sequence recurrent neural network, the generated video frames can maintain the continuity in time sequence, and the content of the previous frame and the next frame naturally connects, and the dynamic information in the output features ensures that the video frames can accurately reflect the dynamic expansion of the input images, thereby laying a foundation for generating complete videos.
[0093] In some embodiments of the present application, as shown in Figure 6 , a specific video frame generation scheme is provided. In S30, the video frames are generated based on the time sequence recurrent neural network and the output features, and specifically include the following steps S31-S33.
[0094] S31: It is judged whether there is a video frame before the current time.
[0095] S32: If there is no video frame before the current time, the output features and the memory state of the output features are input into the time sequence recurrent neural network to generate the video frame at the current time.
[0096] S33: If there is a video frame before the current time, the previous time video frame and the memory state of the previous time video frame are obtained, and the previous time video frame and the memory state of the previous time video frame are input into the time sequence recurrent neural network to generate the video frame at the current time.
[0097] The memory state is a state vector used for storing the output feature information in the time sequence recurrent neural network, and can accurately reflect the core content in the output features. It can be understood that the memory state is equivalent to the understanding and memory of the output features, and it encodes the rules and information contained in the video frames so far, such as the trend and update frequency of the stock trend chart, the moving speed and swing angle of the medical robot, etc. The previous and subsequent video frames are connected by time, action sequence context and trend.
[0098] The following is a specific implementation example of the video frame generation:
[0099] Given the output features z t and the memory state h t of the previous time t, the generation process of the current frame t+1 is represented as: v t+1 =STAG α *(z t ,h t ;θ* );
[0100] The update of the memory state is implemented by a time sequence LSTM: h t+1 = LSTM(v t+1 , h t ); the final output of each frame of the video frame v t is the picture of the target video.
[0101] S40: input the video frame into a video frame sequence, and generate a target video through the video frame sequence.
[0102] For step S40, the scattered and time sequence generated video frames are concatenated into a coherent whole, integrated into a complete video, and the originally isolated frame pictures form a dynamic video with time sequence logic, which fully presents the dynamic process expanded from static images, and realizes the final conversion from images to videos.
[0103] In some embodiments of the present application, after S40, i.e., inputting the video frame into a video frame sequence and generating a target video through the video frame sequence, the following steps are further included:
[0104] Discriminative optimization and post-processing are performed on the video sequence to output a final video.
[0105] Among them, discriminative optimization is a process of using a discriminator such as a PatchGAN discriminator to evaluate the quality of the video sequence and optimizing the video frame according to the evaluation results, which is used to ensure that the generated video is more visually close to the real video. In this embodiment, post-processing includes time sequence super-resolution and adaptive filtering operation, which is used to optimize the details of the video. Specifically, it improves the frame rate to make the video smoother and reduces noise to make the picture clearer. Therefore, post-processing is an important link to improve the final quality of the video. The final video after discriminative optimization and post-processing has higher visual quality, time sequence coherence and clarity, and can better adapt to the demand for high-precision video in the fields of finance, medical treatment and the like.
[0106] In practice, the entire system is optimized using a composite loss function, consisting of an adversarial loss, a CLIP semantic alignment loss, and a temporal consistency loss. The CLIP semantic alignment loss uses the CLIP model to calculate the embedding similarity between the generated video and the text description. This loss ensures that the semantic content of the generated video is consistent with the input text or underlying semantic requirements, avoiding semantic bias. The CLIP (Contrastive Language-Image Pre-training) model is a multimodal pre-trained neural network designed to establish associations between images and natural language. Its core innovation lies in constructing a unified feature space within which images and text can be directly compared and matched, enabling cross-modal understanding and translation. The temporal consistency loss calculates the differences between adjacent frames after optical flow warping to ensure motion coherence between video frames and reduce temporal jumps. The CLIP semantic alignment loss, temporal consistency loss, and adversarial loss form a composite loss function, jointly optimizing the semantic accuracy, temporal coherence, and visual realism of the generated video, balancing multiple performance metrics.
[0107] Among them, the adversarial loss function is:
[0108]
[0109] Among them, the CLIP semantic alignment loss function is:
[0110]
[0111] Among them, the temporal consistency loss function is:
[0112]
[0113] in It is the warping operation of optical flow estimation.
[0114] The following is a specific implementation example of the image-based video generation method of this embodiment:
[0115] The generated video of this implementation example aims to convert 10 continuous medical rehabilitation therapist walking images into a smooth rehabilitation video. For this purpose, a sequence of images including 10 static images showing the process of a person from standing to walking, and a background containing dynamic elements such as trees, is uploaded by the client. The server pre-processes the sequence of images in turn and inputs them into the spatio-temporal architecture generator. The spatio-temporal architecture generator automatically selects a lightweight processing path for the first static image and outputs an initial video frame. For the second to ninth static images, a more complex spatio-temporal modeling is enabled, and for the tenth static image, a lightweight processing path is adopted. At the same time, sparse attention mechanism is used for the background area unrelated to the therapist and his actions, such as trees, to save computation, and video frames are output in turn. The details of the dynamic area, i.e. the therapist, are preserved to a higher degree, and the background processing efficiency is improved, making the overall computational load significantly lower than that of static fixed architecture.
[0116] As can be seen, in the above scheme, the optimal spatio-temporal processing path is automatically determined by differentiable search, eliminating the need for human intervention and improving the adaptability to diverse needs and the efficiency of time series processing. By dynamically generating output features that integrate time series information and generating videos, the temporal coherence and detail authenticity of the generated videos are ensured, balancing the generation quality and computational efficiency, and improving the application value of video generation technology in high-precision scenarios such as video synthesis, content creation, and virtual reality in the financial and medical fields.
[0117] It should be understood that the size of the serial number of each step in the above embodiments does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0118] In an embodiment, the present application provides an image-based video generation device 100, which corresponds to the image-based video generation method described above. As shown in Figure 7 The image-based video generation device 100 includes a receiving module 101, a fusion module 102, a generation module 103, and an output module 104. The functions of each module are described in detail as follows:
[0119] The receiving module 101 is used to receive and pre-process the input static images to generate multi-scale features.
[0120] The fusion module 102 is used to determine the optimal spatio-temporal processing path through the differentiable search of the spatio-temporal architecture generator based on the multi-scale features, and generate output features that integrate time series dynamic information.
[0121] The generation module 103 is used to generate video frames based on the time series recurrent neural network and the output features.
[0122] The output module 104 is configured to input the video frame into a video frame sequence, and generate a target video through the video frame sequence.
[0123] In an embodiment, the receiving module 101 is specifically configured to:
[0124] receive an inputted static image;
[0125] perform standardization processing on the static image to obtain standardized image data;
[0126] perform feature extraction on the standardized image data at different abstraction levels and spatial resolutions to obtain multi-scale features.
[0127] In an embodiment, the fusion module 102 is specifically configured to:
[0128] perform feature encoding on the multi-scale features by using a space-time architecture generator to obtain a feature map of multi-level abstract information from global to local;
[0129] calculate information entropy of the feature map layer by layer in sequence;
[0130] perform differentiable search based on the information entropy in a predefined search space by using a double-layer optimization framework to obtain an optimal space-time processing path;
[0131] fuse space-time information of the feature map by using the optimal space-time processing path to generate an output feature fused with time-series dynamic information.
[0132] In an embodiment, the performing differentiable search based on the information entropy in a predefined search space by using a double-layer optimization framework to obtain an optimal space-time processing path comprises:
[0133] predicting a channel adjustment amount by using a learnable mapping function according to the information entropy;
[0134] determining a target channel number based on the channel adjustment amount and a preset basic channel number;
[0135] obtaining an architecture parameter by differentiable search in a double-layer optimization framework according to the calculated information entropy;
[0136] constructing a set of space-time aggregation operators based on the predefined search space, and generating a candidate operator set with channel constraints according to the target channel number;
[0137] optimizing the architecture parameter by using a loss function, and selecting an architecture parameter with minimum loss from the optimized architecture parameter;
[0138] instantiating the candidate operator set by using the architecture parameter with minimum loss to obtain the optimal space-time processing path.
[0139] In an embodiment, the adopting the optimal space-time processing path to fuse the spatial and temporal information of the feature maps to generate the output feature fused with the temporal dynamic information comprises:
[0140] According to the optimal space-time processing path, the hierarchical spatial and temporal information of the feature maps is extracted;
[0141] The spatial and temporal information is weighted and fused by a gating fusion mechanism according to scales to generate the output feature fused with the temporal dynamic information.
[0142] In an embodiment, the generating module 103 is specifically configured to:
[0143] determine whether there is a video frame before the current moment;
[0144] If there is no video frame before the current moment, the output feature and the memory state of the output feature are input into a time recurrent neural network to generate a video frame at the current moment;
[0145] If there is a video frame before the current moment, a previous moment video frame and a memory state of the previous moment video frame are obtained, and the previous moment video frame and the memory state of the previous moment video frame are input into a time recurrent neural network to generate a video frame at the current moment.
[0146] The specific limitations of the image-based video generation apparatus 100 can refer to the limitations of the image-based video generation method described above, and will not be repeated here. Each module in the above image-based video generation apparatus 100 can be realized by software, hardware and a combination thereof in whole or in part. The above each module can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so as to call and execute the operation corresponding to each module by the processor.
[0147] In one embodiment, a computer device 200 is provided, which can be a server, and its internal structure diagram can be as shown in Figure 8As shown in the figure. The computer device 200 includes a processor 220, a memory and a network interface 250 connected through a system bus 210. Among them, the processor 220 of the computer device is used to provide computing and control capabilities. The memory of the computer device 200 includes a non-volatile and / or volatile storage medium, an internal memory 240. The non-volatile storage medium 230 stores an operating system 231, a computer program 232 and a database 233. The internal memory 240 provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium 230. The network interface 250 of the computer device 200 is used to communicate with the external client through the network connection. The computer program is executed by the processor 220 to realize the function or step of the image-based video generation method server. That is, the processor 220 realizes the following steps when executing the computer program:
[0148] Receiving the input static image and preprocessing to generate multi-scale features;
[0149] Based on the multi-scale features, determining the optimal spatio-temporal processing path through the differentiable search of the spatio-temporal architecture generator, and generating output features that fuse temporal dynamic information;
[0150] Generating video frames based on the temporal recurrent neural network and the output features;
[0151] Inputting the video frames into a video frame sequence, and generating a target video through the video frame sequence.
[0152] In one embodiment, a computer device 300 is provided, which can be a client, and its internal structure diagram can be as shown in the figure. Figure 9 The computer device includes a processor 320, a memory, a network interface 350, a display screen 370 and an input device 360 connected through a system bus 310. Among them, the processor 320 of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium 330 and an internal memory 340. The non-volatile storage medium 330 stores an operating system 331 and a computer program 332. The internal memory provides an environment for the operation of the operating system 331 and the computer program 332 in the non-volatile storage medium 330. The network interface 350 of the computer device 300 is used to communicate with the external server through the network connection. The computer program is executed by the processor 320 to realize the function or step of the image-based video generation method client side. That is, the processor 320 realizes the following steps when executing the computer program 332:
[0153] Receiving the input static image and preprocessing to generate multi-scale features;
[0154] Based on the multi-scale features, the optimal spatiotemporal processing path is determined through a differentiable search of the spatiotemporal architecture generator to generate output features that incorporate temporal dynamic information;
[0155] Generate video frames based on the temporal recurrent neural network and the output features;
[0156] The video frame is input into a video frame sequence, and a target video is generated through the video frame sequence.
[0157] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0158] Receive the input static image and perform preprocessing to generate multi-scale features;
[0159] Based on the multi-scale features, the optimal spatiotemporal processing path is determined through a differentiable search of the spatiotemporal architecture generator to generate output features that incorporate temporal dynamic information;
[0160] Generate video frames based on the temporal recurrent neural network and the output features;
[0161] The video frame is input into a video frame sequence, and a target video is generated through the video frame sequence.
[0162] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can refer to the relevant description in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0163] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, storage, database or other medium used in the embodiments provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0164] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of functional units and modules is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the above-described functions.
[0165] The above-mentioned embodiments are only used to illustrate the technical solutions of the present application, but not limit it. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features. Such modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
Claims
1. An image-based video generation method, characterized in that: include: Receive the input static image and perform preprocessing to generate multi-scale features; Based on the multi-scale features, the optimal spatiotemporal processing path is determined through a differentiable search of the spatiotemporal architecture generator to generate output features that incorporate temporal dynamic information; Generate video frames based on the temporal recurrent neural network and the output features; The video frame is input into a video frame sequence, and a target video is generated through the video frame sequence.
2. The image-based video generation method according to claim 1, characterized in that The receiving of the input static image and preprocessing to generate multi-scale features includes: Receive an input static image; performing standardization processing on the static image to obtain standardized image data; Feature extraction at different abstraction levels and spatial resolutions is performed from the standardized image data to obtain multi-scale features.
3. The image-based video generation method according to claim 2, characterized in that The method of determining the optimal spatiotemporal processing path based on the multi-scale features through a differentiable search of the spatiotemporal architecture generator to generate output features that incorporate temporal dynamic information includes: Using a spatiotemporal architecture generator to perform feature encoding on the multi-scale features, thereby obtaining a feature map of multi-level abstract information from global to local levels; Calculating the information entropy of the feature map layer by layer; A two-layer optimization framework is used to perform a differentiable search based on the information entropy in a predefined search space to obtain the optimal spatiotemporal processing path; The optimal spatiotemporal processing path is used to fuse the feature map with spatiotemporal information to generate the output feature that fuses the temporal dynamic information.
4. The image-based video generation method according to claim 3, characterized in that: The two-layer optimization framework is used to perform a differentiable search based on the information entropy in a predefined search space to obtain the optimal spatiotemporal processing path, including: Predicting a channel adjustment amount using a learnable mapping function based on the information entropy; Determining a target number of channels based on the channel adjustment amount and a preset basic number of channels; According to the information entropy obtained by calculation, the architecture parameters are obtained through differentiable search within a two-level optimization framework; Constructing a spatiotemporal aggregation operator set based on the predefined search space, wherein the spatiotemporal aggregation operator set is configured to generate a candidate operator set with channel constraints according to the number of target channels; Optimizing the architecture parameters by using a loss function, and selecting the architecture parameters with the minimum loss from the optimized architecture parameters; The candidate operator set is instantiated using the architecture parameters with the smallest loss to obtain the optimal spatiotemporal processing path.
5. The image-based video generation method according to claim 4, characterized in that: The adopting the optimal spatiotemporal processing path to fuse the feature map with spatiotemporal information to generate the output feature fused with temporal dynamic information includes: Extracting hierarchical spatiotemporal information from the feature map according to the optimal spatiotemporal processing path; Through the gated fusion mechanism, the spatiotemporal information is weightedly fused according to scale to generate output features that incorporate temporal dynamic information.
6. The image-based video generation method according to claim 5, characterized in that: The generating of the video frame based on the temporal recurrent neural network and the output feature includes: Determine whether there is a video frame before the current moment; If there is no video frame before the current moment, inputting the output feature and the memory state of the output feature into the temporal recurrent neural network to generate the video frame at the current moment; If there is a video frame before the current moment, the video frame at the previous moment and the memory state of the video frame at the previous moment are obtained, and the video frame at the previous moment and the memory state of the video frame at the previous moment are input into the temporal recurrent neural network to generate the video frame at the current moment.
7. The image-based video generation method according to claim 6, characterized in that: After inputting the video frame into a video frame sequence and generating a target video through the video frame sequence, the method further includes: The video sequence is subjected to discriminant optimization and post-processing to output a final video.
8. An image-based video generation device, characterized in that include: A receiving module is used to receive and preprocess the input static image to generate multi-scale features; A fusion module is configured to determine an optimal spatiotemporal processing path based on the multi-scale features through a differentiable search of a spatiotemporal architecture generator, and generate output features that incorporate temporal dynamic information; A generation module, configured to generate a video frame based on a temporal recurrent neural network and the output features; The output module is used to input the video frame into a video frame sequence and generate a target video through the video frame sequence.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the image-based video generation method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the image-based video generation method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Edge-assisted unmanned aerial vehicle cooperative adaptive aerial view sensing method and system
CN122289905A
An edge-assisted drone collaborative and adaptive bird's-eye view perception method and system
CN122289905B