Picture material automatic expansion method and device, equipment and medium

By generating Gaussian noise and hidden vectors for each frame of the extended video, constraining the pictures with text content, and generating occlusion pictures and streaming pictures sequences, the problems of low expansion efficiency and insufficient diversity of picture materials in the prior art are solved, and efficient and diversified picture expansion effects are achieved.

CN120198527APending Publication Date: 2025-06-24PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510283645.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art has shortcomings in handling semantic content retention and diversity generation, and it is difficult to improve the efficiency and diversity of image material expansion without changing image semantics.

Method used

By doing Gaussian noise on each frame of the extended video, a collection of noise pictures is generated and mapped to a preset vector space to obtain the hidden vector. Combining text content and hidden vectors, the noise picture collection is constrained, and the occlusion picture sequence and flow picture sequence are generated, and the picture is expanded based on this information.

Benefits of technology

It improves the efficiency and diversity of image material expansion, ensures that the generated extended videos are consistent in the time dimension, and avoids the problem of semantic inconsistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198527A_ABST
    Figure CN120198527A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image detection, can be applied to business system platforms of financial science and technology, medical health and the like, and discloses a method, a device, equipment and a medium for automatically expanding picture materials, which comprises the following steps of: performing Gaussian noise addition on each frame of picture in a video to be expanded to obtain a noise picture set; selecting a target noise picture from the noise picture set, and mapping the target noise picture to a preset vector space to obtain an implicit vector corresponding to the target noise picture; obtaining text content in a video to be expanded, and performing picture constraint on the noise picture set according to the text content and the implicit vector to obtain an occlusion picture sequence; generating a flow graph sequence corresponding to the noise picture set by using the shielding graph sequence; and performing picture expansion on each frame of noise picture of the to-be-expanded video according to the implicit vector, the shielding graph sequence and the flow graph sequence to obtain an expanded video corresponding to the to-be-expanded video. And the efficiency and diversity of picture material expansion are effectively improved on the basis of not changing semantics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image detection, and particularly to a method, device, equipment and medium for automatically expanding picture materials. Background Art

[0002] In the fields of medical health and fintech, the application scenarios of image material expansion technology are gradually increasing, but there are still many deficiencies in the existing technology in terms of semantic content retention and diversity generation.

[0003] For example, in the field of medical health, the expansion and editing of medical images need to face a large amount of complex and diverse image data. The image data often shows differences due to different shooting angles, lighting conditions, and imaging devices, making it difficult for existing expansion and editing technologies to achieve the diversification of image content while retaining the core features of the lesion area.

[0004] For example, in the field of fintech business, image expansion and editing technologies are mainly applied to the batch production and content creation of marketing materials. For example, in the insurance marketing scenario, insurance companies need to generate diverse marketing materials to meet the needs of different customers. However, existing image expansion technologies are difficult to achieve the diversification of materials while keeping the core semantic information unchanged.

[0005] Existing image expansion and editing technologies mainly perform picture expansion based on manual methods. For example, data augmentation methods such as rotating, flipping, adding filters, and changing hues to pictures will not change the semantic content in the pictures. However, there is generally a need for diversity in picture materials. Without changing the semantics, traditional manual methods cannot achieve the expansion of a large number of picture materials, resulting in low efficiency. Therefore, how to improve the efficiency and diversity of picture material expansion without changing the semantics has become an urgent problem to be solved. Summary of the Invention

[0006] The present invention provides a method, device, equipment and medium for automatically expanding picture materials, and its main purpose is to solve the problems of poor efficiency and few expanded materials in picture material expansion without changing the semantics.

[0007] In a first aspect, to achieve the above object, a method for automatically expanding picture materials provided by the present invention includes:

[0008] Obtain a video to be expanded, and perform Gaussian noise addition on each frame of picture in the video to be expanded to obtain a corresponding set of noise pictures;

[0009] Select a target noise picture from the set of noise pictures, and map the target noise picture to a preset vector space to obtain a latent vector corresponding to the target noise picture;

[0010] Obtain the text content in the video to be extended, and perform picture constraints on the set of noise pictures according to the text content and the hidden vector to obtain a sequence of occlusion maps;

[0011] Generate a sequence of flow maps corresponding to the set of noise pictures by using the sequence of occlusion maps;

[0012] Perform picture extension on each frame of noise picture of the video to be extended according to the hidden vector, the sequence of occlusion maps and the sequence of flow maps to obtain an extended video corresponding to the video to be extended.

[0013] In a second aspect, the present invention also provides an automatic picture material extension device, including:

[0014] A Gaussian noise addition module, configured to obtain a video to be extended and perform Gaussian noise addition on each frame of picture in the video to be extended to obtain a corresponding set of noise pictures;

[0015] A hidden vector generation module, configured to select a target noise picture from the set of noise pictures, map the target noise picture to a preset vector space, and obtain a hidden vector corresponding to the target noise picture;

[0016] An occlusion map generation module, configured to obtain the text content in the video to be extended, and perform picture constraints on the set of noise pictures according to the text content and the hidden vector to obtain a sequence of occlusion maps;

[0017] A flow map sequence generation module, configured to generate a sequence of flow maps corresponding to the set of noise pictures by using the sequence of occlusion maps;

[0018] A picture extension module, configured to perform picture extension on each frame of noise picture of the video to be extended according to the hidden vector, the sequence of occlusion maps and the sequence of flow maps to obtain an extended video corresponding to the video to be extended.

[0019] In a third aspect, the present invention also provides an electronic device, where the electronic device includes:

[0020] At least one processor; and,

[0021] A memory communicatively connected to the at least one processor; wherein,

[0022] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method for automatically extending a picture material as described above.

[0023] Fourthly, the present invention further provides a computer-readable storage medium, in which at least one computer program is stored, and the at least one computer program is executed by a processor in an electronic device to implement the method for automatically expanding picture materials described above.

[0024] In the present invention, Gaussian noise is added to each frame of the picture in the video to be expanded, making the data more diverse, improving the robustness of the model under different noise levels and the generalization ability of the model; after converting the picture data into low-dimensional latent vectors, the data volume is greatly reduced, the computational complexity is reduced, and at the same time, the latent vectors can capture the key features of the pictures, providing richer feature representations and enhancing the feature representation ability; by combining the text content and the latent vectors, the semantic information of the pictures can be more comprehensively understood, enhancing the semantic understanding ability of the pictures and supporting data augmentation; by generating a sequence of occlusion maps, the situation where some areas of the pictures are occluded can be simulated, thereby improving the robustness of the model under occlusion conditions; by converting the sequence of occlusion maps into a sequence of flow maps, the change situation of the occluded parts in the images can be clearly shown, improving the understanding ability of dynamic scenes, and using the information of the sequence of occlusion maps and the sequence of flow maps, the details lost in the noisy pictures can be restored, making the generated expanded video more natural and having consistency in the time dimension, avoiding the problem of semantic inconsistency, and thus improving the efficiency and diversity of picture material expansion. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0026] Figure 1 It is a schematic diagram of an application environment of a method for automatically expanding picture materials in an embodiment of the present invention;

[0027] Figure 2 It is a schematic flowchart of a method for automatically expanding picture materials provided by an embodiment of the present invention;

[0028] Figure 3 It is a schematic flowchart of picture expansion for each frame of noisy pictures of the video to be expanded provided by an embodiment of the present invention;

[0029] Figure 4 It is a schematic diagram of modules of a device for automatically expanding picture materials provided by an embodiment of the present invention;

[0030] Figure 5Schematic structural diagram of an electronic device for implementing a method for automatically expanding picture materials according to an embodiment of the present invention;

[0031] Figure 6 Another schematic structural diagram of an electronic device for implementing a method for automatically expanding picture materials according to an embodiment of the present invention.

[0032] The implementation, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Specific embodiments

[0033] In order to enable those skilled in the art of the present technology to better understand the technical solutions of the present disclosure, and to fully understand how the present disclosure uses technical means to solve technical problems and the implementation process of achieving corresponding technical effects and to implement accordingly, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. The embodiments of the present disclosure and each feature in the embodiments can be combined with each other without conflict, and the formed technical solutions are all within the protection scope of the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure.

[0034] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or equipment.

[0035] An embodiment of the present application provides a method for automatically expanding picture materials. The execution subject of the method for automatically expanding picture materials includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the device provided in the embodiment of the present application. In other words, the method for automatically expanding picture materials can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0036] A method for automatically expanding picture materials according to the present invention can be applied in an application environment such as Figure 1 . Among them, the client communicates with the server through the network. The server can obtain the video to be expanded through the client, perform Gaussian noise addition on each frame of the picture in the video to be expanded, make the data more diverse, improve the robustness of the model under different noise levels and the generalization ability of the model; after converting the picture data into a low-dimensional hidden vector, the data volume is greatly reduced, the computational complexity is reduced, and at the same time the hidden vector can capture the key features of the picture, provide a richer feature representation, and enhance the feature representation ability; by combining the text content and the hidden vector, the semantic information of the picture can be more comprehensively understood, the semantic understanding ability of the picture is enhanced, and data augmentation is supported; by generating a sequence of occlusion maps, the situation where some areas of the picture are occluded can be simulated, thereby improving the robustness of the model under occlusion conditions; by converting the sequence of occlusion maps into a sequence of flow maps, the change situation of the occluded part in the image can be clearly shown, improving the understanding ability of dynamic scenes. Using the information of the sequence of occlusion maps and the sequence of flow maps, the details lost in the noisy picture can be restored, making the generated expanded video more natural and having consistency in the time dimension, avoiding the problem of semantic inconsistency, thereby improving the efficiency and diversity of picture material expansion. Finally, the expanded video is output and fed back to the client. Among them, the client can be, but is not limited to, various personal computers, laptop computers, smartphones, tablet computers, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.

[0037] The following is an explanation of the specification of the present invention. The present invention adopts a technology for automatically expanding picture materials based on a denoising diffusion probability model. By adding Gaussian noise to each frame of the video to be expanded, a corresponding set of noise pictures is obtained. A target noise picture is selected from the set of noise pictures, and the corresponding latent vector of the target noise picture is obtained. According to the text content and latent vector in the noise picture, the set of noise pictures is constrained to obtain a sequence of occlusion pictures and a sequence of flow pictures. According to the latent vector, the sequence of occlusion pictures and the sequence of flow pictures, each frame of the noise picture of the video to be expanded is expanded to obtain an expanded video corresponding to the video to be expanded.

[0038] Referring to Figure 2 As shown, it is a schematic flowchart of a method for automatically expanding picture materials provided by an embodiment of the present invention. In this embodiment, the method for automatically expanding picture materials includes:

[0039] S1. Obtain the video to be expanded, and add Gaussian noise to each frame of the video to be expanded to obtain a corresponding set of noise pictures.

[0040] In the embodiment of the present invention, the video to be expanded may be video data including multiple consecutive frames of pictures. The Gaussian noise addition refers to gradually increasing the noise intensity through a preset number of noise addition steps and transformation slopes to ensure that each frame of picture can gradually transition to a noise state during the noise addition process.

[0041] In the embodiment of the present invention, the video to be expanded may be crawled from a web page through web crawler technology. For example, a video file can be obtained from a relevant social media video platform using a Python crawler; it can also use programming languages such as Python and Java to connect to a preset video database through the HTTP protocol and send an HTTP request to obtain the video data to be expanded within a reference time period.

[0042] Specifically, the video to be expanded may be video material of a medical experiment scenario, an official hospital video, etc., and may also be an advertising video and a product introduction video of a financial product, etc.

[0043] In the embodiment of the present invention, adding Gaussian noise to each frame of the video to be expanded to obtain a corresponding set of noise pictures includes:

[0044] Determine the Gaussian noise intensity corresponding to each noise addition step according to a preset number of noise addition steps to form a Gaussian noise intensity sequence;

[0045] Generate the Gaussian noise corresponding to each noise addition step step by step according to the Gaussian noise intensity sequence;

[0046] Perform Gaussian noise addition to each frame image in the video to be extended according to multiple pieces of the Gaussian noise, and obtain a corresponding set of noise images.

[0047] Among them, the Gaussian noise is calculated through the following formula:

[0048] ε = αt

[0049] Among them, ε represents the Gaussian noise, α represents a preset transformation slope, and t represents the number of noise addition steps.

[0050] Specifically, the number of noise addition steps refers to the total number of times Gaussian noise is added, which can be set to 1000 steps. The transformation slope is used to control the intensity change of noise addition in each step, which can be set to 0.02. By adjusting the number of noise addition steps and the transformation slope, the rate of noise addition and the quality of the noise images can be precisely controlled, ensuring the gradualness and controllability of noise addition.

[0051] In detail, gradually add random noise conforming to the Gaussian distribution to each frame image in the video to be extended according to the number of noise addition steps and the transformation slope. The more steps there are, the greater the added Gaussian noise. Generate a random noise matrix conforming to the Gaussian distribution according to the Gaussian noise intensity, and its size is the same as the image size. Superimpose the generated noise matrix on the current image, and repeat the above process for each frame image until all addition steps are completed. After all addition steps are completed, use the result image corresponding to the preset number of steps as the final noise image, and collect the noise images into a set of noise images.

[0052] Exemplarily, the Gaussian noise addition scenario of the present invention can be applied to medical video data enhancement, training of medical image denoising models, etc. Currently, the amount of medical video data is limited, and there may be significant differences in the video data of different cases, resulting in insufficient or overfitting of model training data. Performing Gaussian noise addition processing on the original video frames to generate diverse noise images can simulate video data under different devices, environments, or transmission conditions, thereby expanding the training dataset and improving the generalization ability of the model. For example, in surgical video analysis, by performing Gaussian noise addition processing on the original video frames to generate video data with different noise intensities for training a surgical action recognition model, the surgical actions can be accurately recognized under different noise conditions, improving the robustness of the surgical assistance system and the clarity and diagnostic accuracy of the images.

[0053] In the embodiments of the present invention, Gaussian noise addition is performed on each frame image in the video to be extended, making the data more diverse and improving the robustness of the model under different noise levels and the generalization ability of the model.

[0054] S2. Select a target noise picture from the set of noise pictures, and map the target noise picture to a preset vector space to obtain a latent vector corresponding to the target noise picture.

[0055] In the embodiments of the present invention, the target noise picture may be the first-frame noise picture in the noise pictures corresponding to the video to be extended, and is used to generate a corresponding latent vector for further processing or analysis in the vector space; the latent vector refers to a low-dimensional numerical vector obtained by mapping the target noise picture to a preset vector space, which is a mathematical representation of the picture features and can capture the key information of the noise picture.

[0056] In the embodiments of the present invention, the mapping of the target noise picture to a preset vector space to obtain a latent vector corresponding to the target noise picture includes:

[0057] Extract the key feature information in the target noise picture through a preset feature extraction model to obtain target features;

[0058] Obtain the feature dimension parameter of the preset vector space, and map the target features to the vector space according to the feature dimension parameter to generate a latent vector corresponding to the target noise picture.

[0059] Specifically, the present invention can use a preset feature extraction model such as a Convolutional Neural Network (CNN) to extract target features from the key feature information in the target noise picture. The target features may include edges, textures, color distributions, etc. According to the feature dimension parameter of the preset vector space, such as the number of dimensions can be set to 128, 256, etc., use a preset image model encoder such as a Vision Transformer Encode (ViT) to convert the target features into a numerical vector in the vector space, ensuring that the dimension of the generated latent vector is consistent with the feature dimension parameter.

[0060] Exemplarily, in the scenario of latent vector generation in the field of medical and health, it can be applied to medical diagnosis. For example, the first-frame noise picture of an endoscopic video can be selected as the target noise picture, and a Convolutional Neural Network (CNN) can be used to extract the lesion area features in the target noise picture, map the extracted features to a 128-dimensional vector space to generate a latent vector, and input the latent vector into a medical lesion classification model to determine the type of lesion.

[0061] The hidden vector generation scenario of the present invention can also be applied to the fintech scenario for identity verification. Select the first-frame noise picture of the user identity verification video as the target noise picture, use the face recognition algorithm to extract the facial features of the user, map the extracted features to a 256-dimensional vector space to generate a hidden vector, and compare the hidden vector with the user hidden vector in the database to complete the identity verification.

[0062] In the embodiment of the present invention, after converting the picture data into a low-dimensional hidden vector, the data volume is greatly reduced, the computational complexity is reduced, and at the same time, the hidden vector can capture the key features of the picture (such as edges, textures, color distributions, etc.), provide a richer feature representation, improve the picture data processing efficiency, and enhance the feature representation ability.

[0063] S3. Obtain the text content in the video to be extended, and perform picture constraints on the noise picture set according to the text content and the hidden vector to obtain an occlusion map sequence.

[0064] In the embodiment of the present invention, the occlusion map sequence refers to a group of occlusion maps arranged in chronological order, indicating the occlusion conditions of different regions or time points in the noise picture. The text content can be obtained from the video to be extended by using artificial text annotation or optical character recognition technology, and the noise picture can be denoised by combining the text vector and the hidden vector through the denoising diffusion probability model to generate an occlusion map sequence.

[0065] In the embodiment of the present invention, the video to be extended can be manually text-annotated, and the annotated pictures can be described in text to obtain the text content in the video to be extended; the present invention can also directly extract the text content in the video to be extended through optical character recognition technology, and thus the text content in the video to be extended can also be obtained.

[0066] Specifically, the obtaining of the text content in the video to be extended includes:

[0067] Perform standardization processing on the video to be extended to obtain a standard video to be extended;

[0068] Use the edge detection method to perform text localization on the standard video to be extended to obtain the corresponding text region;

[0069] Extract the text content in the video to be extended according to the optical character recognition technology for the text region.

[0070] Specifically, the standardization process refers to converting each frame of the video to be expanded into a unified format and size to reduce the impact of different picture sizes on text extraction. Image processing libraries can be used to perform standardization operations such as grayscale conversion and binarization on the pictures. The edges of the pictures are detected through edge detection algorithms (such as Canny edge detection), and the text regions are found through contour analysis. The located text regions are used to extract the text content according to optical character recognition technology, providing semantic support for subsequent picture constraints.

[0071] Among them, since the semantic meaning of the noise images remains unchanged, the text content can also be directly obtained by extracting the noise picture set.

[0072] Exemplarily, in the field of medical and health, the text content can be patient identification information, diagnosis conclusions and indicators, drug prescription data, etc. For example, if a medicine bottle picture needs to be expanded, the picture can be expanded according to the text content such as the medicine name and dosage in the drug prescription data, thus generating various medicine bottle shapes without changing the drug data information. In the field of fintech, such as in the insurance industry, the text content can be the core terms of the insurance policy, claim information summary, etc.

[0073] In the embodiments of the present invention, the step of performing picture constraints on the noise picture set according to the text content and the latent vector to obtain a sequence of occlusion maps includes:

[0074] Encoding the text content using a preset text encoder to obtain a text vector;

[0075] Performing vector fusion on the text vector and the latent vector to obtain a fusion vector;

[0076] Using the fusion vector as a picture constraint condition, and using a preset denoising diffusion probability model to denoise the noise picture set according to the picture constraint condition to obtain a sequence of occlusion maps.

[0077] Specifically, a text encoder such as the BERT Embedding model (BERT) can be used to encode text content into text vectors. Methods such as weighted summation or feature concatenation are used to fuse the text vectors with the latent vectors to obtain fused vectors. Using the fused vectors as constraint conditions, the fused vectors are input into a pre-trained denoising diffusion probability model. The denoising diffusion probability model can be a U-net model, etc., which is usually used for image denoising or image generation tasks. The noisy image is input into the denoising diffusion probability model, and forward propagation is performed through the encoder-decoder structure of the model. That is, in the encoding and denoising stage of the model, high-level features of the noisy image are gradually extracted while reducing the spatial dimension. In the decoding and denoising stage of the model, the high-level features are gradually transformed to obtain the denoised image corresponding to the feature map with a larger spatial dimension. During the denoising process, the fused vectors are used as constraint conditions to generate a sequence of occlusion maps that meet specific constraints. Each occlusion map in the sequence of occlusion maps represents the occlusion information at a specific moment, and each occlusion map contains spatial sequence features. The occlusion map is usually a binary map, indicating which areas are occluded (value is 1) and which areas are not occluded (value is 0).

[0078] In the embodiments of the present invention, by combining text content and latent vectors, the semantic information of the picture can be more comprehensively understood, the semantic understanding ability of the picture can be enhanced, and data augmentation is supported; by generating a sequence of occlusion maps, the situation where some areas in the picture are occluded can be simulated, thereby improving the robustness of the model under occlusion conditions.

[0079] S4. Generate a sequence of flow maps corresponding to the set of noisy images by using the sequence of occlusion maps.

[0080] In the embodiments of the present invention, the sequence of flow maps is used to characterize the dynamic changes of the sequence of occlusion maps in time or space.

[0081] In the embodiments of the present invention, the generating a sequence of flow maps corresponding to the set of noisy images by using the sequence of occlusion maps includes:

[0082] Extract the graph features corresponding to each occlusion map in the sequence of occlusion maps, and generate a sequential encoding for each occlusion map;

[0083] Extract the position features of the sequential encoding, splice the position features and the graph features to obtain the flow map features corresponding to the set of noisy images, and generate a sequence of flow maps according to the flow map features.

[0084] Specifically, the graph features corresponding to each occlusion map can be extracted through a convolutional neural network. The graph features can be edges, textures, shapes, etc. The sequential encoding is used to represent the position order of the occlusion maps in the sequence. The position feature refers to the information extracted from the sequential encoding that can represent the position of the occlusion map in the sequence. The position feature is crucial for understanding the sequential relationship between the occlusion maps. The feature concatenation refers to combining the position feature and the graph feature into a new feature vector, which helps to combine temporal information and spatial information to form a more comprehensive representation. The feature concatenation can be completed through a simple connection operation (such as the concatenate function of numpy or an attention mechanism), thereby obtaining the flow graph features. Then, a series of flow graph sequences representing the change of feature points in the image over time are generated through visualization algorithms, neural network generation, etc.

[0085] In the embodiment of the present invention, generating the sequential encoding for each of the occlusion maps includes:

[0086] Obtain the initial sequential encoding vector for each occlusion map, and initialize a learnable matrix according to the number and dimension of the occlusion maps;

[0087] Multiply the initial sequential encoding vector by the learnable matrix to obtain the sequential encoding for each occlusion map.

[0088] Specifically, different from the traditional sequential encoding (positional encoding) which is directly calculated using sine and cosine functions (sin / cos functions), the present invention can initialize a learnable matrix of k×d, where k is the number of occlusion maps and d is the dimension of the occlusion maps, generate an initial sequential encoding vector for each occlusion map, and multiply the initial sequential encoding vector by the learnable matrix, thereby obtaining the sequential encoding for each occlusion map, simplifying the computational complexity and improving the efficiency of sequential encoding generation.

[0089] In the embodiment of the present invention, by converting the occlusion map sequence into a flow graph sequence, the change situation of the occluded part in the image can be clearly displayed, improving the understanding ability of dynamic scenes. The generated flow graph sequence can be used as the input for subsequent image processing, such as image expansion, etc., to improve the stability and robustness of image expansion.

[0090] S5. Perform image expansion on each frame of noise image of the video to be expanded according to the latent vector, the occlusion map sequence, and the flow graph sequence, to obtain the expanded video corresponding to the video to be expanded.

[0091] In the embodiment of the present invention, corresponding latent space sequence vectors are generated through the latent vector, the occlusion map sequence, and the flow graph sequence, and then a preset decoder uses a transposed convolutional layer or an upsampling layer to gradually map the low-dimensional features back to the high-dimensional image space, thereby generating an expanded video.

[0092] Referring to Figure 3 As shown, in the embodiment of the present invention, the method for expanding each frame of noise picture of the video to be expanded according to the latent vector, the occlusion map sequence and the flow map sequence to obtain the expanded video corresponding to the video to be expanded includes:

[0093] Performing linear and non-linear transformation fusion on each frame of noise picture of the video to be expanded according to the latent vector, the occlusion map sequence and the flow map sequence to obtain a latent space sequence vector;

[0094] Decoding the latent space sequence vector to obtain a target picture sequence;

[0095] Generating the expanded video corresponding to the video to be expanded according to the target picture sequence.

[0096] Specifically, the latent space sequence vector can be shown as the following formula:

[0097] z = m * w(x, n)

[0098] Wherein, z represents the latent space sequence vector, m represents the occlusion map sequence, w represents the linear and non-linear transformation fusion, x represents the latent vector, and n represents the flow map sequence.

[0099] Specifically, the latent space sequence vector is a low-dimensional vector, representing the latent features of the noise picture, the occlusion information and the motion trend of the flow map sequence, and is used as the input of the decoder to generate the target picture; the linear and non-linear transformation fusion includes operations such as weighted summation and activation functions, which can capture the linear and non-linear relationships between vectors.

[0100] Specifically, the flow map sequence is concatenated with the latent vector in terms of features, and the concatenated vector is subjected to linear transformation and non-linear transformation fusion through a fusion network. The fused feature vector is multiplied by the noise picture corresponding to the occlusion map sequence to obtain the latent space sequence vector. Then, a preset decoder such as a 3D Causal Decoder extracts features from the latent space sequence vector to obtain the latent space features corresponding to the latent space sequence vector. Then, according to the stacking module (Dec Stage) in the decoder, the latent space features are upsampled by a factor of two, and the above steps are repeated to gradually map the low-dimensional features back to the high-dimensional image space, thereby generating a target picture sequence, and video synthesis is performed on the target picture sequence to obtain the expanded video corresponding to the video to be expanded.

[0101] Exemplarily, in the field of healthcare, the video to be extended can be a real-time recording in an operating room, and the extended video can be a surgical process simulation, generating a detailed simulation of surgical operations. For example, a real-time recording of a coronary artery bypass grafting surgery can be extended to generate a microscopic view of blood vessel suturing, showing the path of the suture and the dynamic effect of blood flow restoration, thereby helping to understand the surgical details; it can also be an electrocardiogram waveform, which can be extended to generate a 3D simulation of the heart beating, combining electrocardiogram data to show the dynamic process of heart contraction and relaxation, thereby helping doctors and patients to more intuitively understand the changes in health data.

[0102] In the field of fintech technology, the video to be extended can be the interface operation of a user submitting a claim application. Extending this can generate a dynamic process of background review, risk assessment, and claim calculation, and mark the time nodes of each link, which can be used for customer education or employee training, thereby improving process transparency.

[0103] In the embodiments of the present invention, by using the information of the occlusion map sequence and the flow map sequence, the details lost in the noisy picture can be restored, making the generated extended video more natural, having consistency in the time dimension, avoiding the problem of semantic inconsistency, and thus improving the efficiency and diversity of picture material expansion.

[0104] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0105] As Figure 4 shown, it is a functional module diagram of a picture material automatic expansion device provided by an embodiment of the present invention.

[0106] In the embodiments of the present disclosure, a picture material automatic expansion device is provided. The picture material automatic expansion device corresponds one-to-one with the method of picture material automatic expansion in the above embodiments. As Figure 4 shown, the picture material automatic expansion device 100 can be installed in an electronic device. According to the functions implemented, the picture material automatic expansion device 100 includes a Gaussian noise addition module 101, a latent vector generation module 102, an occlusion map generation module 103, a flow map sequence generation module 104, and a picture expansion module 105. The detailed description of each functional module is as follows:

[0107] The Gaussian noise addition module 101 is used to obtain the video to be extended, and perform Gaussian noise addition on each frame picture in the video to be extended to obtain a corresponding set of noisy pictures;

[0108] The latent vector generation module 102 is configured to select a target noise picture from the set of noise pictures, map the target noise picture to a preset vector space, and obtain a latent vector corresponding to the target noise picture;

[0109] The occlusion map generation module 103 is configured to obtain the text content in the video to be extended, perform picture constraints on the set of noise pictures according to the text content and the latent vector, and obtain a sequence of occlusion maps;

[0110] The flow map sequence generation module 104 is configured to generate a sequence of flow maps corresponding to the set of noise pictures by using the sequence of occlusion maps;

[0111] The picture extension module 105 is configured to perform picture extension on each frame of noise picture of the video to be extended according to the latent vector, the sequence of occlusion maps, and the sequence of flow maps, and obtain an extended video corresponding to the video to be extended.

[0112] In one embodiment, when the Gaussian noise addition module 101 performs Gaussian noise addition on each frame of picture in the video to be extended to obtain a corresponding set of noise pictures, it is configured to:

[0113] Determine the Gaussian noise intensity corresponding to each noise addition step according to a preset number of noise addition steps, and form a sequence of Gaussian noise intensities;

[0114] Gradually generate Gaussian noise corresponding to each noise addition step according to the sequence of Gaussian noise intensities;

[0115] Perform Gaussian noise addition on each frame of picture in the video to be extended according to multiple pieces of the Gaussian noise respectively, and obtain a corresponding set of noise pictures.

[0116] In one embodiment, when the latent vector generation module 102 performs mapping the target noise picture to a preset vector space to obtain a latent vector corresponding to the target noise picture, it is configured to:

[0117] Extract key feature information in the target noise picture through a preset feature extraction model to obtain target features;

[0118] Obtain the feature dimension parameter of the preset vector space, map the target features to the vector space according to the feature dimension parameter, and generate a latent vector corresponding to the target noise picture.

[0119] In one embodiment, when the occlusion map generation module 103 performs obtaining the text content in the video to be extended, it is configured to:

[0120] Perform normalization processing on the video to be extended to obtain a standard video to be extended;

[0121] Use an edge detection method to perform text localization on the standard video to be extended, and obtain the corresponding text region;

[0122] Extract the text content in the video to be extended according to the optical character recognition technology for the text region.

[0123] In one embodiment, when the occlusion map generation module 103 executes the operation of performing picture constraint on the noise picture set according to the text content and the latent vector to obtain an occlusion map sequence, it is used for:

[0124] Use a preset text encoder to encode the text content to obtain a text vector;

[0125] Perform vector fusion on the text vector and the latent vector to obtain a fusion vector;

[0126] Use the fusion vector as a picture constraint condition, and use a preset denoising diffusion probability model to denoise the noise picture set according to the picture constraint condition to obtain an occlusion map sequence.

[0127] In one embodiment, when the flow graph sequence generation module 104 executes the operation of generating a flow graph sequence corresponding to the noise picture set by using the occlusion map sequence, it is used for:

[0128] Extract the graph features corresponding to each occlusion map in the occlusion map sequence, and generate a sequential encoding for each occlusion map;

[0129] Extract the position features of the sequential encoding, splice the position features and the graph features to obtain the flow graph features corresponding to the noise picture set, and generate a flow graph sequence according to the flow graph features.

[0130] In one embodiment, when the flow graph sequence generation module 104 executes the operation of generating a sequential encoding for each occlusion map, it is used for:

[0131] Obtain the initial sequential encoding vector of each occlusion map, and initialize a learnable matrix according to the number and dimension of the occlusion maps;

[0132] Multiply the initial sequential encoding vector by the learnable matrix to obtain the sequential encoding of each occlusion map.

[0133] In one embodiment, when the picture extension module 105 executes the operation of performing picture extension on each frame of noise picture of the video to be extended according to the latent vector, the occlusion map sequence and the flow graph sequence to obtain an extended video corresponding to the video to be extended, it is used for:

[0134] Perform linear and non-linear transformation fusion on each frame of noise picture of the video to be extended according to the hidden vector, the occlusion map sequence and the flow map sequence to obtain a hidden space sequence vector;

[0135] Decode the hidden space sequence vector to obtain a target picture sequence;

[0136] Generate an extended video corresponding to the video to be extended according to the target picture sequence.

[0137] In the present invention, the specific limitations on a picture material automatic extension device can refer to the limitations on a method for automatic extension of picture materials above, which will not be elaborated here. Each module in the above picture material automatic extension device can be implemented in whole or in part by software, hardware and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0138] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 5 shown. The computer device includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a method for automatic extension of picture materials.

[0139] In one embodiment, a computer device is provided. The computer device can be a client, and its internal structure diagram can be as Figure 6 shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of a method for automatic extension of picture materials.

[0140] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:

[0141] Obtain the video to be extended, perform Gaussian noise addition on each frame picture in the video to be extended, and obtain a corresponding set of noise pictures;

[0142] Select a target noise picture from the set of noise pictures, map the target noise picture to a preset vector space, and obtain a latent vector corresponding to the target noise picture;

[0143] Obtain the text content in the video to be extended, and perform picture constraint on the set of noise pictures according to the text content and the latent vector to obtain a sequence of occlusion maps;

[0144] Generate a sequence of flow maps corresponding to the set of noise pictures by using the sequence of occlusion maps;

[0145] Perform picture extension on each frame of noise picture of the video to be extended according to the latent vector, the sequence of occlusion maps, and the sequence of flow maps, and obtain an extended video corresponding to the video to be extended.

[0146] In several embodiments provided by the present invention, it should be understood that the disclosed devices and apparatuses can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.

[0147] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware, or in the form of a combination of hardware and software functional modules.

[0148] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be encompassed by the present invention. Any associated drawing marks in the claims should not be regarded as limiting the claimed rights.

[0149] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention.

[0150] In some embodiments of the present embodiment, a computer-readable storage medium is provided, on which a computer program is stored, characterized in that when the computer program is executed by a processor, the steps of the method described in the above embodiment are implemented.

[0151] The readable storage medium of the present invention stores a computer program, and when the computer program is executed by a processor of an electronic device, it can implement:

[0152] Obtain the video to be extended, perform Gaussian noise addition on each frame picture in the video to be extended, and obtain a corresponding set of noise pictures;

[0153] Select a target noise picture from the set of noise pictures, map the target noise picture to a preset vector space, and obtain a latent vector corresponding to the target noise picture;

[0154] Obtain the text content in the video to be extended, and perform picture constraint on the set of noise pictures according to the text content and the latent vector to obtain a sequence of occlusion maps;

[0155] Generate a sequence of flow maps corresponding to the set of noise pictures by using the sequence of occlusion maps;

[0156] Perform picture extension on each frame of noise picture of the video to be extended according to the latent vector, the sequence of occlusion maps and the sequence of flow maps, and obtain an extended video corresponding to the video to be extended.

[0157] It should be noted that for the functions or steps that the above computer-readable storage medium or computer device can implement, reference can be made to the relevant descriptions on the server side and the client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0158] The computer-readable storage medium can also store at least one computer executable program / instruction. The computer executable program / instruction is, for example, a computer-readable instruction. The computer-readable storage medium includes but is not limited to, for example, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The computer-readable storage medium may include, for example, read-only memory (ROM), hard disk, flash memory, etc. For example, a non-transitory computer-readable storage medium can be connected to a computing device such as a computer. Then, when the computing device runs the computer-readable instructions stored on the computer-readable storage medium, the various methods described above can be performed.

[0159] In addition, the computer device may further include (but is not limited to) a data bus, an input / output (I / O) bus, a display, and input / output devices (such as a keyboard, a mouse, a speaker, etc.).

[0160] The processor can communicate with external devices via a wired or wireless network through an I / O bus.

[0161] In one embodiment, the at least one computer-executable instruction may also be compiled into or form a software product / computer program product, where when one or more computer-executable instructions are run by a processor, they perform the various functions and / or method steps in the embodiments described in this technology.

[0162] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it may include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application may include non-volatile and / or volatile memories. Non-volatile memories may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0163] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In practical applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.

[0164] In the embodiments provided in the present disclosure, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0165] The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

[0166] It should be noted that if non-company software tools or components appear in the embodiments of the present application, they are only used for illustrative introduction and do not represent actual use.

Claims

1. A method for automatically expanding picture materials, characterized in that: The method comprises: Obtain a video to be expanded, and perform Gaussian noise addition on each frame of the video to be expanded to obtain a corresponding noise picture set; Selecting a target noise picture from the noise picture set, mapping the target noise picture to a preset vector space, and obtaining a latent vector corresponding to the target noise picture; Acquire text content in the video to be expanded, and perform image constraints on the noise image set according to the text content and the latent vector to obtain an occlusion image sequence; Generating a stream graph sequence corresponding to the noise picture set using the occlusion graph sequence; Each noise picture frame of the video to be expanded is expanded according to the latent vector, the occlusion map sequence and the flow map sequence to obtain an expanded video corresponding to the video to be expanded.

2. The method for automatically expanding picture materials according to claim 1, characterized in that: The step of performing Gaussian noise addition on each frame of the video to be expanded to obtain a corresponding noise picture set includes: Determine the Gaussian noise intensity corresponding to each noise adding step according to a plurality of preset noise adding step numbers to form a Gaussian noise intensity sequence; According to the Gaussian noise intensity sequence, gradually generate Gaussian noise corresponding to each noise addition step; Gaussian noise is performed on each frame of the video to be extended according to the multiple Gaussian noises to obtain a corresponding noise picture set.

3. The method for automatically expanding picture materials according to claim 1, characterized in that: Mapping the target noise image to a preset vector space to obtain a latent vector corresponding to the target noise image includes: Extract key feature information from the target noise image using a preset feature extraction model to obtain target features; Acquire a feature dimension parameter of a preset vector space, map the target feature to the vector space according to the feature dimension parameter, and generate a latent vector corresponding to the target noise image.

4. The method for automatically expanding picture materials according to claim 1, characterized in that: The performing picture constraints on the noise picture set according to the text content and the latent vector to obtain an occlusion picture sequence includes: Encoding the text content using a preset text encoder to obtain a text vector; Performing vector fusion on the text vector and the latent vector to obtain a fused vector; The fusion vector is used as a picture constraint condition, and a preset denoising diffusion probability model is used to denoise the noisy picture set according to the picture constraint condition to obtain an occlusion map sequence.

5. The method for automatically expanding picture materials according to claim 1, characterized in that: The step of generating a stream graph sequence corresponding to the noise picture set by using the occlusion graph sequence includes: Extracting the image features corresponding to each occlusion image in the occlusion image sequence, and generating a sequential code for each occlusion image; Extract the position features of the sequential encoding, perform feature splicing on the position features and the graph features to obtain the flow graph features corresponding to the noise picture set, and generate a flow graph sequence according to the flow graph features.

6. The method for automatically expanding picture materials according to claim 5, characterized in that: The step of generating a sequential code for each of the occlusion maps comprises: Get the initial sequential encoding vector of each occlusion map and initialize a learnable matrix according to the number and dimension of the occlusion maps; The initial sequential encoding vector is multiplied by the learnable matrix to obtain the sequential encoding of each occlusion map.

7. The method for automatically expanding picture materials according to claim 1, characterized in that: The step of performing picture expansion on each noise picture frame of the video to be expanded according to the latent vector, the occlusion map sequence, and the flow map sequence to obtain an expanded video corresponding to the video to be expanded includes: According to the latent vector, the occlusion map sequence and the flow map sequence, each frame of the noise picture of the video to be expanded is subjected to linear and nonlinear transformation fusion to obtain a latent space sequence vector; Decoding the latent space sequence vector to obtain a target picture sequence; An extended video corresponding to the to-be-extended video is generated according to the target picture sequence.

8. A device for automatically expanding picture materials, characterized in that: The device comprises: A Gaussian noise adding module is used to obtain a video to be expanded, and perform Gaussian noise adding on each frame of the video to be expanded to obtain a corresponding noise picture set; A latent vector generation module, used to select a target noise picture from the noise picture set, map the target noise picture to a preset vector space, and obtain a latent vector corresponding to the target noise picture; An occlusion map generation module is used to obtain text content in the video to be expanded, and perform image constraints on the noise image set according to the text content and the latent vector to obtain an occlusion map sequence; A flow graph sequence generation module, used to generate a flow graph sequence corresponding to the noise picture set using the occlusion graph sequence; The picture expansion module is used to expand each frame of the noise picture of the video to be expanded according to the latent vector, the occlusion map sequence and the flow map sequence to obtain an expanded video corresponding to the video to be expanded.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method for automatically expanding image materials as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, a method for automatically expanding image materials as described in any one of claims 1 to 7 is implemented.