Image sequence construction method and device, equipment and medium
By receiving initial image and text information, extracting and combining image and text feature vectors, performing denoising processing, and finally generating target image sequences through the decoding module, the problem of difficult to achieve efficient fusion of dynamic video details and multimodal information in the prior art is solved, and high-quality and flexible image generation video is achieved.
Patent Information
- Application Number
- CN202510042955.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-09
AI Technical Summary
In the prior art, in the image generation video task, it is difficult to simultaneously realize the efficient fusion of dynamic video detail fidelity and multimodal information, resulting in insufficient quality and flexibility of generating videos.
By receiving the initial image and text information, the image is divided into multiple image segments, the image and text feature vectors are extracted through the first and second network structures, combined and denoising, and finally the target image sequence is generated by the decoding module.
It realizes efficient combination of multimodal information, improves the fidelity and dynamic consistency of generated image sequences, and improves the quality and applicability of image-generated videos.
Smart Images

Figure CN119963429A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence technology and medical health, and in particular to an image sequence construction method, device, equipment and storage medium. Background Art
[0002] With the rapid development of computer vision technology, the technology of image-to-video generation has shown broad application prospects in the fields of healthcare and financial technology. In the field of healthcare, this technology is widely used in scenarios such as medical image dynamicization, surgical simulation, and rehabilitation training; in the field of financial technology, dynamic video technology helps advertising marketing, user experience optimization, and financial scenario simulation. However, the existing technology still faces many shortcomings in the actual application of these fields, which are mainly reflected in the following aspects:
[0003] Currently, the technical paths to solve the problem of image-generated video are mainly divided into two types: generation methods based on physical models and methods based on generative models:
[0004] The generation method based on physical models combines deep learning with physical rules to more realistically restore the tissue movement in medical images or the dynamic evolution of financial scenarios. However, due to the high complexity of its model and the large requirements for computing resources, the efficiency of real-time generation of high-resolution videos is low, which makes it difficult to meet the needs of real-time diagnosis in medical health and rapid material generation in financial marketing.
[0005] Generative models such as Generative Adversarial Networks (GANs) and Variational Autoencoders (VAEs) have become a hot topic of current research due to their fast generation speed. However, in practical applications, these methods are weak in detail processing, long-term consistency, and dynamics. For example, the boundary fuzziness problem in the dynamic generation of medical images and the inability to accurately capture the characteristics of time series changes in financial scenario simulations have limited their practical application effects.
[0006] In the field of medical health, dynamic medical image generation often requires the combination of multimodal data (such as MRI, CT, ultrasound, etc.) to achieve an accurate description of the patient's pathological changes. However, the existing technology has limited ability to integrate and process multimodal information, making it difficult to fully reflect the correlation between different modal data in dynamic videos.
[0007] In the field of financial technology, dynamic videos usually need to integrate multimodal information such as text, time series data and images to show complex market trends or financial products. However, the existing technology does not make full use of external information, making it difficult to accurately show the dynamic characteristics of financial scenarios in video generation.
[0008] In healthcare, dynamic imaging requires precise control of the movement and changes of the lesion area, but existing methods are insufficient in their ability to model and control the movement of specific areas, and the generated images are difficult to meet the needs of personalized diagnosis and treatment.
[0009] In FinTech, existing generative models make limited use of first-frame information and external control inputs, resulting in the generated video content being difficult to adapt to diverse user needs and lacking flexibility and targeting. Summary of the invention
[0010] The main purpose of the present invention is to provide an image sequence construction method, device, equipment and storage medium, aiming to solve the technical problem that in the prior art, in the image generation video task, it is difficult to simultaneously achieve the detail fidelity of dynamic video and the efficient fusion of multimodal information, resulting in insufficient quality and flexibility of the generated video.
[0011] To achieve the above object, the present invention provides an image sequence construction method, comprising:
[0012] receiving initial image and text information;
[0013] Dividing the initial image into a plurality of image segments;
[0014] Processing the plurality of image segments through a first network structure to obtain an image feature vector;
[0015] Processing the text information through a second network structure to obtain a text feature vector;
[0016] Merging the image feature vector, the text feature vector and the noise information to generate an initial fusion vector;
[0017] Performing denoising processing on the initial fusion vector to obtain a denoising result;
[0018] Decoding is performed on the denoising result to generate a target image sequence.
[0019] Furthermore, to achieve the above object, the present invention provides an image sequence construction device, comprising:
[0020] A data receiving module, used for receiving initial image and text information;
[0021] An image preprocessing module, used for dividing the initial image into a plurality of image segments;
[0022] A first feature extraction module, configured to process the plurality of image segments through a first network structure to obtain an image feature vector;
[0023] A second feature extraction module, used for processing the text information through a second network structure to obtain a text feature vector;
[0024] A multimodal fusion module, used for merging the image feature vector, the text feature vector and the noise information to generate an initial fusion vector;
[0025] A denoising module, used for performing denoising processing on the initial fusion vector to obtain a denoising result;
[0026] The decoding module is used to perform decoding processing on the denoising result to generate a target image sequence.
[0027] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer device, which includes a memory, a processor, and an image sequence construction program stored in the memory and executable on the processor, wherein the image sequence construction program implements the steps of the image sequence construction method described above when executed by the processor.
[0028] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which an image sequence construction program is stored, and when the image sequence construction program is executed by a processor, the steps of the image sequence construction method described above are implemented.
[0029] Beneficial effects: The present invention relates to the fields of artificial intelligence technology and medical health, and discloses a method for constructing an image sequence, including: receiving an initial image and text information, dividing the initial image into multiple image segments, extracting an image feature vector through a first network structure, extracting a text feature vector through a second network structure, fusing the image feature vector, the text feature vector and the noise information to generate an initial fusion vector, performing denoising on the initial fusion vector to obtain a denoising result, and generating a target image sequence by decoding the denoising result. The present invention realizes the efficient combination of multimodal information by fusing the image feature vector and the text feature vector, and at the same time utilizes iterative denoising to optimize detail expression, improves the fidelity of the generated image sequence, and improves the dynamic consistency and detail clarity of the generated video in combination with the multi-layer feature restoration network of the decoding module, effectively improving the quality and applicability of the image-generated video. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:
[0031] Figure 1 A schematic diagram of an application environment of an image sequence construction method according to an embodiment of the present invention;
[0032] Figure 2 A schematic diagram of a flow chart of an embodiment of a method for constructing an image sequence according to the present invention;
[0033] Figure 3A schematic diagram of functional modules of a preferred embodiment of the image sequence construction device of the present invention;
[0034] Figure 4 A schematic diagram of the structure of a computer device in one embodiment of the present invention;
[0035] Figure 5 FIG. 4 is another schematic diagram of the structure of a computer device in one embodiment of the present invention. DETAILED DESCRIPTION
[0036] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0037] The image sequence construction method provided by the embodiment of the present invention can be applied in the following aspects: Figure 1 In an application environment, the user terminal communicates with the server terminal through a network. The server terminal can receive the initial image and text information through the user terminal, divide the initial image into multiple image segments, extract the image feature vector through the first network structure, extract the text feature vector through the second network structure, fuse the image feature vector, the text feature vector and the noise information to generate an initial fusion vector, perform denoising on the initial fusion vector, obtain the denoising result, and generate the target image sequence by decoding the denoising result. The present invention realizes the efficient combination of multimodal information through the fusion of image feature vectors and text feature vectors, and at the same time optimizes the detail performance by iterative denoising, improves the fidelity of the generated image sequence, and improves the dynamic consistency and detail clarity of the generated video in combination with the multi-layer feature restoration network of the decoding module, effectively improving the quality and applicability of the image-generated video. Among them, the user terminal can be but is not limited to various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server terminal can be implemented with an independent server or a server cluster composed of multiple servers. The present invention is described in detail below through specific embodiments.
[0038] See also Figure 2 , Figure 2 This is a flow chart of an embodiment of the image sequence construction method provided by the present invention. It should be noted that although the logical order is shown in the flow chart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0039] like Figure 2 As shown, the image sequence construction method proposed by the present invention includes the following steps:
[0040] S10, receiving initial image and text information;
[0041] In this embodiment, the system receives image data as the initial frame, supports multiple image formats (such as JPEG, PNG, TIFF, etc.), and normalizes it into a standard image format (such as RGB three-channel format) that is suitable for subsequent processing. An image processing library (such as OpenCV or Pillow) is used to read and convert the input image to ensure that the image resolution and color mode match the subsequent network structure.
[0042] Text information is an important input for guiding the generation of video content. It supports users to enter free text or select predefined template text through the keyboard, and performs basic legality checks on the input content (such as character encoding, length range, etc.). Text information is collected through the front-end input box or drop-down menu, converted into a standard encoding format (such as UTF-8), and at the same time detects whether the input character is empty or exceeds the length limit.
[0043] In addition, you can also automatically generate text descriptions from images using the following formula:
[0044] T = NLP_Model(I)
[0045] Where: T represents the generated text description (text vector representation); I represents the input image feature vector; NLP_Model is the mapping function of the natural language processing model (such as Transformer).
[0046] The image feature extraction module is used to convert the input image I into an image feature vector. Using the pre-trained Transformer model, a text vector T corresponding to the image content is generated through the attention mechanism. Finally, T is converted into human-readable text.
[0047] The image and text information need to be synchronously bound when transmitted to the backend to ensure that the corresponding relationship between the two is not disrupted in the subsequent feature extraction and fusion processing stages. During the transmission process, a unique identifier (such as UUID) is used to mark and bind the image and text, and the backend stores the corresponding relationship between the two through an associated database or memory cache.
[0048] Example description: In the field of medical health, users can upload CT scan images of patients as initial images and enter text information describing the location or change trend of lesions. The system receives and normalizes these data to provide support for the subsequent dynamic generation of videos showing the evolution of lesions, helping doctors to observe the changes in the disease more intuitively and improve diagnostic efficiency.
[0049] By receiving the initial image and text information, the diversity and integrity of the input data are ensured, providing a reliable input basis for the subsequent extraction of image features and text features, and improving the adaptability of data processing and the accuracy of subsequent generated results.
[0050] S20, dividing the initial image into a plurality of image segments;
[0051] In this embodiment, the initial image is divided into a plurality of image segments of equal size according to a fixed size, and each image segment represents a part of the original image. The size of the division is usually set to N×N (such as 16×16 or 32×32) to adapt to subsequent network processing. The image is cropped using an image processing tool (such as OpenCV), and the image segments are extracted block by block according to the specified size and stored in an array format.
[0052] The divided image segments are arranged from left to right and from top to bottom to form a sequence representation of the image segments. The rows and columns of the image are traversed using two-dimensional coordinates, image blocks are extracted according to the set cropping size, and stored in an array or matrix form to ensure that the order corresponds to the spatial layout.
[0053] For images that do not meet the fixed-size division requirements (e.g., the length and width are not integer multiples of N×N), the image edges need to be padded to make their size meet the division requirements. Use methods such as zero padding or mirror padding to expand the edge area of the image so that the image width and height are both integer multiples of the division size.
[0054] Normalize each image segment to adjust the distribution of pixel values to make it suitable for subsequent network structure processing requirements. The normalization formula is:
[0055]
[0056] in: is the standardized image segment; I i is the original image segment; μ and σ are the mean and standard deviation of the pixels in the image segment, respectively.
[0057] Example: In the field of healthcare, the initial image can be a medical image of a patient (such as a CT image, MRI image, or ultrasound image). By dividing these medical images into multiple fixed-size image segments, the key areas in the image can be decomposed into standardized input units.
[0058] Lesion detection: Large images from CT scans can be divided into small segments, and subsequent models can analyze each segment piece by piece to see if it contains lesion features.
[0059] Multimodal data fusion: By combining with text descriptions in the patient's medical records (such as diagnosis reports, medical history summaries), image fragments can be used as input features to generate video sequences that dynamically display changes in lesions.
[0060] Image visualization assistance: The segmented segments are reorganized into dynamic sequences to visualize the progression of the disease, such as generating dynamic videos of tumor volume changes.
[0061] For example, after a CT image of size 512×512 is input into the system, it is divided into 256 image segments of size 32×32. Each segment is standardized and labeled, combined with a text description (such as "the tumor is located in the lower left lung"), and is used in subsequent steps to generate a dynamic simulation video of the lesion changes.
[0062] By dividing the initial image into multiple image segments, a standardized input unit is provided for the subsequent feature extraction step. At the same time, the spatial structure information of the image is retained, the model's ability to capture detailed features is enhanced, and it can adapt to image input requirements of different resolutions.
[0063] S30, processing the multiple image segments through a first network structure to obtain an image feature vector;
[0064] In this embodiment, the divided multiple image segments are input into a first network structure (such as ViT-Encoder or CNN) to perform layer-by-layer feature extraction. Each image segment is encoded as an embedding vector x i , processed by convolutional layers or linear mapping layers in the network to extract low-level visual features (such as edges, textures).
[0065] The first network structure consists of multiple feature extraction modules, which extract global and local features of image segments layer by layer and build spatial associations between image segments. The global dependency between image segments is modeled using the self-attention mechanism:
[0066] X=LN(X+Attention(Q,K,V)),X=LN(X+FFN(X))
[0067] Where: Q, K, V are query, key and value matrices obtained by linear transformation. FFN is a feed-forward network used to further extract deep features.
[0068] For the input feature X (a feature vector with dimension d), the calculation formula of LN is as follows:
[0069]
[0070] Where: X represents the input feature that needs to be normalized;
[0071] represents the feature mean;
[0072] represents the characteristic variance;
[0073] ∈ represents a small constant, which is used to prevent the denominator from being zero;
[0074] γ represents a learnable scaling parameter (used to adjust the normalized feature distribution);
[0075] β represents a learnable offset parameter (used to adjust the range of normalized feature values).
[0076] The final result is a normalized feature with a mean of β and a variance controlled by γ.
[0077] The extracted image segment features are summarized through a feature integration module (such as feature pooling or weighted summation) to generate an image feature vector. The segment features are compressed into a unified vector representation using global average pooling (GAP) or weighted fusion operations.
[0078] Example description:
[0079] In the field of healthcare, for example, analyzing tumors in CT images: the initial CT image is divided into multiple 32×32 image segments.
[0080] The first network structure (such as ViT) processes the image fragments: preliminarily extracts local features such as tumor edges and density changes in the fragments; builds associations between fragments through a global attention mechanism to identify the overall morphology of the tumor area.
[0081] Finally, an image feature vector containing tumor region features is output for subsequent fusion analysis with diagnostic text information to generate a dynamic lesion change video.
[0082] By extracting features from image fragments through the first network structure, local and global image features can be effectively captured, improving the model's ability to understand image spatial dependencies and semantic information, and laying the foundation for the fusion of image and text information in subsequent steps.
[0083] S40, processing the text information through a second network structure to obtain a text feature vector;
[0084] In this embodiment, the input text information is segmented, the text is divided into multiple independent words or phrases, and a part-of-speech tag is added to each word or phrase. The text is segmented using a natural language processing tool. The segmentation result retains the position information of each word or phrase and its order in the original text. Part-of-speech tagging is completed using a pre-trained language model (such as BERT or SpaCy).
[0085] The words or phrases after word segmentation are converted into high-dimensional embedding vectors through the embedding module. Use the embedding layer or pre-trained embedding model (such as Word2Vec, GloVe or BERT) to map each word to a vector space of fixed dimension. The embedding vector retains the semantic information and contextual relationship of the word.
[0086] The word embedding vector is processed through the context feature extraction module to capture the semantic dependencies between different words or phrases. The Multi-Head Attention mechanism is used to calculate the correlation between words.
[0087] Attention mechanism formula:
[0088]
[0089] Where: Q, K, V represent query, key and value matrices respectively, which are obtained by linear transformation of word embedding vectors; d k Indicates the dimension of the key vector, used for normalization.
[0090] The context features are normalized and processed by the feedforward network to generate the global semantic feature representation of the text information. The context features are transformed nonlinearly using the feedforward neural network (FFN):
[0091] X=LN(X+FFN(X))
[0092] Use normalization (such as LayerNormalization, LN) to stabilize feature distribution.
[0093] Example: In the healthcare field, text information can be a patient's diagnosis report or medical record summary:
[0094] Example of text content: The patient described "mild inflammation of the right lung accompanied by coughing symptoms."
[0095] The text is segmented into "right lung", "appearance", "mild inflammation", "accompanied by", and "cough symptoms". A pre-trained language model (such as BERT) is used to generate an embedding vector for each word. The semantic dependencies between words are calculated through the attention mechanism to extract the key points in the text (such as "right lung", "mild inflammation"). Finally, a global semantic feature vector of the text is generated to guide the generation of dynamically changing videos of medical images (such as visualization of the development process of inflammation).
[0096] By processing text information through the second network structure, the semantic dependencies of words or phrases in the text can be effectively captured and global semantic feature representation can be generated. This not only enhances the model's ability to understand text information, but also provides high-quality semantic representation for subsequent steps of fusion with image features.
[0097] S50, merging the image feature vector, the text feature vector and the noise information to generate an initial fusion vector;
[0098] In this embodiment, the image feature vectors and text feature vectors are standardized to ensure that the feature vectors of different modalities have consistent distribution characteristics to facilitate subsequent fusion.
[0099] Use the normalization formula:
[0100]
[0101] Where: μ is the mean of the eigenvector, σ 2 is the variance. ∈ is a small constant to prevent the denominator from being zero.
[0102] The normalized image feature vector and text feature vector both have zero mean and unit variance.
[0103] Generate noise information that follows a standard normal distribution. The dimension of the noise is consistent with the image feature vector and the text feature vector. Generate noise by random sampling:
[0104] Z noise ~N(0,1)
[0105] Where Z noise is the noise vector.
[0106] Based on the statistical characteristics of image features, text features and noise (such as variance, information entropy), weighting factors are generated to control the fusion ratio between modalities.
[0107] The weighting factor is calculated using the following formula:
[0108]
[0109] Where: Var(X i ) is the variance of mode i; H(X i ) is the information entropy of mode i. Weighting factor w i Indicates the weight ratio of each modality in the fusion.
[0110] The image feature vector, text feature vector and noise information are concatenated in a weighted manner to generate a merged vector.
[0111] Weighted splicing formula:
[0112] X merged =w img ·X img +w text ·X text +w noise ·Z noise
[0113] Where: X img represents the image feature vector; X text Represents the text feature vector; Z noise is the noise information; w img , w text , w noise Represents the corresponding weighting factor.
[0114] Perform linear transformation on the merged vector to further extract fusion features and generate the initial fusion vector. Linear transformation formula:
[0115] X fusion =W fusion ·X merged
[0116] Where: W fusion is the fusion weight matrix.
[0117] By standardizing, weighted fusion and linearly transforming image feature vectors, text feature vectors and noise information, we can not only effectively balance the importance of multimodal features, but also generate an initial fusion vector of unified representation, improve the accuracy and robustness of multimodal information fusion, and provide high-quality input for subsequent denoising and decoding.
[0118] S60, performing denoising processing on the initial fusion vector to obtain a denoising result;
[0119] In this embodiment, the initial fusion vector is input into a denoising module to extract noise information and feature information from the initial fusion vector.
[0120] The initial fusion vector is used as a high-dimensional input tensor Z0, which contains mixed information of image, text and noise modalities. A feature extraction unit (such as a convolutional neural network or an encoder structure) is used to perform a preliminary analysis on Z0 to extract noise components and effective features.
[0121] In the denoising module, multiple iterations of denoising are performed on the initial fusion vector, and each iteration optimizes the feature information and reduces the noise information. The inverse diffusion process in the diffusion model is used to iterate based on the following formula:
[0122]
[0123] Where: Z t is the feature vector of the current time step; ∈ is the predicted noise information; βt , α t are predefined parameters in the diffusion process.
[0124] According to the changes in feature information and noise information during each iteration, the parameters of the denoising module are dynamically adjusted. By analyzing the residual noise after each iteration, the learning rate, activation function or convolution kernel size in the denoising network (such as the U-Net model) is adjusted.
[0125] Dynamic adjustment formula:
[0126]
[0127] Where: Param t is the current parameter of the denoising module; η is the learning rate; is the gradient of the current iteration.
[0128] In each iteration, the features are optimized while gradually reducing the noise component. Use the loss function to optimize feature generation:
[0129]
[0130] Where: L ∈ Represents the loss function in the denoising process, which is used to measure the real noise (∈) and the model prediction noise (∈ pred The smaller the value, the closer the noise predicted by the model is to the real noise, and the better the denoising effect is.
[0131] E represents the mathematical expectation, which is the average value of all samples. The loss function calculates the average error in the entire training set or batch to ensure that the model has global optimization performance.
[0132] || || 2 represents the square of the Euclidean norm, which is the sum of the squares of the vector.
[0133] ∈ represents the real noise, ∈ pred To predict the noise.
[0134] After completing multiple iterations of denoising, the denoised feature vector is output, which contains the pure feature information of the image mode and the text mode. clean The effective features in the initial fusion vector are retained, most of the noise components are eliminated, and it has the characteristics of high signal-to-noise ratio and multimodal information fusion.
[0135] By performing multiple iterative denoising processes on the initial fusion vector, the key information in the image and text modalities can be effectively retained, while redundant noise is eliminated, and the signal-to-noise ratio and expression ability of the feature vector are improved. This denoising process provides high-quality input for the subsequent decoding stage, significantly enhancing the generation effect of the target image sequence.
[0136] S70, performing decoding processing on the denoising result to generate a target image sequence.
[0137] In this embodiment, the denoising result output by the denoising module is input into the decoding module for further processing. The denoising result is represented in the form of a high-dimensional feature vector, which contains information of image and text modalities. The denoising result is received through the input interface of the decoding module, and the multimodal features are prepared for decoding.
[0138] In the decoding module, the correlation between the multimodal features in the denoising results is analyzed through the attention mechanism to generate multimodal joint features. The relationship between the image modality features and the text modality features is calculated using a cross-modal attention mechanism (such as Cross-Attention):
[0139] Z0=Z clean +LN(CrossAttn(Z clean , X T , X T ))
[0140] Where: Z clean Represents the image modality features in the denoising result. T Represents the text modality feature in the denoising result. CrossAttn represents the cross-modal attention mechanism, which is used to capture the interactive relationship between modalities. Output multimodal joint feature Z0, which enhances the collaborative information between modalities.
[0141] Normalize the multimodal joint features to make the feature distribution consistent. Use layer normalization (LN) to process the joint features:
[0142] Z normalized =LN(Z0)
[0143] The normalization process involves subtracting the mean and dividing by the standard deviation to ensure that the joint features have a stable distribution.
[0144] The normalized multimodal joint features are decoded step by step through the layer-by-layer feature restoration network in the decoding module. The decoding module uses a layer-by-layer feature restoration network (such as a multi-layer perceptron, MLP) to gradually restore the joint features:
[0145] X out =MLP(Z normalized +LN(FFN(Z normalized )))
[0146] Among them: MLP is a multi-layer perceptron, which is used to decode features. FFN is a feed-forward network, which is used to enhance the expressiveness of features.
[0147] The decoded multimodal joint features are converted into pixel values to generate a preliminary result of the image sequence. The decoded features are converted into pixel values through feature mapping operations to generate a preliminary image frame sequence. The pixel value range is usually limited to [0,1] or [0,255] to conform to the conventional image format.
[0148] Perform quality enhancement processing on the preliminary image frame sequence to generate a final image sequence that meets the target resolution requirements and target pixel value range requirements. Use a super-resolution module (such as SRGAN or ESRGAN) to increase the resolution of the preliminary image sequence. Adjust the pixel value range and distribution characteristics to ensure that the output image sequence meets the expected quality.
[0149] Example: In healthcare, decoding processing can be used to generate dynamically changing videos of patient imaging data:
[0150] Input: preliminary diagnostic images of the patient (such as CT or MRI) and related text descriptions (such as "the lesion area has increased").
[0151] Decoding: Extract dynamic change information of the lesion area from the denoising results. Generate lesion evolution video frames through the decoding module to show the dynamic changes of the lesion area over time.
[0152] Output: Dynamic video shows the gradual spread of the lesion area, providing clinicians with intuitive support for disease analysis.
[0153] Through the above steps, multimodal information can be efficiently extracted from the denoising results, and high-quality image sequences that meet the expected resolution and pixel value range requirements can be generated through joint feature enhancement and layer-by-layer feature restoration. This not only improves the quality and detail of the generated images, but also enhances the collaborative consistency of multimodal information in the generated video.
[0154] The present invention relates to the fields of artificial intelligence technology and medical health, and discloses a method for constructing an image sequence, including: receiving an initial image and text information, dividing the initial image into multiple image segments, extracting an image feature vector through a first network structure, extracting a text feature vector through a second network structure, fusing the image feature vector, the text feature vector and noise information to generate an initial fusion vector, performing denoising on the initial fusion vector to obtain a denoising result, and generating a target image sequence by decoding the denoising result. The present invention realizes the efficient combination of multimodal information by fusing the image feature vector and the text feature vector, and at the same time optimizes the detail expression by iterative denoising, improves the fidelity of the generated image sequence, and improves the dynamic consistency and detail clarity of the generated video by combining the multi-layer feature restoration network of the decoding module, effectively improving the quality and applicability of the image-generated video.
[0155] In one embodiment, the above S30 includes:
[0156] S301, performing normalization processing on the multiple image segments, standardizing pixel values of the multiple image segments, and adjusting the multiple image segments into input units of a fixed size;
[0157] S302, inputting the normalized multiple image segments into the embedding module of the first network structure, processing each image segment through linear mapping, generating a corresponding image embedding vector, and adding position information based on the position of each image segment to the image embedding vector;
[0158] S303, extracting image features including global dependencies and local dependencies of the multiple image segments from the image embedding vector through the multi-layer feature extraction module of the first network structure;
[0159] S304: Aggregate the image features of the plurality of image segments through the integration module of the first network structure to generate the image feature vector.
[0160] In this embodiment, the pixel values of each image segment are standardized so that the image segments are consistent at the same scale. The pixel value range is normalized to ensure that the pixel values are distributed in a set interval (such as [0,1]). All image segments are processed using a normalization algorithm: the mean and standard deviation of each image segment are calculated; the mean is subtracted from each pixel value and divided by the standard deviation.
[0161] Resize image segments to a fixed size (e.g. 16×16) to fit the network input.
[0162] The normalized image fragments are input into the embedding module, and an embedding vector is generated through linear mapping. The embedding module performs a matrix transformation on the input image fragments, converting the two-dimensional image data into a one-dimensional vector of fixed dimension. During the linear mapping process, the embedding module applies a pre-trained weight matrix to complete the conversion. Position information based on the position of the image fragment is added to the generated embedding vector (Position Embedding) to enhance the model's perception of spatial position information.
[0163] Features containing global dependencies and local details are extracted from the embedded vector through a multi-layer feature extraction module. Using the multi-layer structure of the feature extraction module, each layer contains an attention mechanism and a feedforward network: the attention mechanism captures global dependencies; the feedforward network extracts local details of the image fragment.
[0164] The embedding vector is updated layer by layer, and finally a high-quality vector containing global and local features is output.
[0165] The feature vectors of all image segments are aggregated through the integration module to generate a unified image feature vector. Multiple feature vectors are integrated using feature aggregation methods (such as average pooling or weighted pooling). The integrated feature vector maintains the global semantic and spatial information of the image segment, providing a unified input for subsequent steps.
[0166] In the financial field, it can be applied to the processing of market trend charts or heat maps. The input market trend chart is divided into multiple small areas, each of which corresponds to an image segment. By normalizing these segments, the pixel value differences between regions can be effectively eliminated, and the global dependencies and local change trends between regions can be extracted. Finally, the features of each segment are integrated to generate a unified image feature vector, which is used to predict market changes or generate a graphical trend report.
[0167] Through the above steps, this embodiment can efficiently extract global and local features in image segments, generate a unified image feature vector, and provide high-quality input for subsequent multimodal fusion. It solves the problem of inconsistent expression of image features and improves the global relevance and local accuracy of features.
[0168] In one embodiment, the above S40 includes:
[0169] S401, performing word segmentation processing on the text information, dividing the text information into multiple independent words or phrases, and adding a part-of-speech tag to each word or phrase;
[0170] S402, inputting the words or phrases after the word segmentation processing into the embedding module of the second network structure, and generating a corresponding text embedding vector for each word or phrase through an embedding operation;
[0171] S403, inputting the text embedding vector into the context feature extraction module of the second network structure, extracting the semantic dependency between different words or phrases in the text information through a multi-head attention mechanism, and generating context semantic features;
[0172] S404: Process the contextual semantic features through the feedforward network in the second network structure to generate a text feature vector containing a global semantic representation of the text information.
[0173] In this embodiment, the input text information is divided into multiple independent words or phrases, and a part-of-speech tag is added to each word or phrase. The text is segmented using a word segmentation algorithm (such as a dictionary-based word segmentation or a machine learning word segmentation). A part-of-speech tag (such as a noun, a verb, etc.) is added to each word or phrase after segmentation to enrich the semantic information. The output result is a text list containing the word segmentation results and the part-of-speech tag.
[0174] The segmented words or phrases are input into the embedding module of the second network structure, and the corresponding text embedding vector is generated through the embedding operation. Each word or phrase is mapped to a vector of fixed dimension using a pre-trained word embedding model (such as Word2Vec, GloVe, or Transformer-based model). The embedding module supports dynamic word embedding updates to capture context-related semantic changes.
[0175] The generated text embedding vector is input into the context feature extraction module, and the semantic dependency between different words or phrases is extracted through the multi-head attention mechanism. The multi-head attention mechanism captures the semantic relationship between words by calculating the weighted scores of Query, Key and Value. The output result is a text feature vector containing context information, reflecting the semantic relevance of each word in the sentence.
[0176] The contextual semantic features are processed by the feedforward network to generate a text feature vector containing global semantic representation. The contextual semantic features are transformed nonlinearly by the feedforward network to extract global semantic features. The feedforward network enhances the expressiveness of semantic features through activation functions (such as ReLU) and finally generates a text feature vector of global semantic representation.
[0177] This embodiment can efficiently capture the semantic relationship in text information through word segmentation, text embedding generation, context feature extraction and global semantic representation generation, and provide high-quality text feature vectors for subsequent multimodal information fusion. At the same time, it improves the context relevance of text features and enhances the richness and consistency of semantic expression.
[0178] In one embodiment, the above S50 includes:
[0179] S501, performing standardization processing on the image feature vector and the text feature vector, and adjusting the image feature vector and the text feature vector to vector forms with the same distribution characteristics;
[0180] S502, generating noise information with the same dimension as the image feature vector and the text feature vector, wherein the noise information obeys a standard normal distribution;
[0181] S503, merging the standardized image feature vector, the text feature vector and the noise information in a weighted concatenation manner to generate a merged vector;
[0182] S504: Perform a linear transformation operation on the merged vector to generate the initial fusion vector.
[0183] In this embodiment, the image feature vector and the text feature vector are standardized and adjusted to vector forms with the same distribution characteristics. The mean and standard deviation of each vector are calculated. The element values of the vectors are mapped to the same distribution range using a standardization method to ensure that the different modal features have a consistent feature distribution before fusion.
[0184] Generate a noise information that follows a standard normal distribution, whose dimension is consistent with the image feature vector and the text feature vector. Generate noise data randomly using a standard normal distribution. The generation of noise data takes into account the specific dimension of each modality feature to ensure that the impact of noise on the overall fusion feature is uniform and stable.
[0185] The normalized image feature vector, text feature vector and noise information are combined by weighted concatenation to generate a combined vector. The weight factor of each modality feature and noise is set to control the contribution ratio of each modality to the fusion vector. The weighted feature vectors are concatenated along the specified dimension to form a combined vector.
[0186] Apply linear transformation to the merged vector to generate the initial fused vector. Adjust the dimension and feature distribution of the merged vector through matrix transformation. Use the trained transformation weights and bias values to perform linear transformation on the vector to optimize the expressiveness of the fused features.
[0187] This embodiment can effectively fuse multimodal features and generate an initial fusion vector with stronger expressive power through standardization, noise generation, weighted concatenation and linear transformation. It solves the problem of inconsistent distribution of multimodal features and improves the diversity and stability of generated features by introducing noise information.
[0188] In one embodiment, the above S503 includes:
[0189] S5031, determining the variance and information entropy of the image modality feature based on the distribution statistics of the image feature vector;
[0190] S5032, determining the variance and information entropy of the text modal feature based on the distribution statistics of the text feature vector;
[0191] S5033, determining an influencing factor of a noise mode based on the amplitude and distribution characteristics of the noise information;
[0192] S5034, generating a weighting factor based on the variance and information entropy of the image modality feature, the variance and information entropy of the text modality feature, and the influence factor of the noise modality;
[0193] S5035: Perform weighted merging of the image feature vector, the text feature vector and the noise information according to the weighting factor to generate a merged vector.
[0194] In this embodiment, the distribution characteristics of the image feature vector are analyzed, and its variance (reflecting the discrete degree of the data) and information entropy (quantifying the uncertainty of the information) are calculated. The variance of the feature vector is calculated using statistical methods:
[0195]
[0196] Among them, x i is the eigenvalue and μ is the mean.
[0197] Calculate information entropy:
[0198]
[0199] Among them, P(x i ) is the probability distribution of the eigenvalues.
[0200] The same method as the image modality is used to analyze the distribution characteristics of the text feature vector and calculate its variance and information entropy. The distribution of the text feature vector is derived from the context information, and special attention should be paid to whether there is a semantic deviation in the distribution. The standard formula is used to perform statistical calculations on the feature values to ensure consistency with the metric standard of the image modality features.
[0201] Evaluate the impact of noise data on the fusion results and generate impact factors by analyzing the amplitude and distribution characteristics of the noise information. The noise data follows a standard normal distribution, and its amplitude can be controlled by randomly generated parameters. The noise impact factor is calculated using the amplitude and distribution parameters to ensure that the contribution of the noise does not obscure the real features.
[0202] The statistical information of image features, text features and noise is integrated to generate weighting factors to control the contribution ratio of different modal features. The weight benchmark of each modality is set according to the variance and information entropy. The noise impact factor is combined with the weight benchmark to generate the final weighting factor.
[0203] The standardized image feature vector, text feature vector and noise information are merged according to the weighting factor. The weighting factor is applied to each modal feature vector for weighting operation, and the weighted feature vectors are concatenated into a merged vector.
[0204] This embodiment can achieve efficient fusion of multimodal features by analyzing the distribution statistics of image features, text features and noise information and reasonably allocating weighting factors, thereby improving the expressiveness of multimodal features and enhancing the robustness of fused features to noise.
[0205] In one embodiment, the above S60 includes:
[0206] S601, inputting the initial fusion vector into a denoising module, and extracting feature information and noise information from the initial fusion vector through a feature extraction unit in the denoising module;
[0207] S602, performing multiple iterative denoising processes in a denoising module, optimizing feature information in each iterative denoising process, and gradually reducing noise information;
[0208] S603, adjusting the parameter configuration in the denoising module according to the changes of the feature information and the noise information in each iterative denoising process;
[0209] S604, after completing multiple iterations of denoising processing, a denoising result is obtained.
[0210] In this embodiment, the initial fusion vector contains target feature information and noise information, and the denoising module needs to distinguish and extract these two parts. The feature extraction unit in the denoising module performs feature decomposition on the fusion vector through a convolution operation or an attention mechanism. Specific parameter settings are used to extract the target feature information and isolate the noise information to form feature components and noise components.
[0211] Iterative denoising is a step-by-step optimization process, where each iteration reduces noise information through noise estimation and removal mechanisms. In each iteration, the noise information is predicted by a denoising network (such as a U-Net-based denoising network) and subtracted from the fusion vector. At the same time, the target feature information is optimized to enhance its expressiveness.
[0212] Dynamically adjust the parameters of the denoising module to adapt the denoising strength according to the changes in features and noise. After each iteration, calculate the change in feature information and noise information, and update the learning rate or network weight of the denoising module. Use the trained parameter model to adjust the denoising weight distribution of each iteration.
[0213] After multiple iterations, the denoising module outputs the final denoising result, which contains the optimized feature information. In the last iteration, the residual noise is reduced as much as possible to within the preset tolerance range. The output denoising result has clear feature information and can be used in the subsequent decoding step.
[0214] This embodiment can effectively remove noise information in the initial fusion vector through feature extraction, multiple iterations of denoising and dynamic parameter adjustment, while retaining and optimizing target feature information. The denoising result provides high-quality feature input for subsequent decoding steps, enhancing the stability and consistency of the final generated content.
[0215] In one embodiment, the above S70 includes:
[0216] S701, inputting the denoising result into a decoding module;
[0217] S702, in the decoding module, analyzing the correlation between the multimodal features in the denoising result by using an attention mechanism to generate a multimodal joint feature;
[0218] S703, performing a normalization operation on the multimodal joint feature;
[0219] S704, decoding the normalized multimodal joint features step by step through the layer-by-layer feature restoration network in the decoding module;
[0220] S705, converting the decoded multimodal joint features into pixel values through a feature mapping operation to generate a preliminary result of the image sequence;
[0221] S706, performing quality enhancement processing on the preliminary result of the image sequence to generate a target image sequence that meets the target resolution requirement and the target pixel value range requirement.
[0222] In this embodiment, the decoding module receives the result after denoising as the initial input of the decoding process. The denoising result is directly mapped to the input vector of the decoding module, ensuring that the decoding module can start the decoding process with high-quality feature information.
[0223] The attention mechanism is used to analyze the correlation between different modal features (such as image modality and text modality) and extract multi-modal joint features. Using the multi-head attention mechanism, the weight distribution of each modal feature as Query, Key and Value is calculated to generate a joint feature vector. The joint feature vector reflects the interactive information between the modalities.
[0224] The generated multimodal joint features are normalized to ensure the stability of data distribution. The LayerNormalization method is applied to normalize each dimension in the joint features to reduce the impact of distribution inconsistency on the decoding process.
[0225] The decoding module reconstructs the multimodal joint features step by step through a layer-by-layer feature restoration network. The decoding network consists of multiple stacked decoding layers, and each layer gradually restores the details of the joint features. The output of each layer serves as the input of the next layer, gradually refining the feature information and enhancing the resolution and integrity of the decoding results.
[0226] The decoded joint features are mapped to the pixel values of the image sequence to form a preliminary decoding result. The feature mapping operation maps the information in the feature space to the pixel space through a fully connected layer or a convolutional layer to generate a preliminary image sequence.
[0227] The quality enhancement module is used to optimize the resolution and pixel value range of the preliminary image sequence. Super-resolution reconstruction methods or pixel value correction strategies are applied to improve the clarity and consistency of the image. The pixel value range of the image is adjusted to meet the preset requirements.
[0228] This embodiment can efficiently generate the target image sequence from the denoising results through step-by-step decoding and quality enhancement. The attention mechanism improves the interactive ability of multimodal features, the layer-by-layer feature restoration network optimizes the detail performance of the decoding process, and the quality enhancement process ensures that the generated image sequence meets the target resolution and pixel value range requirements, providing a guarantee for the generation of high-quality image sequences.
[0229] In one embodiment, an image sequence construction device is provided, and the image sequence construction device corresponds one-to-one to the image sequence construction method in the above embodiment. Figure 3 , Figure 3 The functional module diagram of a preferred embodiment of the image sequence construction device of the present invention is as follows: data receiving module 10, image preprocessing module 20, first feature extraction module 30, second feature extraction module 40, multimodal fusion module 50, denoising module 60 and decoding module 70. The functional modules are described in detail as follows:
[0230] The data receiving module 10 is used to receive the initial image and text information;
[0231] An image preprocessing module 20, configured to divide the initial image into a plurality of image segments;
[0232] A first feature extraction module 30, configured to process the plurality of image segments through a first network structure to obtain an image feature vector;
[0233] A second feature extraction module 40, used for processing the text information through a second network structure to obtain a text feature vector;
[0234] A multimodal fusion module 50, used to merge the image feature vector, the text feature vector and the noise information to generate an initial fusion vector;
[0235] A denoising module 60, configured to perform denoising processing on the initial fusion vector to obtain a denoising result;
[0236] The decoding module 70 is used to perform decoding processing on the denoising result to generate a target image sequence.
[0237] In one embodiment, the first feature extraction module 30 is specifically used to:
[0238] Performing normalization processing on the multiple image segments, standardizing pixel values of the multiple image segments, and adjusting the multiple image segments into input units of a fixed size;
[0239] Inputting the normalized multiple image segments into the embedding module of the first network structure, processing each image segment through linear mapping to generate a corresponding image embedding vector, and adding position information based on the position of each image segment to the image embedding vector;
[0240] Extracting image features including global dependencies and local dependencies of the plurality of image segments from the image embedding vector through a multi-layer feature extraction module of the first network structure;
[0241] The image features of the multiple image segments are aggregated and processed through the integration module of the first network structure to generate the image feature vector.
[0242] In one embodiment, the second feature extraction module 40 is specifically configured to:
[0243] Performing word segmentation processing on the text information, dividing the text information into multiple independent words or phrases, and adding a part-of-speech tag to each word or phrase;
[0244] Inputting the segmented words or phrases into the embedding module of the second network structure, and generating corresponding text embedding vectors for each word or phrase through embedding operation;
[0245] Inputting the text embedding vector into the context feature extraction module of the second network structure, extracting the semantic dependency between different words or phrases in the text information through a multi-head attention mechanism, and generating context semantic features;
[0246] The contextual semantic features are processed through a feedforward network in the second network structure to generate a text feature vector containing a global semantic representation of the text information.
[0247] In one embodiment, the multimodal fusion module 50 is specifically configured to:
[0248] Performing standardization processing on the image feature vector and the text feature vector to adjust the image feature vector and the text feature vector to vector forms with the same distribution characteristics;
[0249] Generate noise information having the same dimension as the image feature vector and the text feature vector, wherein the noise information obeys a standard normal distribution;
[0250] The standardized image feature vector, the text feature vector and the noise information are combined in a weighted concatenation manner to generate a combined vector;
[0251] A linear transformation operation is performed on the merged vector to generate the initial fusion vector.
[0252] In one embodiment, the multimodal fusion module 50 is specifically configured to:
[0253] Based on the distribution statistics of the image feature vector, the variance and information entropy of the image modality features are determined;
[0254] Based on the distribution statistics of text feature vectors, the variance and information entropy of text modal features are determined;
[0255] Determine the influencing factors of noise modes based on the amplitude and distribution characteristics of noise information;
[0256] Generate a weighting factor based on the variance and information entropy of the image modality feature, the variance and information entropy of the text modality feature, and an influence factor of the noise modality;
[0257] The image feature vector, the text feature vector and the noise information are weighted and combined according to the weighting factor to generate a combined vector.
[0258] In one embodiment, the denoising module 60 is specifically configured to:
[0259] Inputting the initial fusion vector into a denoising module, and extracting feature information and noise information from the initial fusion vector through a feature extraction unit in the denoising module;
[0260] In the denoising module, multiple iterative denoising processes are performed, and the feature information is optimized in each iterative denoising process, and the noise information is gradually reduced;
[0261] In each iterative denoising process, the parameter configuration in the denoising module is adjusted according to the changes in feature information and noise information;
[0262] After completing multiple iterations of denoising, the denoising result is obtained.
[0263] In one embodiment, the decoding module 70 is specifically configured to:
[0264] Inputting the denoising result into a decoding module;
[0265] In the decoding module, the correlation between the multimodal features in the denoising result is analyzed by an attention mechanism to generate a multimodal joint feature;
[0266] Performing a normalization operation on the multimodal joint features;
[0267] The normalized multimodal joint features are decoded step by step through a layer-by-layer feature restoration network in the decoding module;
[0268] The decoded multimodal joint features are converted into pixel values through feature mapping operations to generate preliminary results of the image sequence;
[0269] A quality enhancement process is performed on the preliminary results of the image sequence to generate a target image sequence that meets the target resolution requirements and the target pixel value range requirements.
[0270] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external user terminal through a network connection. When the computer program is executed by the processor, the functions or steps on the server side of an image sequence construction method are realized.
[0271] In one embodiment, a computer device is provided. The computer device may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps of a user side of an image sequence construction method
[0272] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the following steps are implemented:
[0273] receiving initial image and text information;
[0274] Dividing the initial image into a plurality of image segments;
[0275] Processing the plurality of image segments through a first network structure to obtain an image feature vector;
[0276] Processing the text information through a second network structure to obtain a text feature vector;
[0277] Merging the image feature vector, the text feature vector and the noise information to generate an initial fusion vector;
[0278] Performing denoising processing on the initial fusion vector to obtain a denoising result;
[0279] Decoding is performed on the denoising result to generate a target image sequence.
[0280] In one embodiment, a computer readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0281] receiving initial image and text information;
[0282] Dividing the initial image into a plurality of image segments;
[0283] Processing the plurality of image segments through a first network structure to obtain an image feature vector;
[0284] Processing the text information through a second network structure to obtain a text feature vector;
[0285] Merging the image feature vector, the text feature vector and the noise information to generate an initial fusion vector;
[0286] Performing denoising processing on the initial fusion vector to obtain a denoising result;
[0287] Decoding is performed on the denoising result to generate a target image sequence.
[0288] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can refer to the relevant descriptions on the server side and the user side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0289] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0290] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0291] It should be noted that if software tools or components other than those of the Company appear in the embodiments of the present application, they are only used for illustration and do not represent actual use. The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the above-mentioned embodiments, a person of ordinary skill in the art should understand that the technical solutions described in the above-mentioned embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents; and these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A method for constructing an image sequence, characterized in that: The following steps are involved: receiving initial image and text information; Dividing the initial image into a plurality of image segments; Processing the plurality of image segments through a first network structure to obtain an image feature vector; Processing the text information through a second network structure to obtain a text feature vector; Merging the image feature vector, the text feature vector and the noise information to generate an initial fusion vector; Performing denoising processing on the initial fusion vector to obtain a denoising result; Decoding is performed on the denoising result to generate a target image sequence.
2. The image sequence construction method according to claim 1, characterized in that: Processing the plurality of image segments by a first network structure to obtain an image feature vector includes: Performing normalization processing on the multiple image segments, standardizing pixel values of the multiple image segments, and adjusting the multiple image segments into input units of a fixed size; Inputting the normalized multiple image segments into the embedding module of the first network structure, processing each image segment through linear mapping to generate a corresponding image embedding vector, and adding position information based on the position of each image segment to the image embedding vector; Extracting image features including global dependencies and local dependencies of the plurality of image segments from the image embedding vector through a multi-layer feature extraction module of the first network structure; The image features of the multiple image segments are aggregated and processed through the integration module of the first network structure to generate the image feature vector.
3. The image sequence construction method according to claim 1, characterized in that: Processing the text information through the second network structure to obtain a text feature vector includes: Performing word segmentation processing on the text information, dividing the text information into multiple independent words or phrases, and adding a part-of-speech tag to each word or phrase; Inputting the segmented words or phrases into the embedding module of the second network structure, and generating corresponding text embedding vectors for each word or phrase through embedding operation; Inputting the text embedding vector into the context feature extraction module of the second network structure, extracting the semantic dependency between different words or phrases in the text information through a multi-head attention mechanism, and generating context semantic features; The contextual semantic features are processed through a feedforward network in the second network structure to generate a text feature vector containing a global semantic representation of the text information.
4. The image sequence construction method as claimed in claim 1, characterized in that: The image feature vector, the text feature vector and the noise information are combined to generate an initial fusion vector, including: Performing standardization processing on the image feature vector and the text feature vector to adjust the image feature vector and the text feature vector to vector forms with the same distribution characteristics; Generate noise information having the same dimension as the image feature vector and the text feature vector, wherein the noise information obeys a standard normal distribution; The standardized image feature vector, the text feature vector and the noise information are combined in a weighted concatenation manner to generate a combined vector; A linear transformation operation is performed on the merged vector to generate the initial fusion vector.
5. The image sequence construction method according to claim 4, characterized in that: The standardized image feature vector, the text feature vector and the noise information are combined in a weighted concatenation manner to generate a combined vector, including: Based on the distribution statistics of the image feature vector, the variance and information entropy of the image modality features are determined; Based on the distribution statistics of text feature vectors, the variance and information entropy of text modal features are determined; Determine the influencing factors of noise modes based on the amplitude and distribution characteristics of noise information; Generate a weighting factor based on the variance and information entropy of the image modality feature, the variance and information entropy of the text modality feature, and an influence factor of the noise modality; The image feature vector, the text feature vector and the noise information are weighted and combined according to the weighting factor to generate a combined vector.
6. The image sequence construction method according to claim 1, characterized in that: Performing denoising processing on the initial fusion vector to obtain a denoising result includes: Inputting the initial fusion vector into a denoising module, and extracting feature information and noise information from the initial fusion vector through a feature extraction unit in the denoising module; In the denoising module, multiple iterative denoising processes are performed, and the feature information is optimized in each iterative denoising process, and the noise information is gradually reduced; In each iterative denoising process, the parameter configuration in the denoising module is adjusted according to the changes in feature information and noise information; After completing multiple iterations of denoising, the denoising result is obtained.
7. The image sequence construction method according to claim 1, characterized in that: Performing decoding processing on the denoising result to generate a target image sequence, including: Inputting the denoising result into a decoding module; In the decoding module, the correlation between the multimodal features in the denoising result is analyzed by an attention mechanism to generate a multimodal joint feature; Performing a normalization operation on the multimodal joint features; The normalized multimodal joint features are decoded step by step through a layer-by-layer feature restoration network in the decoding module; The decoded multimodal joint features are converted into pixel values through feature mapping operations to generate preliminary results of the image sequence; A quality enhancement process is performed on the preliminary results of the image sequence to generate a target image sequence that meets the target resolution requirements and the target pixel value range requirements.
8. An image sequence construction device, characterized in that: The image sequence construction device comprises: A data receiving module, used for receiving initial image and text information; An image preprocessing module, used for dividing the initial image into a plurality of image segments; A first feature extraction module, configured to process the plurality of image segments through a first network structure to obtain an image feature vector; A second feature extraction module, used for processing the text information through a second network structure to obtain a text feature vector; A multimodal fusion module, used for merging the image feature vector, the text feature vector and the noise information to generate an initial fusion vector; A denoising module, used for performing denoising processing on the initial fusion vector to obtain a denoising result; The decoding module is used to perform decoding processing on the denoising result to generate a target image sequence.
9. A computer device, characterized in that: The computer device includes a memory, a processor, and an image sequence construction program stored in the memory and executable on the processor. When the image sequence construction program is executed by the processor, the steps of the image sequence construction method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: An image sequence construction program is stored on the storage medium, and when the image sequence construction program is executed by the processor, the steps of the image sequence construction method according to any one of claims 1 to 7 are implemented.