Image-text content matching generation system and method and storage medium
By combining the text recognition model of the HSFs algorithm and the CRNN model and the improved generative adversarial network model of the multi-model voting algorithm, the problem that the existing technology cannot accurately identify multimodal natural input information is solved, and the accurate understanding of user intentions and high-quality effect of matching graphic and text content is achieved.
Patent Information
- Application Number
- CN202510324357.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art cannot accurately identify the natural input information of multimodality and deeply extract complex features in the multimodal input, resulting in inaccurate understanding of user intentions.
The text recognition model is constructed and pre-trained by the HSFs algorithm combined with the CRNN model. The optimal matching set is determined through the improved generative adversarial network model of the multi-model voting algorithm, and typesetting and reconstruction is carried out through the graphic and text generator based on greedy strategies.
It realizes accurate identification and feature extraction of multimodal natural input information, accurately understands user intentions and preferences, and improves the quality and layout effect of graphic and text content matching.
Smart Images

Figure CN120220162A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of electrical data processing, and particularly relates to a graphic and text content matching and generating system, method, and storage medium. Background Art
[0002] With the wide application of multimedia content, the presentation mode combining graphics and text has become increasingly important in fields such as news reporting, social media, advertising, and education publishing. However, there are still many deficiencies in existing graphic and text matching and typesetting technologies. Traditional methods mainly rely on manual selection and typesetting, which are inefficient and error-prone. In recent years, with the development of multimodal technologies, graphic and text matching models based on deep learning have gradually emerged.
[0003] Chinese Patent CN110020411B discloses a graphic and text content generating method and device. The method can, in response to a user's text input operation, determine the text input by the user, display the text in a display interface, then determine keywords based on the text, obtain pictures matching the keywords, and generate and display in the display interface a graphic and text content containing the text and the pictures. However, the existing method determines candidate keywords through simple word segmentation and semantic analysis, and cannot accurately identify multimodal natural input information and deeply extract complex features in multimodal input, resulting in inaccurate understanding of user intentions. To address the above problems, we propose a graphic and text content matching and generating system, method, and storage medium. Summary of the Invention
[0004] The purpose of the present invention is to provide a graphic and text content matching and generating system, method, and storage medium for the deficiencies of the existing technology, and solve the problem that the existing method determines candidate keywords through simple word segmentation and semantic analysis, and cannot accurately identify multimodal natural input information and deeply extract complex features in multimodal input, resulting in inaccurate understanding of user intentions.
[0005] The present invention is implemented as follows. A graphic and text content matching and generating method, the graphic and text content matching and generating method includes:
[0006] Obtain natural input information, construct and pre-train a text recognition model using the HSFs algorithm combined with the CRNN model, and perform recognition and analysis on the natural input information based on the pre-trained text recognition model to obtain a text output vector;
[0007] In response to the text output vector, use the text output vector as an index to traverse a pre-constructed picture database, grab at least one set of similar picture sets that match the text output vector based on a preset graphic and text similarity threshold, and use a generative adversarial network model improved by a multi-model voting algorithm to determine an optimal matching set;
[0008] Load the text output vector and the optimal matching set, extract the text output vector and the optimal matching set features through the graphic and text generator, fuse the text output vector and the optimal matching set features, and the graphic and text generator performs typesetting reconstruction on the text output vector and the optimal matching set based on the greedy strategy, and outputs the typesetting generation result.
[0009] Preferably, the method for constructing and pre-training a text recognition model by using the HSFs algorithm in combination with the CRNN model specifically includes:
[0010] Taking the CRNN model as the initial model of the text recognition model, the initial model of the text recognition model includes an input layer, a convolutional layer, a recurrent layer, a transcription layer, and an output layer. The convolutional layer consists of three layers of convolution, and there are three layers in the recurrent layer. The convolutional layer, the recurrent layer, and the transcription layer form the CRNN architecture network;
[0011] Introduce the HSFs algorithm and the global attention layer into the CRNN architecture network, introduce a multi-layer perceptron MLP in front of the convolutional layer of the CRNN architecture network, replace the transcription layer with a BERT pre-training model, and introduce a Transformer encoder and three groups of BILSTM layers between the CRNN architecture network and the output layer to complete the construction of the text recognition model;
[0012] Use the sample retrieval regular expression in combination with the Scrapy crawler framework to crawl multi-modal data sources. Among them, the multi-modal data sources include text data, voice data, and video data. Perform data cleaning, word segmentation annotation, and decoding preprocessing on the multi-modal data sources, and divide the preprocessed multi-modal data sources into a training set and a test set;
[0013] Load the pre-constructed text recognition model, and set the activation function, loss function, hyperparameters, and parameter optimizer of the text recognition model;
[0014] Obtain the training set, and iteratively train the text recognition model with the training set. The text recognition model extracts training text features based on the training set and generates training output vectors based on the training text features;
[0015] Based on the loss function, judge the loss value between the training output vector and the training true vector, and adjust the hyperparameters of the text recognition model through the genetic algorithm, and output the converged text recognition model;
[0016] Obtain the test set, use the test set as the input, execute the text recognition model, output the test accuracy, and judge whether the test accuracy exceeds the preset accuracy threshold. If the test accuracy exceeds the preset accuracy threshold, output the converged text recognition model.
[0017] Preferably, when judging the loss value between the training output vector and the training true vector based on the loss function, the loss value calculation formula is as follows:
[0018] L all = L cr + α1L bi + α2L tr (1)
[0019] Among them, L all , L cr , L bi , L tr respectively represent the total loss of the text recognition model, the global attention loss, the text discrimination loss, and the encoding loss. α1 and α2 are the hyperparameters of the BILSTM layer and the encoder respectively;
[0020] The global attention loss is:
[0021]
[0022] Among them, E cr represents the data expectation of the global attention loss value, A(x), are the true value and the predicted value extracted from the training text features by the CRNN architecture network respectively;
[0023] The text discrimination loss is:
[0024]
[0025] Among them, represents the cosine similarity between the training real vector and the training output vector after the training text features are processed by the BILSTM layer, and φ represents the softmax temperature;
[0026] The encoding loss is:
[0027]
[0028] Among them, represents the IOU value of the Transformer encoder, C(i), are the true encoding label and the predicted encoding label after the training output vector is processed by the Transformer encoder respectively.
[0029] Preferably, the method for recognizing and analyzing natural input information based on the pre-trained text recognition model specifically includes:
[0030] Loading natural input information, the multi-layer perceptron MLP in the text recognition model discriminates the type of natural input information to determine whether the type of natural input information is text data;
[0031] If the type of natural input information is text data, the BERT pre-trained model in the CRNN architecture network is triggered. The BERT pre-trained model extracts word vectors, sentence vectors, and position vectors from the natural input information, and splices the word vectors, sentence vectors, and position vectors to obtain a text splicing sequence;
[0032] If the type of natural input information is non-text data, the Mel filter is used to filter the natural input information, and the semantic information in the natural input information is extracted and recognized based on the HSFs algorithm. The high-dimensional semantic information is continuously mapped to a low-dimensional feature space, and a semantic extraction sequence is output;
[0033] Load the text splicing sequence or the semantic extraction sequence. The convolutional layer in the CRNN architecture network performs convolutional fusion on the features of the text splicing sequence or the semantic extraction sequence to extract local features and obtain a local feature set;
[0034] Through the global attention layer, the local feature set is recognized, and the global features in the local feature set are extracted to obtain the global feature sequences of the splicing word features, sentence features, position features, emotional features, and preference features;
[0035] Obtain the global feature sequence. The BILSTM layer extracts the context semantic information of the global feature sequence and fuses the context semantic information as the extracted text output vector. The text output vector is sent to the Transformer encoder, and the Transformer encoder encodes and labels it based on the probability of the text output vector in the picture type to obtain the encoded and labeled text output vector.
[0036] Preferably, the method for determining the optimal matching set by the generative adversarial network model improved by the multi-model voting algorithm includes:
[0037] Load the similar picture set, traverse the picture entities, picture types, and picture attributes of the similar picture set, perform picture alignment based on the timestamp alignment algorithm, and perform a temporal relationship mapping on the picture entities. The generative adversarial network model improved by the multi-model voting algorithm extracts the picture entity features, and calculates the attribute correlation value between the picture entity and the text output vector based on the DBSCAN algorithm;
[0038] The picture entity features and the picture type features are weighted and combined into a grayscale feature vector, and the grayscale correlation value with the text output vector is calculated based on the principal component analysis method;
[0039] Load the attribute correlation value and the grayscale correlation value, and use the emotional features and preference features of the text output vector as constraints. The generative adversarial network model improved by the multi-model voting algorithm performs weighted summation on the attribute correlation value and the grayscale correlation value to calculate the comprehensive decoding value of the picture entity;
[0040] Using the encoded label value of the text output vector as the screening threshold, retain the top M image entities whose comprehensive decoding value is greater than the encoded label value, integrate the M image entities, and set them as the optimal matching set.
[0041] Preferably, the comprehensive decoding value of the image entity is calculated by the following formula:
[0042]
[0043] where p j (q j , r j ) is the comprehensive decoding value of the image entity, P is the number of image entities in the similar image set, W j is the image entity feature, x i is the text output vector, q j , r j are the attribute association value and the grayscale association value respectively, λ1 and λ2 are the sentiment feature and preference feature of the text output vector respectively, σ(·) is the activation function of the generative adversarial network model improved by the multi-model voting algorithm, B j represents the bias term of the generative adversarial network model improved by the multi-model voting algorithm, q0 represents the initial weight of the image entity, ε j is the distance threshold of the ε-neighborhood based on the DBSCAN algorithm, h j , h j represent the image type feature and the grayscale feature vector respectively.
[0044] Preferably, the method for the graphic generator to perform layout reconstruction on the text output vector and the optimal matching set based on the greedy strategy includes:
[0045] Load the text output vector and the optimal matching set, and the graphic generator performs initial layout on the text output vector based on the concatenated word feature, sentence feature, and position feature in the text output vector to obtain the initial layout result;
[0046] The graphic generator sorts the image entities in the optimal matching set in descending order according to the comprehensive decoding value, and uses the greedy strategy in the Kruskal algorithm to gradually select the optimal image entities, and imports the optimal image entities into the initial layout result to complete the improvement of the initial layout result;
[0047] Load the improved initial layout result, solve the minimum spanning tree problem of the text output vector and the optimal matching set based on the greedy strategy, select the minimum-weight spanning tree step by step to construct the optimal solution, and dynamically adjust the layout of the improved initial layout result to output the layout generation result.
[0048] On the other hand, the present invention also provides a graphic and text content matching generation system, and the graphic and text content matching generation system includes:
[0049] An information collection module, which is used to obtain natural input information, constructs and pre-trains a text recognition model by using the HSFs algorithm combined with the CRNN model, and performs recognition and analysis on the natural input information based on the pre-trained text recognition model to obtain a text output vector;
[0050] A picture indexing module, in response to the text output vector, uses the text output vector as an index to traverse a pre-constructed picture database, grabs at least one set of similar picture sets that match the text output vector based on a preset picture-text similarity threshold, and determines the optimal matching set by using a generative adversarial network model improved by a multi-model voting algorithm;
[0051] A layout generation module, which is used to load the text output vector and the optimal matching set, extracts the features of the text output vector and the optimal matching set through a picture-text generator, fuses the features of the text output vector and the optimal matching set, and the picture-text generator reconstructs the layout of the text output vector and the optimal matching set based on a greedy strategy, and outputs a layout generation result.
[0052] Preferably, the picture indexing module includes:
[0053] A picture database, which is used to store picture entities associated with the text output vector, transform the formats of the picture entities, and label the length, angle, shape, position, direction, area, volume, saturation, and hue of the picture entities;
[0054] A database indexing unit, in response to the text output vector, uses the text output vector as an index to traverse a pre-constructed picture database, and grabs at least one set of similar picture sets that match the text output vector based on a preset picture-text similarity threshold;
[0055] A picture-text matching unit, which is used to load the text output vector and determine the optimal matching set by using a generative adversarial network model improved by a multi-model voting algorithm.
[0056] On the other hand, the present invention also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the picture-text content matching and generating method as described above is implemented.
[0057] Compared with the prior art, the embodiments of the present application mainly have the following beneficial effects:
[0058] In the embodiments of the present invention, the HSFs algorithm is combined with the CRNN model to construct and pre-train a text recognition model. Based on the pre-trained text recognition model, the natural input information is recognized and analyzed, so as to ensure the accurate recognition of multi-modal natural input information while extracting the user's emotional features and preference features, ensuring the accurate understanding of the user's intentions and preferences, and combining the generative adversarial network model improved by the multi-model voting algorithm to determine the optimal matching set, which can capture the semantic association between text and image more comprehensively, optimize the layout effect of text and image, and thus generate a higher-quality typesetting generation result. It overcomes the problem that the existing method determines candidate keywords through simple word segmentation and semantic analysis, and cannot accurately recognize multi-modal natural input information and deeply extract complex features in multi-modal input, resulting in inaccurate understanding of the user's intentions.
[0059] In the embodiments of the present invention, a text recognition model using the HSFs algorithm combined with the CRNN model is provided. By introducing the HSFs algorithm, the BERT pre-trained model, the Transformer encoder, and the BILSTM layer, the text recognition model significantly improves the feature extraction ability and semantic understanding ability for multi-modal data sources such as text, audio, and video. By integrating multi-modal data sources and performing in-depth preprocessing, the model can better adapt to complex input scenarios, improve the recognition accuracy, and enable the finally output text recognition model to efficiently and accurately recognize the text information in multi-modal input, with wide applicability and strong adaptability.
[0060] In the embodiments of the present invention, when the text recognition model recognizes and analyzes natural input information, through the MLP type discrimination and the processing strategies for different types of input, the text recognition model can support various input forms such as text and speech. Through the BERT pre-trained model and the global attention layer, the text recognition model can extract semantic information, emotional features, and user preference features from the input data. Through the Transformer encoder, the probability of the text output vector on the picture type is used to encode the label, and the encoded label ensures a high degree of semantic consistency between the text and the picture, making the finally generated text and picture content more natural and coordinated.
[0061] In the embodiments of the present invention, by calculating the attribute association value through the DBSCAN algorithm, the generative adversarial network model can effectively quantify the semantic correlation between the picture and the text, improving the matching accuracy. And by calculating the gray correlation value through PCA, the generative adversarial network model can effectively quantify the visual similarity between the picture and the text, providing an important basis for the subsequent comprehensive decoding value calculation. Introducing emotional features and preference features as constraint conditions can generate matching results that better meet the user's needs and improve the user experience. Description of the Drawings
[0062] Figure 1 It is a schematic diagram of the implementation process of the method for generating graphic and text content matching provided by the present invention.
[0063] Figure 2 It shows a schematic diagram of the implementation process of the method for constructing and pre-training a text recognition model by using the HSFs algorithm in combination with the CRNN model.
[0064] Figure 3 It shows a schematic diagram of the implementation process of the method for recognizing and analyzing natural input information based on the pre-trained text recognition model.
[0065] Figure 4 It shows a schematic diagram of the implementation process of the method for determining the optimal matching set by using the multi-model voting algorithm to improve the generative adversarial network model.
[0066] Figure 5 It shows a schematic diagram of the implementation process of the method for the graphic and text generator to perform typesetting reconstruction on the text output vector and the optimal matching set based on the greedy strategy.
[0067] Figure 6 It shows a schematic diagram of the structure of the graphic and text content matching generation system. Detailed implementation manners
[0068] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification and claims of this application or the above drawings are used to distinguish different objects and not to describe a specific order.
[0069] Referring to "embodiment" herein means that the specific features, structures or characteristics described in connection with the embodiment can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0070] Existing methods determine candidate keywords through simple word segmentation and semantic analysis, and cannot accurately identify multimodal natural input information and deeply extract complex features in multimodal input, resulting in inaccurate understanding of user intentions. To address the above problems, we propose a graphic and text content matching generation system, method, and storage medium. Briefly, when the method is implemented, natural input information is first obtained, and a text recognition model is constructed and pre-trained using the HSFs algorithm combined with the CRNN model. Based on the pre-trained text recognition model, the natural input information is recognized and analyzed, and the text output vector is used as an index to traverse the pre-constructed picture database. The generative adversarial network model improved by the multi-model voting algorithm is used to determine the optimal matching set. Finally, the text output vector and the optimal matching set features are extracted by a graphic and text generator, and the text output vector and the optimal matching set features are fused. The graphic and text generator reconstructs the layout of the text output vector and the optimal matching set based on a greedy strategy, and outputs the layout generation result. In the embodiments of the present invention, a text recognition model is constructed and pre-trained using the HSFs algorithm combined with the CRNN model, and the natural input information is recognized and analyzed based on the pre-trained text recognition model, so as to ensure accurate recognition of multimodal natural input information while extracting the user's emotional features and preference features, ensuring accurate understanding of user intentions and preferences, and combining the generative adversarial network model improved by the multi-model voting algorithm to determine the optimal matching set, which can more comprehensively capture the semantic association between text and image, optimize the layout effect of graphics and text, and thus generate a higher-quality layout generation result. It overcomes the problem that existing methods determine candidate keywords through simple word segmentation and semantic analysis, cannot accurately identify multimodal natural input information and deeply extract complex features in multimodal input, resulting in inaccurate understanding of user intentions.
[0071] An embodiment of the present invention provides a graphic and text content matching generation method. Figure 1 The implementation process schematic diagram of the graphic and text content matching generation method is shown. The graphic and text content matching generation method specifically includes:
[0072] Step S10, obtain natural input information, construct and pre-train a text recognition model using the HSFs algorithm combined with the CRNN model, and recognize and analyze the natural input information based on the pre-trained text recognition model to obtain a text output vector;
[0073] It should be noted that the natural input information includes but is not limited to text data, voice data, and video data. When obtaining natural input information, text data collection devices (keyboards, touch screens, scanners, web crawlers), voice data collection terminals (microphones, headphones, speakers, voice recorders, smartphones, tablets), video collection devices (ordinary cameras, high-definition cameras, infrared cameras, smart TVs, intelligent security cameras), etc. can be used to collect natural input information.
[0074] Step S20: In response to the text output vector, using the text output vector as an index, traverse the pre-constructed picture database, grab at least one set of similar picture sets that match the text output vector based on a preset text-picture similarity threshold, and use a generative adversarial network model improved by a multi-model voting algorithm to determine the optimal matching set;
[0075] Step S30: Load the text output vector and the optimal matching set, extract the features of the text output vector and the optimal matching set through a text-picture generator, and fuse the features of the text output vector and the optimal matching set. The text-picture generator performs typesetting reconstruction on the text output vector and the optimal matching set based on a greedy strategy, and outputs a typesetting generation result.
[0076] In this embodiment, extracting the features of the text output vector and the optimal matching set through a text-picture generator and fusing the features of the text output vector and the optimal matching set refer to the process of the text-picture generator's understanding of text semantics and picture features. The text-picture generator can be a LaBSE model, a Fotor model, a PicMonkey model, or a GeekerX tool.
[0077] In the embodiment of the present invention, an HSFs algorithm is combined with a CRNN model to construct and pre-train a text recognition model. Based on the pre-trained text recognition model, natural input information is recognized and analyzed, so as to ensure accurate recognition of multi-modal natural input information while extracting the user's emotional features and preference features, ensure accurate understanding of the user's intentions and preferences, and combine a generative adversarial network model improved by a multi-model voting algorithm to determine the optimal matching set, which can more comprehensively capture the semantic association between text and images, optimize the layout effect of text and pictures, and thus generate a higher-quality typesetting generation result. It overcomes the problem that the existing method determines candidate keywords through simple word segmentation and semantic analysis, cannot accurately recognize multi-modal natural input information and deeply extract complex features in multi-modal input, resulting in inaccurate understanding of the user's intentions.
[0078] The embodiment of the present invention provides a method for constructing and pre-training a text recognition model by combining an HSFs algorithm with a CRNN model, Figure 2 The schematic diagram of the implementation process of the method for constructing and pre-training a text recognition model by combining an HSFs algorithm with a CRNN model is shown. The method for constructing and pre-training a text recognition model by combining an HSFs algorithm with a CRNN model specifically includes:
[0079] Step S101: Use the CRNN model as the initial model of the text recognition model. The initial model of the text recognition model includes an input layer, a convolutional layer, a recurrent layer, a transcription layer, and an output layer. The convolutional layer consists of three layers of convolution, and the recurrent layer is provided with three layers. The convolutional layer, the recurrent layer, and the transcription layer form a CRNN architecture network;
[0080] In the embodiments of the present invention, the convolution kernels of the three-layer convolution in the convolution layer have a size of 3×3, the number of nodes in the input layer is 3-8, the number of nodes in the output layer is 2, and the number of nodes in the recurrent layer can be gradually increased by the trial-and-error method. The same test set is repeatedly trained multiple times, and the most suitable number of nodes is selected according to the principle of minimum error.
[0081] Step S102, introduce the HSFs algorithm and the global attention layer into the CRNN architecture network, introduce a multi-layer perceptron MLP before the convolution layer of the CRNN architecture network, replace the transcription layer with the BERT pre-trained model, and introduce a Transformer encoder and three groups of BILSTM layers between the CRNN architecture network and the output layer to complete the construction of the text recognition model;
[0082] It should be noted that the HSFs (Hierarchical Semantic Features) algorithm can extract hierarchical semantic features in speech data and video data. Compared with traditional convolution feature extraction methods, HSFs can more effectively capture the semantic information of speech data and video data, thereby improving the model's recognition ability for complex speech data and video data. The powerful language model ability of BERT can provide richer semantic understanding for text recognition. Replacing the traditional transcription layer with BERT can significantly improve the model's ability to generate text sequences. Especially when dealing with multi-modal inputs, it can better understand the context information of the text. Introducing a Transformer encoder and three groups of BILSTM layers between the CRNN architecture and the output layer further enhances the model's ability to model long text sequences. The self-attention mechanism of Transformer can capture long-distance dependencies in the text, while BILSTM can better process the bidirectional information of the sequence. This combination significantly improves the performance of the model.
[0083] Step S103, use the sample retrieval regular expression combined with the Scrapy crawler framework to crawl multi-modal data sources, where the multi-modal data sources include text data, speech data, and video data. Perform data cleaning, word segmentation annotation, and decoding preprocessing on the multi-modal data sources, and divide the preprocessed multi-modal data sources into a training set and a test set;
[0084] In this embodiment, various types of text information can be obtained through web crawlers, such as news articles, social media posts, forum discussions, etc. These text data can provide rich language features and semantic information. Speech data obtained from sources such as speech recognition systems, podcasts, and audio files can support natural language processing and speech recognition tasks. By crawling video data from video websites, a large amount of visual information can be obtained, which is very important for tasks such as image recognition and video classification. And through regular expressions and other data processing techniques, the crawled data can be cleaned to remove noise, duplicate data, and irrelevant information, thereby improving the quality of the data. Perform unified format conversion and standardization processing on data from different sources to make it suitable for subsequent analysis and modeling.
[0085] Step S104, load the pre-built text recognition model, and set the activation function, loss function, hyperparameters, and parameter optimizer of the text recognition model;
[0086] In this embodiment, the learning rate hyperparameter can be 0.02 - 0.04, the parameter optimizer can be the Adam optimizer, the loss function can be the cross-entropy loss function, absolute value loss function, mean squared error loss function, and the activation function can be a non-linear activation function, Maxout function, Sigmoid function, Tanh function.
[0087] Step S105, obtain the training set, and iteratively train the text recognition model using the training set. The text recognition model extracts training text features based on the training set and generates training output vectors based on the training text features;
[0088] In the embodiment of the present invention, through iterative training, the model can continuously learn and adjust its own parameters, gradually improving its ability to recognize complex texts. This dynamic learning process ensures that the model can continuously optimize its performance during training.
[0089] Step S106, based on the loss function, determine the loss value between the training output vector and the training true vector, and adjust the hyperparameters of the text recognition model through the genetic algorithm. Output the converged text recognition model. By adjusting the hyperparameters of the text recognition model through the genetic algorithm, the inefficiency of traditional optimization methods (such as grid search) can be effectively avoided. The genetic algorithm can quickly search for the optimal combination of hyperparameters by simulating the natural selection process, improving the performance of the model;
[0090] In the embodiment of the present invention, when determining the loss value between the training output vector and the training true vector based on the loss function, the loss value calculation formula is as follows:
[0091] L all = L cr + α1L bi + α2Ltr (1)
[0092] Among them, L all , L cr , L bi , L tr respectively represent the total loss of the text recognition model, the global attention loss, the text discrimination loss, and the encoding loss. α1 and α2 are the hyperparameters of the BILSTM layer and the encoder respectively. In this embodiment, the hyperparameters of the BILSTM layer and the encoder can be 0.01 - 0.04;
[0093] Traditional loss functions (such as cross-entropy loss, mean square error, etc.) usually only focus on the error calculation of a single dimension. In contrast, the total loss of the text recognition model can measure the performance of the model from multiple perspectives by combining the global attention loss, the text discrimination loss, and the encoding loss. This multi-dimensional loss calculation method can more comprehensively reflect the performance of the model on different tasks, improving the adaptability and robustness of the model to complex data. The introduction of the encoding loss enables the model to better optimize the feature encoding process, ensuring that the feature vectors output by the encoder have stronger expressiveness and distinguishability. By combining the encoding loss, the model can more efficiently convert the input data into a compact and information-rich feature representation, thereby improving the overall performance.
[0094] The global attention loss is:
[0095]
[0096] Among them, E cr represents the data expectation of the global attention loss value, A(x), are the true value and the predicted value extracted from the training text features by the CRNN architecture network respectively;
[0097] The text discrimination loss is:
[0098]
[0099] Among them, represents the cosine similarity between the training real vector and the training output vector after the training text features are processed by the BILSTM layer, φ represents the softmax temperature, and i and N represent the training samples and the total number of samples respectively;
[0100] The encoding loss is:
[0101]
[0102] Among them, represents the IOU value of the Transformer encoder, C(i), respectively represent the true encoded label and the predicted encoded label after the training output vector is processed by the Transformer encoder.
[0103] Step S107: Obtain a test set. Using the test set as the input, execute the text recognition model and output the test accuracy. By comparing with a preset accuracy threshold, it can be determined whether the model meets the requirements of actual applications. This threshold setting provides a clear goal for optimizing the model, ensuring that the model can continuously improve its performance during the training process. In this embodiment, the accuracy threshold can be 0.85 - 0.95.
[0104] Step S108: Determine whether the test accuracy exceeds the preset accuracy threshold.
[0105] Step S109: If the test accuracy exceeds the preset accuracy threshold, output the converged text recognition model.
[0106] If the test accuracy does not exceed the preset accuracy threshold, return to step S105 and continue to iteratively train the text recognition model using the training set.
[0107] In the embodiment of the present invention, a text recognition model that combines the HSFs algorithm with the CRNN model is provided. By introducing the HSFs algorithm, the BERT pre-training model, the Transformer encoder, and the BILSTM layer, the text recognition model significantly improves its feature extraction ability and semantic understanding ability for multi-modal data sources such as text, audio, and video. By integrating multi-modal data sources and performing in-depth preprocessing, the model can better adapt to complex input scenarios, improve the recognition accuracy, and enable the finally output text recognition model to efficiently and accurately recognize text information in multi-modal inputs, with wide applicability and strong adaptability.
[0108] The embodiment of the present invention provides a method for recognizing and analyzing natural input information based on a pre-trained text recognition model. Figure 3 The schematic diagram of the implementation process of the method for recognizing and analyzing natural input information based on a pre-trained text recognition model is shown. The method for recognizing and analyzing natural input information based on a pre-trained text recognition model specifically includes:
[0109] Step S201: Load natural input information, and the multi-layer perceptron (MLP) in the text recognition model discriminates the type of natural input information.
[0110] In the embodiments of the present invention, a multi-layer perceptron (MLP) is used to discriminate the type of natural input information, and it can quickly identify whether the input data is text, speech, video, etc. This type discrimination mechanism provides a clear direction for the subsequent processing flow, ensuring that the subsequent modules can adopt appropriate processing strategies for different types of data. The introduction of the MLP enables the model to quickly respond to different types of inputs, reduces unnecessary waste of computing resources, and improves the operating efficiency of the entire system.
[0111] Step S202, determine whether the type of natural input information is text data;
[0112] It should be noted that by determining whether the input information is text data, the system can clarify the specific path of subsequent processing. If it is text data, it directly enters the text processing module; if it is not text data, it enters the non-text processing module. This clear path division enables the system to efficiently process different types of data and avoids chaos during the processing.
[0113] Step S203, if the type of natural input information is text data, trigger the BERT pre-trained model in the CRNN architecture network. The BERT pre-trained model extracts word vectors, sentence vectors, and position vectors from the natural input information, and concatenates the word vectors, sentence vectors, and position vectors to obtain a text concatenation sequence;
[0114] In the embodiments of the present invention, the bidirectional Transformer architecture of BERT can capture bidirectional context information in the text, making the model's understanding of the text deeper. This context awareness ability is particularly important for the semantic understanding of complex texts. By extracting features through the BERT pre-trained model, it can significantly reduce the time and computing resource consumption of training the model from scratch, and at the same time improve the efficiency and accuracy of text processing.
[0115] Step S204, if the type of natural input information is non-text data, use a Mel filter to filter the natural input information, and extract and identify the semantic information in the natural input information based on the HSFs algorithm. Continuously map the high-dimensional semantic information to a low-dimensional feature space, and output a semantic extraction sequence;
[0116] In this embodiment, the Mel filter is widely used in speech processing and can extract spectral features in the speech signal, providing a basis for subsequent semantic extraction. Extracting high-dimensional semantic information based on the HSFs algorithm and mapping it to a low-dimensional feature space can effectively reduce the data dimension while retaining key semantic information. This processing method enables the model to better process non-text inputs and convert them into recognizable feature sequences.
[0117] Step S205: Load the text splicing sequence or semantic extraction sequence. The convolutional layer in the CRNN architecture network performs convolutional fusion on the features of the text splicing sequence or semantic extraction sequence to extract local features, obtaining a local feature set.
[0118] Step S206: Identify the local feature set through the global attention layer, extract the global features in the local feature set, obtaining the global feature sequences of the spliced word features, sentence features, position features, emotional features, and preference features. Through the global attention layer, the model can extract the emotional features and user preference features in the input data. These features are crucial for understanding the user's intention and generating personalized outputs.
[0119] Step S207: Obtain the global feature sequence. The BILSTM layer extracts the context semantic information of the global feature sequence and fuses the context semantic information as the extracted text output vector, and sends the text output vector to the Transformer encoder. The Transformer encoder encodes and labels it based on the probability of the text output vector in the picture type, obtaining the encoded and labeled text output vector. By fusing the context semantic information into the text output vector, the model can generate a compact and information-rich feature representation. This feature representation can better capture the semantic features of the input data and provide a basis for subsequent processing.
[0120] In the embodiment of the present invention, when the text recognition model recognizes and analyzes natural input information, through the MLP type discrimination and the processing strategies for different types of inputs, the text recognition model can support various input forms such as text and speech. Through the BERT pre-training model and the global attention layer, the text recognition model can extract the semantic information, emotional features, and user preference features in the input data. Through the Transformer encoder, it encodes and labels based on the probability of the text output vector in the picture type. The encoding and labeling ensure a high degree of consistency between the text and the picture at the semantic level, making the finally generated graphic and text content more natural and coordinated. The encoding and labeling is a representation method that associates the text output vector with the picture type. By encoding and labeling based on the probability of the text output vector in the picture type through the Transformer encoder, the text output vector is given probability labels related to the picture type, and these labels reflect the matching degree between the text content and the picture type. This encoding method enables the text features to be directly compared with the picture type, providing a quantitative basis for subsequent screening.
[0121] The embodiment of the present invention provides a method for determining the optimal matching set by using a multi-model voting algorithm to improve the generative adversarial network model. Figure 4The figure shows a schematic implementation process diagram of a method for determining an optimal matching set using a generative adversarial network model improved by a multi-model voting algorithm. The method for determining an optimal matching set using a generative adversarial network model improved by a multi-model voting algorithm specifically includes:
[0122] Step S301: Load the similar picture set, traverse the picture entities, picture types, and picture attributes of the similar picture set, perform picture alignment based on the timestamp alignment algorithm, perform a temporal relationship mapping on the picture entities, extract picture entity features using a generative adversarial network model improved by a multi-model voting algorithm, and calculate the attribute association value between the picture entity and the text output vector based on the DBSCAN algorithm.
[0123] It should be noted that when traversing the pre-constructed picture database with the text output vector as the index and grabbing at least one set of similar picture sets that match the text output vector based on a preset picture-text similarity threshold, using the encoded label in the text output vector as the index, the picture-text similarity threshold is set to 0.5 - 1. This step is to quickly grab picture entities in the picture database that have a relatively high similarity to the encoded label in the text output vector, so that multiple pictures similar to the text output vector can be grabbed at one time, achieving batch processing and further improving the processing speed. In addition, the setting of the picture-text similarity threshold is flexible and can be adjusted according to specific requirements. For example, if a more strict match is needed, the threshold can be increased; if more relevant pictures are desired, the threshold can be decreased. This retrieval method based on vectors and thresholds is applicable to various scenarios, including but not limited to image search, content recommendation, visual question answering, etc., and has strong generality and scalability.
[0124] In this embodiment, the introduction of the timestamp alignment algorithm and the temporal relationship mapping can handle pictures in dynamic scenarios, such as video frames or time series images. This processing method enables the model to better understand the temporal correlation of pictures and improves the adaptability to dynamic content. DBSCAN is a density-based clustering algorithm that can effectively identify the attribute association between picture entities and text output vectors. By calculating the attribute association value, the model can quantify the semantic correlation between pictures and text, providing an important basis for subsequent screening.
[0125] Step S302: Weightedly combine the picture entity features and the picture type features into a grayscale feature vector, and calculate the grayscale association value with the text output vector based on the principal component analysis method.
[0126] In the embodiments of the present invention, a gray-scale feature vector is generated by weighted combination of the picture entity features and the type features. This weighted combination method can comprehensively consider the content and type information of the picture, making the feature vector more representative. Principal Component Analysis (PCA) can effectively reduce the dimension and extract key features. By calculating the gray-scale correlation value through PCA, the model can further quantify the visual similarity between the picture and the text.
[0127] Step S303: Load the attribute correlation value and the gray-scale correlation value, and use the emotional features and preference features of the text output vector as constraints. Through the generative adversarial network model improved by the multi-model voting algorithm, weight the sum of the attribute correlation value and the gray-scale correlation value to calculate the comprehensive decoding value of the picture entity.
[0128] By using the multi-model voting algorithm to weight the sum of the attribute correlation value and the gray-scale correlation value, the model can comprehensively consider the semantic relevance and visual similarity to generate the comprehensive decoding value. This comprehensive consideration method enables the model to more comprehensively evaluate the matching degree between the picture and the text. And with the emotional features and preference features of the text output vector as constraints, the model can better understand the user's intentions and preferences, thereby generating a matching result that better meets the user's needs.
[0129] Step S304: Use the encoded label value of the text output vector as the screening threshold, retain the top M picture entities with the comprehensive decoding value greater than the encoded label value, and integrate the M picture entities to set them as the optimal matching set.
[0130] In this embodiment, the comprehensive decoding value of the picture entity is calculated by the following formula:
[0131]
[0132] where p j (q j , r j ) is the comprehensive decoding value of the picture entity, P is the number of picture entities in the similar picture set, W j is the picture entity feature, x i is the text output vector, q j , r j are the attribute correlation value and the gray-scale correlation value respectively, λ1 and λ2 are the emotional features and preference features of the text output vector respectively, σ(·) is the activation function of the generative adversarial network model improved by the multi-model voting algorithm, which can be the sigmoid activation function, B j represents the bias term of the generative adversarial network model improved by the multi-model voting algorithm, q0 represents the initial weight of the picture entity, ε j is the distance threshold of the ε neighborhood based on the DBSCAN algorithm, which can be 0.02 - 0.05, h j , respectively represent the picture type feature and the grayscale feature vector.
[0133] In the embodiment of the present invention, the DBSCAN algorithm is used to calculate the attribute correlation value, and the generative adversarial network model can effectively quantify the semantic correlation between pictures and texts, improving the matching accuracy. By calculating the grayscale correlation value through PCA, the generative adversarial network model can effectively quantify the visual similarity between pictures and texts, providing an important basis for the subsequent comprehensive decoding value calculation. Introducing the emotion feature and preference feature as constraint conditions can generate a matching result that better meets the user's needs and improves the user experience.
[0134] The embodiment of the present invention provides a method for a graphic generator to perform typesetting reconstruction on a text output vector and an optimal matching set based on a greedy strategy. Figure 5 The figure shows a schematic implementation flow diagram of the method for a graphic generator to perform typesetting reconstruction on a text output vector and an optimal matching set based on a greedy strategy. The method for a graphic generator to perform typesetting reconstruction on a text output vector and an optimal matching set based on a greedy strategy specifically includes:
[0135] Step S401, load the text output vector and the optimal matching set. The graphic generator performs an initial typesetting on the text output vector based on the concatenated word feature, sentence feature, and position feature in the text output vector to obtain an initial typesetting result.
[0136] In the embodiment of the present invention, based on the concatenated word feature, sentence feature, and position feature in the text output vector, an initial typesetting is performed on the text output vector. This initial typesetting takes into account the internal structure and features of the text, enabling the typesetting result to roughly reflect the logical and semantic relationships of the text. By directly using the features in the text output vector for typesetting, the complexity of designing the typesetting logic from scratch is avoided, improving the typesetting efficiency. The initial typesetting result can provide a reasonable framework for subsequent picture insertion, reducing the complexity of subsequent adjustments.
[0137] Step S402, the graphic generator sorts the picture entities in the optimal matching set in descending order according to the comprehensive decoding value, and gradually selects the optimal picture entities using the greedy strategy in the Kruskal algorithm, and imports the optimal picture entities into the initial typesetting result to complete the improvement of the initial typesetting result.
[0138] It should be noted that sorting the picture entities according to the comprehensive decoding value from high to low can ensure the selection of the picture that best matches the text content. This sorting method is based on the semantic association and visual similarity between the picture and the text, ensuring the accuracy of picture insertion. The greedy strategy in the Kruskal algorithm gradually selects the optimal picture entities, avoiding the complex calculation of the global optimal solution, while ensuring the efficient selection of the local optimal solution. Importing the optimal picture entities into the initial layout result step by step can dynamically adjust the layout, ensuring the coordination between the picture and the text. This step-by-step optimization method can effectively avoid typesetting conflicts caused by one-time insertion.
[0139] Step S403: Load the improved initial layout result, solve the text output vector and the minimum spanning tree problem of the optimal matching set based on the greedy strategy, select the spanning tree with the minimum weight to gradually construct the optimal solution, dynamically adjust the layout of the improved initial layout result, and output the layout generation result.
[0140] In this embodiment, solving the minimum spanning tree problem based on the greedy strategy can optimize the graphic and text layout from a global perspective. By selecting the spanning tree with the minimum weight to gradually construct the optimal solution, the graphic and text generator can dynamically adjust the layout of the text and pictures, ensuring the coordination and beauty of the overall layout. For example, during the layout process, the model can dynamically adjust the size and position of the pictures to avoid excessive overlap or blank areas.
[0141] On the other hand, the embodiment of the present invention also provides a graphic and text content matching generation system. Figure 6 The structural schematic diagram of the graphic and text content matching generation system is shown. The graphic and text content matching generation system specifically includes:
[0142] An information collection module 100, which is used to obtain natural input information, construct and pre-train a text recognition model by using the HSFs algorithm in combination with the CRNN model, and perform recognition and analysis on the natural input information based on the pre-trained text recognition model to obtain a text output vector;
[0143] A picture index module 200, in response to the text output vector, uses the text output vector as an index to traverse the pre-constructed picture database, grabs at least one set of similar picture sets that match the text output vector based on a preset graphic and text similarity threshold, and determines the optimal matching set by using a generative adversarial network model improved by the multi-model voting algorithm;
[0144] A layout generation module 300, which is used to load the text output vector and the optimal matching set, extract the features of the text output vector and the optimal matching set through a graphic and text generator, fuse the features of the text output vector and the optimal matching set, and the graphic and text generator performs layout reconstruction on the text output vector and the optimal matching set based on the greedy strategy, and outputs the layout generation result.
[0145] In this embodiment, the picture indexing module 200 includes:
[0146] A picture database 210, which is used to store picture entities associated with the text output vector, transform the formats of the picture entities, and label the length, angle, shape, position, direction, area, volume, saturation, and hue of the picture entities;
[0147] A database indexing unit 220, in response to the text output vector, uses the text output vector as an index to traverse the pre-constructed picture database, and grabs at least one set of similar picture sets that match the text output vector based on a preset picture-text similarity threshold;
[0148] A picture-text matching unit 230, which is used to load the text output vector and determine the optimal matching set by using a generative adversarial network model improved by a multi-model voting algorithm.
[0149] In another aspect of the present invention, a computer-readable storage medium is further provided. The readable storage medium stores a computer program, and when the computer program is executed by a processor, the picture-text content matching generation method described above is implemented.
[0150] The computer-readable storage medium herein (for example, a memory) can be a volatile memory or a non-volatile memory, or can include both a volatile memory and a non-volatile memory. By way of example and not limitation, the non-volatile memory can include a read-only memory (ROM), a programmable ROM (PROM), an electrically programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory can include a random access memory (RAM), and this RAM can act as an external cache memory. By way of example and not limitation, the RAM can be obtained in various forms, such as synchronous RAM (DRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct Rambus RAM (DRRAM). The storage devices of the disclosed aspects are intended to include, but are not limited to, these and other suitable types of memories.
[0151] In summary, the present invention provides a graphic and text content matching and generating system, method and storage medium. In the embodiments of the present invention, the HSFs algorithm is combined with the CRNN model to construct and pre-train a text recognition model, and the natural input information is recognized and analyzed based on the pre-trained text recognition model, so as to ensure accurate recognition of multi-modal natural input information while extracting the user's emotional features and preference features, ensuring accurate understanding of the user's intentions and preferences, and combining the generative adversarial network model improved by the multi-model voting algorithm to determine the optimal matching set, which can more comprehensively capture the semantic association between text and image, optimize the layout effect of graphics and text, and thus generate a higher-quality typesetting generation result. It overcomes the problem that the existing method determines candidate keywords through simple word segmentation and semantic analysis, and cannot accurately recognize multi-modal natural input information and deeply extract complex features in multi-modal input, resulting in inaccurate understanding of the user's intentions.
[0152] It should be noted that for the foregoing embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps may be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0153] In several embodiments provided by the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above-mentioned unit division is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection between devices or units can be in the form of telecommunications or other forms.
[0154] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0155] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting the protection scope of the invention. Obviously, the described embodiments are only partial embodiments of the present invention, rather than all embodiments. Based on these embodiments, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art can still, without conflict and without creative efforts, combine, add or delete the features in the embodiments of the present invention according to the circumstances or make other adjustments, so as to obtain different technical solutions that essentially do not deviate from the concept of the present invention, and these technical solutions also belong to the scope of protection of the present invention.
Claims
1. A method for matching and generating image and text content, characterized in that: The image-text content matching generation method comprises: Obtain natural input information, use the HSFs algorithm combined with the CRNN model to build and pre-train a text recognition model, and analyze the natural input information based on the pre-trained text recognition model to obtain a text output vector; In response to the text output vector, the pre-built image database is traversed with the text output vector as an index, and at least one set of similar images that match the text output vector is captured based on a preset image-text similarity threshold, and the optimal matching set is determined using a generative adversarial network model improved by a multi-model voting algorithm; Load the text output vector and the optimal matching set, extract the features of the text output vector and the optimal matching set through the image-text generator, and fuse the features of the text output vector and the optimal matching set. The image-text generator reconstructs the layout of the text output vector and the optimal matching set based on a greedy strategy and outputs the layout generation result.
2. The method for matching and generating image and text content according to claim 1, characterized in that: The method of using the HSFs algorithm in combination with the CRNN model to construct and pre-train a text recognition model specifically includes: The CRNN model is used as the initial model of the text recognition model. The initial model of the text recognition model includes an input layer, a convolution layer, a circulation layer, a transcription layer, and an output layer. The convolution layer is composed of three layers of convolutions, and the circulation layer is provided with three layers. The convolution layer, the circulation layer, and the transcription layer constitute a CRNN architecture network; The HSFs algorithm and global attention layer are introduced into the CRNN architecture network, and the multi-layer perceptron MLP is introduced before the convolution layer of the CRNN architecture network. The BERT pre-trained model is used to replace the transcription layer. The Transformer encoder and three groups of BILSTM layers are introduced between the CRNN architecture network and the output layer to complete the construction of the text recognition model. Using sample retrieval regular expressions combined with the Scrapy crawler framework to crawl multimodal data sources, where the multimodal data sources include text data, voice data, and video data, performing data cleaning, word segmentation and annotation, and decoding preprocessing on the multimodal data sources, and dividing the preprocessed multimodal data sources into a first training set and a first test set; Load the pre-built text recognition model, set the activation function, loss function, hyperparameters, and parameter optimizer of the text recognition model; Obtaining a first training set, and iteratively training a text recognition model using the first training set, wherein the text recognition model extracts training text features based on the first training set, and generates a training output vector based on the training text features; Based on the loss function, the loss value between the training output vector and the training true vector is determined, and the hyperparameters of the text recognition model are adjusted through the genetic algorithm to output a converged text recognition model. Obtain a first test set, use the first test set as input, execute the text recognition model, output the test accuracy, and determine whether the test accuracy exceeds a preset accuracy threshold. If the test accuracy exceeds the preset accuracy threshold, output a converged text recognition model.
3. The method for matching and generating image and text content according to claim 2, characterized in that: When the loss value of the training output vector and the training true vector is determined based on the loss function, the loss value calculation formula is as follows: L all =L cr +α1L bi +α2L tr (1) Among them, L all , L cr , L bi , L tr They represent the total loss of the text recognition model, global attention loss, text discrimination loss, and encoding loss, respectively. α1 and α2 represent the BILSTM layer hyperparameters and encoder hyperparameters, respectively. The global attention loss is: Among them, E cr represents the data expectation of the global attention loss value, A(x), They are the true value and predicted value extracted by the training text features through the CRNN architecture network; The text discrimination loss is: in, represents the cosine similarity between the training true vector and the training output vector after the training text feature is processed by the BILSTM layer, and φ represents the softmax temperature; The encoding loss is: in, represents the IOU value of the Transformer encoder, C(i), They represent the true coding label and predicted coding label of the training output vector after being processed by the Transformer encoder.
4. The method for matching and generating image and text content according to claim 2, characterized in that: The method for identifying and analyzing natural input information based on a pre-trained text recognition model specifically includes: Load the natural input information, and the multi-layer perceptron MLP in the text recognition model determines the type of the natural input information to determine whether the natural input information type is text data; If the natural input information type is text data, the BERT pre-training model in the CRNN architecture network is triggered. The BERT pre-training model extracts the word vector, sentence vector, and position vector from the natural input information, and concatenates the word vector, sentence vector, and position vector to obtain a text concatenation sequence. If the natural input information type is non-text data, the Mel filter is used to filter the natural input information, and the semantic information in the natural input information is extracted and identified based on the HSFs algorithm, the high-dimensional semantic information is continuously mapped to the low-dimensional feature space, and the semantic extraction sequence is output; Load the text splicing sequence or semantic extraction sequence, and the convolution layer in the CRNN architecture network performs convolution fusion on the features of the text splicing sequence or semantic extraction sequence to extract local features and obtain a local feature set; The global attention layer is used to identify the local feature set, extract the global features from the local feature set, and obtain the global feature sequence of concatenated word features, sentence features, position features, sentiment features, and preference features; The global feature sequence is obtained. The BILSTM layer extracts the contextual semantic information of the global feature sequence and fuses the contextual semantic information as the extracted text output vector. The text output vector is sent to the Transformer encoder. The Transformer encoder encodes the text output vector based on the probability of the image type to obtain the text output vector after encoding the label.
5. The method for matching and generating image and text content according to claim 1, characterized in that: The method for determining the optimal matching set using a generative adversarial network model improved by a multi-model voting algorithm comprises: Load similar image sets, traverse the image entities, image types, and image attributes of the similar image sets, align images based on the timestamp alignment algorithm, and map the temporal relationships of image entities. Use the generative adversarial network model improved by the multi-model voting algorithm to extract image entity features, and calculate the attribute association value between the image entity and the text output vector based on the DBSCAN algorithm. The image entity features and image type features are weighted and combined into a grayscale feature vector, and the grayscale correlation value with the text output vector is calculated based on the principal component analysis method; Load the attribute association value and grayscale association value, and use the sentiment and preference features of the text output vector as constraints. Use the generative adversarial network model improved by the multi-model voting algorithm to perform weighted summation of the attribute association value and grayscale association value to calculate the comprehensive decoding value of the image entity. The encoded label value of the text output vector is used as the screening threshold, the first M image entities whose comprehensive decoding values are greater than the encoded label value are retained, and the M image entities are integrated and set as the optimal matching set.
6. The method for matching and generating image and text content according to claim 5, characterized in that: The comprehensive decoding value of the picture entity is calculated by the following formula: Among them, p j (q j ,r j ) is the comprehensive decoding value of the picture entity, P is the number of picture entities in the similar picture set, and W j is the image entity feature, x i is the text output vector, q j ,r j are the attribute association value and grayscale association value, respectively. λ1 and λ2 are the sentiment feature and preference feature of the text output vector, respectively. σ(·) is the activation function of the generative adversarial network model improved by the multi-model voting algorithm. j represents the bias term of the generative adversarial network model improved by the multi-model voting algorithm, q0 represents the initial weight of the image entity, and ε j is the distance threshold of the ε neighborhood based on the DBSCAN algorithm, h j , Represent the image type features and grayscale feature vectors respectively.
7. The method for matching and generating image and text content according to claim 6, characterized in that: The method for the image-text generator to reconstruct the layout of the text output vector and the optimal matching set based on the greedy strategy includes: The text output vector and the best matching set are loaded, and the graphic text generator initially typesets the text output vector based on the concatenated word features, sentence features, and position features in the text output vector to obtain the initial typesetting result; The image and text generator sorts the image entities in the optimal matching set according to the comprehensive decoding value from high to low, and gradually selects the optimal image entity using the greedy strategy in the Kruskal algorithm. The optimal image entity is imported into the initial typesetting result to improve the initial typesetting result. Load the improved initial typesetting results, solve the text output vector and the optimal matching set minimum spanning tree problem based on the greedy strategy, select the spanning tree with the smallest weight to gradually build the optimal solution, dynamically adjust the layout of the improved initial typesetting results, and output the typesetting generation results.
8. A system for matching and generating graphic content, used to implement the graphic content matching and generating method as claimed in any one of claims 1 to 7, characterized in that: The image-text content matching generation system includes: The information acquisition module is used to obtain natural input information, use the HSFs algorithm combined with the CRNN model to build and pre-train a text recognition model, and analyze the natural input information based on the pre-trained text recognition model to obtain a text output vector; The image indexing module responds to the text output vector, takes the text output vector as an index, traverses the pre-built image database, grabs at least one set of similar images that match the text output vector based on a preset image-text similarity threshold, and uses a generative adversarial network model improved by a multi-model voting algorithm to determine the optimal matching set; The typesetting generation module is used to load the text output vector and the optimal matching set, extract the features of the text output vector and the optimal matching set through the graphic generator, and fuse the features of the text output vector and the optimal matching set. The graphic generator reconstructs the text output vector and the optimal matching set based on the greedy strategy and outputs the typesetting generation result.
9. The image-text content matching generation system according to claim 8, characterized in that: The image index module includes: The image database is used to store image entities associated with the text output vector, transform the format of the image entities, and annotate the length, angle, shape, position, direction, area, volume, saturation and hue of the image entities; A database indexing unit, in response to the text output vector, uses the text output vector as an index, traverses a pre-built image database, and captures at least one set of similar images that match the text output vector based on a preset image-text similarity threshold; The image-text matching unit is used to load the text output vector and determine the optimal matching set using the generative adversarial network model improved by the multi-model voting algorithm.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for matching and generating graphic and text content as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Methods and equipment for generating text and image content
CN110020411B
Cited By
Intelligent contract analysis method and system based on multi-modal feature fusion algorithm
CN120822115A
A Smart Contract Analysis Method and System Based on Multimodal Feature Fusion Algorithm
CN120822115B