Multi-modal content identification method and system based on image and text information fusion
By setting a preprocessing strategy to process image and text information, generating a text-image mapping table and fusing it, the problem of insufficient utilization of information complementarity in multimodal content recognition is solved, and accurate and efficient recognition of multimodal content is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-27
- Publication Date
- 2026-03-10
AI Technical Summary
Existing multimodal content recognition technologies cannot provide reasonable preprocessing strategies and are unable to fully utilize and understand the complementarity between multimodal information, resulting in inaccurate content information recognition and low efficiency.
By setting a first preprocessing strategy and a second preprocessing strategy, real-time image and text information are preprocessed to generate image evaluation values and text evaluation values. Correlation analysis is performed to construct a text-image mapping table and generate a comprehensive fusion result, outputting multimodal content recognition results.
It achieves comprehensive and accurate identification of multimodal content, improving identification efficiency and accuracy.
Smart Images

Figure CN121640231A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of multimodal content recognition technology, and in particular to a multimodal content recognition method based on the fusion of image and text information. Background Technology
[0002] In today's era of digital information explosion, multimodal content recognition technology plays an increasingly important role in many fields. However, current multimodal content recognition technology still has some shortcomings. It cannot provide reasonable preprocessing strategies to mine effective features, nor can it fully utilize and understand the complementarity and correlation between multimodal information, making it difficult to meet the needs of accurate and efficient processing of content information. Summary of the Invention
[0003] To address the aforementioned technical issues, this application provides a multimodal content recognition method based on the fusion of image and text information. By setting a first preprocessing strategy and a second preprocessing strategy, several real-time sub-images and real-time sub-texts are obtained and correlation analysis is performed to obtain a text-image mapping table and several single fusion results. A comprehensive fusion result is generated and the content recognition result is output. By utilizing the correlation between multimodal information, comprehensive and accurate recognition of multimodal content can be achieved.
[0004] Some embodiments of this application provide a multimodal content recognition method based on the fusion of image and text information, including: The system acquires real-time image information and real-time text information of the subject to be identified, generates image evaluation value based on real-time image information, generates text evaluation value based on real-time text information, and sets a first preprocessing strategy and a second preprocessing strategy respectively. The real-time image information is preprocessed according to the first preprocessing strategy to obtain several real-time sub-images and corresponding image features. The real-time text information is preprocessed according to the second preprocessing strategy to obtain several real-time sub-texts and corresponding text features. A correlation analysis is performed on several real-time sub-images and several real-time sub-texts. Based on the analysis results, a text-image mapping table is obtained, and several single fusion results are generated. Generate a comprehensive fusion result based on all individual fusion results, and output the multimodal content recognition result.
[0005] In some embodiments of this application, generating image evaluation values based on real-time image information and generating text evaluation values based on real-time text information includes: Several image evaluation indicators are pre-defined; The image sub-evaluation value for each image evaluation index is generated based on real-time image information, and the image evaluation value is obtained by combining the weight coefficients of the image evaluation index. Several text evaluation indicators are pre-defined; The text sub-evaluation value for each text evaluation indicator is generated based on real-time text information, and the text evaluation value is obtained by combining the weight coefficients of the text evaluation indicators.
[0006] In some embodiments of this application, a first preprocessing strategy and a second preprocessing strategy are respectively set, including: A first preset image evaluation value range, a second preset image evaluation value range, a third preset image evaluation value range, and a fourth preset image evaluation value range are preset. When the image evaluation value is within the first preset image rating value range, the fourth preset first preprocessing strategy is selected as the first preprocessing strategy for real-time image information. When the image evaluation value is within the second preset image rating value range, the third preset first preprocessing strategy is selected as the first preprocessing strategy for real-time image information. When the image evaluation value is within the third preset image rating value range, the second preset first preprocessing strategy is selected as the first preprocessing strategy for real-time image information. When the image evaluation value is within the fourth preset image rating value range, the first preset first preprocessing strategy is selected as the first preprocessing strategy for real-time image information. The first preset text evaluation value range, the second preset text evaluation value range, the third preset text evaluation value range, and the fourth preset text evaluation value range are preset. When the text evaluation value is within the first preset text rating value range, the fourth preset second preprocessing strategy is selected as the second preprocessing strategy for real-time text information. When the text evaluation value is within the second preset text rating value range, the third preset second preprocessing strategy is selected as the second preprocessing strategy for real-time text information. When the text evaluation value is within the third preset text rating value range, the second preset second preprocessing strategy is selected as the second preprocessing strategy for real-time text information. When the text evaluation value is within the fourth preset text rating value range, the first preset second preprocessing strategy is selected as the second preprocessing strategy for real-time text information.
[0007] In some embodiments of this application, the real-time image information is preprocessed according to a first preprocessing strategy to obtain several real-time sub-images and corresponding image features, and the real-time text information is preprocessed according to a second preprocessing strategy to obtain several real-time sub-texts and corresponding text features, including: The real-time image information is preprocessed according to the first preprocessing strategy to obtain several preprocessed real-time sub-images. Each real-time sub-image is input into a pre-built image feature extraction model to obtain the image features of each real-time sub-image; The real-time text information is preprocessed according to the second preprocessing strategy to obtain several preprocessed real-time sub-texts. Each real-time subtext is input into a pre-built text feature extraction model to obtain the text features of each real-time subtext.
[0008] In some embodiments of this application, before performing correlation analysis on several real-time sub-images and several real-time sub-texts, the following steps are also included: The image features of adjacent real-time sub-images are correlated to obtain several first correlation coefficients; The first comprehensive correlation coefficient of adjacent real-time sub-images is generated based on several first correlation coefficients and the weight coefficients of corresponding image features; Real-time sub-images with a first comprehensive correlation coefficient greater than a preset first comprehensive correlation coefficient threshold are stitched together, and image features of several same dimensions of the stitched real-time sub-images are fused to obtain the stitched sub-image and the corresponding fused image features. Construct a sub-image sequence based on the stitched sub-images and the unstitched real-time sub-images; By associating the text features of adjacent real-time sub-texts, several second association coefficients are obtained; The second comprehensive correlation coefficient of adjacent real-time sub-texts is generated based on several second correlation coefficients and the weight coefficients of the corresponding text features. Real-time sub-texts with a second comprehensive correlation coefficient greater than a preset second comprehensive correlation coefficient threshold are concatenated, and several text features of the same dimension of the concatenated real-time sub-texts are fused to obtain concatenated sub-texts and corresponding fused text features. Construct a subtext sequence based on the concatenated subtext and the unconcatenated real-time subtext; Extract keywords from each subtext in the subtext sequence and generate query conditions for the corresponding subtext, wherein the query conditions include whether there is a correlation between the keywords and the subtext. Based on the query conditions of each sub-text, traverse each sub-image in the sub-image sequence and its corresponding image features, filter out the sub-images that meet the query conditions of each sub-text, calculate the corresponding first relevance, and construct the sub-image sub-sequence associated with each sub-text according to the first relevance. The first correlation degree is calculated based on the number of keywords that are related to the same sub-image, the corresponding initial correlation degree, and the weight coefficient of the corresponding keyword.
[0009] In some embodiments of this application, correlation analysis is performed on several real-time sub-images and several real-time sub-texts, including: Randomly select one sub-image from the sub-image sub-sequence associated with each sub-text, and sort the selected sub-images according to the order of the sub-texts in the sub-text sequence to obtain a sub-image sequence to be analyzed; Generate several sub-image sequences to be analyzed; The second correlation degree of the corresponding sub-image sequence to be analyzed is generated based on the first correlation degree between all sub-images and associated sub-texts in the same sub-image sequence to be analyzed and the weight coefficient of the corresponding sub-texts. Based on the logical relationship between adjacent subtexts in the subtext sequence, a first correlation threshold is set between adjacent subimages in each subimage sequence to be analyzed, and a first correction coefficient is generated based on the first correlation degree between the subimage and the corresponding subtext in each subimage sequence to be analyzed. The first correlation threshold is corrected according to the first correction coefficient to obtain the second correlation threshold; Calculate the actual correlation value between adjacent sub-images in the same sub-image sequence to be analyzed, subtract the actual correlation value from the corresponding second correlation threshold to obtain the actual correlation value difference, and perform mean processing on the mean processing result to obtain the second correction coefficient; The second correlation degree of the corresponding sub-image sequence to be analyzed is corrected according to the second correction coefficient to obtain the third correlation degree; Each sub-image in the sub-image sequence with the highest third correlation is set as the main sub-image of each associated sub-text; The main sub-image is compared and analyzed with other sub-images in the sub-image sequence of the associated sub-text to obtain the image difference degree. Sub-images with an image difference degree greater than a preset difference degree threshold are set as secondary sub-images of the corresponding sub-text.
[0010] In some embodiments of this application, a text-image mapping table is obtained based on the analysis results, including: Construct a corresponding subtext-subimage mapping table for the main and secondary subimages of each subtext; The subtext-subimage mapping table includes several main feature information associated with the main subimage and its corresponding subtext, as well as secondary feature information associated with the secondary subimage and its corresponding subtext. Construct a text-image mapping table based on the subtext-subimage mapping table of all subtexts and the order in which the subtexts are arranged.
[0011] In some embodiments of this application, several single fusion results are generated, including: For each sub-text and the main sub-image, generate several first fusion feature information, and form a basic fusion vector based on the several first fusion feature information; For each sub-text and secondary sub-image, several secondary feature information is generated to form several second fusion feature information. After dynamically allocating weights through an attention mechanism, the second fusion is performed with the basic fusion vector to generate multimodal fusion features. Sequence fusion features are obtained by performing sequence modeling on all multimodal fusion features based on a bidirectional long short-term memory network; The sequence fusion features are reduced in dimensionality and classified based on a fully connected layer, and several single fusion results are output.
[0012] In some embodiments of this application, a comprehensive fusion result is generated based on all individual fusion results, and a multimodal content recognition result is output, including: The weighted summation of the individual fusion results corresponding to all sub-texts is used to obtain the comprehensive fusion vector; Principal component analysis is performed on the integrated vector to extract principal component features; The principal component features are input into a pre-trained multimodal classification model to generate the final integrated fusion result; The integrated results are structured and organized according to a preset output format, and the output is a multimodal content recognition result.
[0013] In some embodiments of this application, a multimodal content recognition system based on the fusion of image and text information is also included: The acquisition module is used to acquire real-time image information and real-time text information of the subject to be identified, generate image evaluation value based on real-time image information, generate text evaluation value based on real-time text information, and set a first preprocessing strategy and a second preprocessing strategy respectively. The preprocessing module is used to preprocess real-time image information according to the first preprocessing strategy to obtain several real-time sub-images and corresponding image features, and to preprocess real-time text information according to the second preprocessing strategy to obtain several real-time sub-texts and corresponding text features. The association module is used to perform association analysis on several real-time sub-images and several real-time sub-texts, obtain a text-image mapping table based on the analysis results, and generate several single fusion results. The generation module is used to generate a comprehensive fusion result based on all individual fusion results and output the multimodal content recognition result.
[0014] The multimodal content recognition method based on image and text information fusion in this application has the following advantages compared with the prior art: By setting a first preprocessing strategy and a second preprocessing strategy, several real-time sub-images and real-time sub-texts are obtained and correlation analysis is performed to obtain a text-image mapping table and several single fusion results. A comprehensive fusion result is generated and the content recognition result is output. By utilizing the correlation between multimodal information, comprehensive and accurate recognition of multimodal content is achieved. Attached Figure Description
[0015] Figure 1 This is a flowchart illustrating the multimodal content recognition method based on the fusion of image and text information in an embodiment of this application. Detailed Implementation
[0016] The specific embodiments of this application will be described in further detail below with reference to the accompanying drawings and examples. The following examples are used to illustrate this application, but are not intended to limit the scope of this application.
[0017] In the description of this application, it should be understood that the terms "center", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application.
[0018] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0019] In the description of this application, it should be noted that, unless otherwise expressly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.
[0020] like Figure 1 As shown in the figure, the multimodal content recognition method based on image and text information fusion in this application includes: Step S101: Obtain real-time image information and real-time text information of the subject to be identified, generate image evaluation value based on real-time image information, generate text evaluation value based on real-time text information, and set the first preprocessing strategy and the second preprocessing strategy respectively. Step S102: Preprocess the real-time image information according to the first preprocessing strategy to obtain several real-time sub-images and corresponding image features; preprocess the real-time text information according to the second preprocessing strategy to obtain several real-time sub-texts and corresponding text features. Step S103: Perform correlation analysis on several real-time sub-images and several real-time sub-texts, obtain a text-image mapping table based on the analysis results, and generate several single fusion results; Step S104: Generate a comprehensive fusion result based on all individual fusion results, and output the multimodal content recognition result.
[0021] In this embodiment, the real-time image information includes static image data and dynamic video data of the subject to be identified, and the real-time text information includes structured text data and unstructured text data of the subject to be identified. The structured text data includes data with a fixed format such as the subject's attribute tags, classification information, and key parameters, while the unstructured text data covers free-format text content such as the subject's descriptive text, context, and user comments.
[0022] In some embodiments of this application, generating image evaluation values based on real-time image information and generating text evaluation values based on real-time text information includes: Several image evaluation indicators are pre-defined; The image sub-evaluation value for each image evaluation index is generated based on real-time image information, and the image evaluation value is obtained by combining the weight coefficients of the image evaluation index. Several text evaluation indicators are pre-defined; The text sub-evaluation value for each text evaluation indicator is generated based on real-time text information, and the text evaluation value is obtained by combining the weight coefficients of the text evaluation indicators.
[0023] In this embodiment, the image evaluation metrics cover dimensions such as sharpness, contrast, color saturation, target outline completeness, coverage of key features (object outline, text region), and area. The text evaluation metrics cover dimensions such as semantic completeness, grammatical correctness, keyword density, and domain relevance.
[0024] In this embodiment, actual values of various image evaluation indicators associated with real-time image information are generated and compared with corresponding preset standard values. The greater the difference from the preset standard value, the smaller the weighted image evaluation value, that is, the lower the clarity, contrast, color saturation, and target outline integrity, and the larger the coverage and area of key features, and vice versa.
[0025] In this embodiment, the actual values of each text evaluation index are obtained by analyzing real-time text information and compared with the corresponding preset standard values. The greater the difference from the preset standard value, the smaller the text evaluation value obtained after quantification, that is, the lower the semantic integrity, grammatical correctness, and domain relevance, and the higher the keyword density, and vice versa.
[0026] In this embodiment, by calculating image evaluation values and text evaluation values, a foundation is laid for selecting a reasonable preprocessing strategy and extracting features, thereby achieving deep fusion of image and text information and accurate recognition of multimodal content, ensuring the efficient and stable operation of the multimodal content recognition process.
[0027] In some embodiments of this application, a first preprocessing strategy and a second preprocessing strategy are respectively set, including: A first preset image evaluation value range, a second preset image evaluation value range, a third preset image evaluation value range, and a fourth preset image evaluation value range are preset. When the image evaluation value is within the first preset image rating value range, the fourth preset first preprocessing strategy is selected as the first preprocessing strategy for real-time image information. When the image evaluation value is within the second preset image rating value range, the third preset first preprocessing strategy is selected as the first preprocessing strategy for real-time image information. When the image evaluation value is within the third preset image rating value range, the second preset first preprocessing strategy is selected as the first preprocessing strategy for real-time image information. When the image evaluation value is within the fourth preset image rating value range, the first preset first preprocessing strategy is selected as the first preprocessing strategy for real-time image information. The first preset text evaluation value range, the second preset text evaluation value range, the third preset text evaluation value range, and the fourth preset text evaluation value range are preset. When the text evaluation value is within the first preset text rating value range, the fourth preset second preprocessing strategy is selected as the second preprocessing strategy for real-time text information. When the text evaluation value is within the second preset text rating value range, the third preset second preprocessing strategy is selected as the second preprocessing strategy for real-time text information. When the text evaluation value is within the third preset text rating value range, the second preset second preprocessing strategy is selected as the second preprocessing strategy for real-time text information. When the text evaluation value is within the fourth preset text rating value range, the first preset second preprocessing strategy is selected as the second preprocessing strategy for real-time text information.
[0028] In this embodiment, the first preset image evaluation value range < the second preset image evaluation value range < the third preset image evaluation value range < the fourth preset image evaluation value range, the first preset first preprocessing strategy < the second preset first preprocessing strategy < the third preset first preprocessing strategy < the fourth preset first preprocessing strategy, the first preset text evaluation value range < the second preset text evaluation value range < the third preset text evaluation value range < the fourth preset text evaluation value range, and the first preset second preprocessing strategy < the second preset second preprocessing strategy < the third preset second preprocessing strategy < the fourth preset second preprocessing strategy. The aforementioned preset image evaluation value ranges, text evaluation value ranges, and corresponding preset preprocessing strategies are all pre-set, and the correspondence is set based on the historical preprocessing process and feature extraction results.
[0029] In this embodiment, the first preprocessing strategy includes data cleaning parameters, noise suppression parameters, geometric correction parameters, brightness balance parameters, image enhancement parameters, and image segmentation parameters; the second preprocessing strategy includes data cleaning parameters, text normalization parameters, word segmentation parameters, semantic density parameters, semantic enhancement parameters, and feature adaptation parameters. In this embodiment, the data cleaning parameters in the first preprocessing strategy include setting a pixel threshold to filter invalid image data; the noise suppression parameters include dynamically adjusting the size of the filtering window according to the noise intensity; the geometric correction parameters include correcting geometric distortion caused by shooting angle deviation; the brightness balance parameters include optimizing the overall brightness uniformity of the image; the image enhancement parameters include sharpening key features; and the image segmentation parameters include dividing the preprocessed image into several independent image regions.
[0030] In this embodiment, the differences between the various preset first preprocessing strategies mainly lie in the specific values of each parameter and the processing intensity. For example, the fourth preset first preprocessing strategy sets the pixel threshold to the lowest, uses the maximum filtering window for the noise suppression parameter, enables multi-level perspective transformation for the geometric correction parameter, maximizes the dynamic range adjustment of the brightness balance parameter, achieves the highest sharpening intensity for the image enhancement parameter, and uses the finest region division granularity for the image segmentation parameter; conversely, the opposite is true. Through this differentiated parameter configuration, real-time image information of different qualities can receive the most suitable preprocessing, thereby laying the foundation for accurate extraction of subsequent image features.
[0031] In this embodiment, the data cleaning parameters in the second preprocessing strategy include removing irrelevant characters, duplicate content, and data with format errors; the text standardization parameters include unifying the format; the word segmentation parameters include segmenting continuous text sequences into sub-word units; the semantic density parameters include quantifying the semantic condensation of text information; the semantic enhancement parameters include performing contextual semantic expansion on the segmented text units; and the feature adaptation parameters include standardizing and dimensionality reduction processing on the extracted text features.
[0032] In this embodiment, the differences between different preset second preprocessing strategies are also reflected in the configuration strength of each parameter. For example, the fourth preset second preprocessing strategy has the most stringent filtering rules for data cleaning parameters, text standardization parameters that cover the entire domain terminology database, word segmentation parameters that enable the maximum domain dictionary matching range, semantic density parameters that tilt their calculation weights towards core semantic units, semantic enhancement parameters that have the greatest expansion depth, and feature adaptation parameters that have the highest dimensionality reduction. Conversely, the opposite is true. Through this hierarchical parameter adjustment mechanism, accurate preprocessing of real-time text information of different qualities is achieved, ensuring the effectiveness and reliability of text feature extraction.
[0033] In some embodiments of this application, real-time image information is preprocessed according to a first preprocessing strategy to obtain several real-time sub-images and corresponding image features, and real-time text information is preprocessed according to a second preprocessing strategy to obtain several real-time sub-texts and corresponding text features, including: The real-time image information is preprocessed according to the first preprocessing strategy to obtain several preprocessed real-time sub-images. Each real-time sub-image is input into a pre-built image feature extraction model to obtain the image features of each real-time sub-image; The real-time text information is preprocessed according to the second preprocessing strategy to obtain several preprocessed real-time sub-texts. Each real-time subtext is input into a pre-built text feature extraction model to obtain the text features of each real-time subtext.
[0034] In this embodiment, a real-time sub-image is an independent image unit obtained by dividing the preprocessed real-time image information into regions using image segmentation parameters.
[0035] In this embodiment, the image features include edge contour features, texture features, color features, and semantic features corresponding to the real-time sub-image, and the text features include basic statistical features, word embedding features, and deep semantic features corresponding to the real-time sub-text.
[0036] In this embodiment, the image feature extraction model and the text feature extraction model are obtained by performing feature annotation based on preprocessed historical image information and historical text information, and by training a neural network based on the annotated historical image features and historical text features.
[0037] In this embodiment, edge contour features describe the outline shape and boundary direction of the subject in the sub-image; texture features reflect the texture thickness, direction and density of the sub-image surface; color features capture the color composition and distribution rules of the sub-image; and semantic features obtain a semantic vector containing category attributes, scene information and abstract concepts.
[0038] In this embodiment, real-time subtext is an independent text segment obtained by dividing the preprocessed real-time text information into semantic units through text segmentation processing.
[0039] In this embodiment, basic statistical features refer to basic morphological features; word embedding features refer to capturing the semantic association and contextual distribution patterns of words; deep semantic features refer to generating high-level semantic vectors that include contextual semantics, sentiment tendencies, and domain knowledge. Through the extraction of the above multi-dimensional text features, the semantic connotation of the subject to be identified in real-time text information can be deeply explored, providing comprehensive text feature support for subsequent deep fusion with image features.
[0040] In some embodiments of this application, before performing correlation analysis on several real-time sub-images and several real-time sub-texts, the following steps are also included: The image features of adjacent real-time sub-images are correlated to obtain several first correlation coefficients; The first comprehensive correlation coefficient of adjacent real-time sub-images is generated based on several first correlation coefficients and the weight coefficients of corresponding image features; Real-time sub-images with a first comprehensive correlation coefficient greater than a preset first comprehensive correlation coefficient threshold are stitched together, and image features of several same dimensions of the stitched real-time sub-images are fused to obtain the stitched sub-image and the corresponding fused image features. Construct a sub-image sequence based on the stitched sub-images and the unstitched real-time sub-images; By associating the text features of adjacent real-time sub-texts, several second association coefficients are obtained; The second comprehensive correlation coefficient of adjacent real-time sub-texts is generated based on several second correlation coefficients and the weight coefficients of the corresponding text features. Real-time sub-texts with a second comprehensive correlation coefficient greater than a preset second comprehensive correlation coefficient threshold are concatenated, and several text features of the same dimension of the concatenated real-time sub-texts are fused to obtain concatenated sub-texts and corresponding fused text features. Construct a subtext sequence based on the concatenated subtext and the unconcatenated real-time subtext; Extract keywords from each subtext in the subtext sequence and generate query conditions for the corresponding subtext, wherein the query conditions include whether there is a correlation between the keywords and the subtext. Based on the query conditions of each sub-text, traverse each sub-image in the sub-image sequence and its corresponding image features, filter out the sub-images that meet the query conditions of each sub-text, calculate the corresponding first relevance, and construct the sub-image sub-sequence associated with each sub-text according to the first relevance. The first correlation degree is calculated based on the number of keywords that are related to the same sub-image, the corresponding initial correlation degree, and the weight coefficient of the corresponding keyword.
[0041] In this embodiment, the image is segmented into several real-time sub-images and real-time sub-texts during the preprocessing stage. Before the association stage, by calculating the association coefficients of adjacent sub-images and adjacent sub-texts, sub-image groups and sub-text groups with strong semantic associations can be effectively identified, thereby enabling splicing and feature fusion, laying a more comprehensive structured data foundation for subsequent image and text association.
[0042] In this embodiment, the first correlation coefficient is calculated based on the matching degree of edge contours, the similarity of texture features, the distance metric of color features, and the cosine similarity of semantic feature vectors in image features. The first comprehensive correlation coefficient is obtained by weighted summation. The second correlation coefficient is obtained by quantifying the word embedding vector similarity, deep semantic vector distance, core semantic unit overlap rate, and domain attribute matching degree in text features.
[0043] In this embodiment, the first comprehensive correlation coefficient threshold and the second comprehensive correlation coefficient threshold are the minimum correlation coefficients that indicate the correlation between adjacent sub-images and adjacent sub-texts, respectively.
[0044] In this embodiment, image features of the same dimension are fused using a weighted average method to obtain more accurate and comprehensive fused image features. Text feature fusion adopts a semantic density-weighted splicing strategy, where subtexts with higher semantic density account for a larger proportion of the fused features.
[0045] In this embodiment, each subtext includes at least one keyword that can reflect the core theme and key information of the subtext. For example, when the keyword of the subtext is "red", the query condition is to filter subimages containing the visual feature of red; when the keyword is "production date", the query condition is to filter subimages containing text areas and whose text content is related to "production date".
[0046] In this embodiment, the initial correlation degree refers to the basic matching degree between the image features of a sub-image and the keywords of a sub-text. For example, if the keyword of a sub-text is "red sneakers", where "sneakers" is set as the entity keyword with a weight of 0.6 and "red" is set as the attribute keyword with a weight of 0.4, when the red component accounts for 60% of the color features of a sub-image and the matching degree between the semantic feature vector and "sneakers" is 0.8, the initial correlation degree can be calculated as (0.6×0.8 + 0.4×0.6) = 0.72.
[0047] In this embodiment, the greater the number of associated keywords, the higher the initial relevance, and the higher the weight coefficient, the greater the first relevance.
[0048] In this embodiment, the same sub-image appears only in the sub-image sub-sequence of a sub-text, that is, in the first sub-text with the highest relevance.
[0049] In this embodiment, by splicing related real-time sub-images and real-time sub-texts to construct sub-image sequences and sub-text sequences, the originally scattered local features are organized into more accurate, comprehensive, and context-related structured feature sequences. The sub-image sub-sequence of each sub-text is determined, realizing the preliminary association and localization of text and image features. This provides accurate association objects and structured data support for the subsequent deep fusion of multimodal features.
[0050] In some embodiments of this application, correlation analysis is performed on several real-time sub-images and several real-time sub-texts, including: Randomly select one sub-image from the sub-image sub-sequence associated with each sub-text, and sort the selected sub-images according to the order of the sub-texts in the sub-text sequence to obtain a sub-image sequence to be analyzed; Generate several sub-image sequences to be analyzed; The second correlation degree of the corresponding sub-image sequence to be analyzed is generated based on the first correlation degree between all sub-images and associated sub-texts in the same sub-image sequence to be analyzed and the weight coefficient of the corresponding sub-texts. Based on the logical relationship between adjacent subtexts in the subtext sequence, a first correlation threshold is set between adjacent subimages in each subimage sequence to be analyzed, and a first correction coefficient is generated based on the first correlation degree between the subimage and the corresponding subtext in each subimage sequence to be analyzed. The first correlation threshold is corrected according to the first correction coefficient to obtain the second correlation threshold; Calculate the actual correlation value between adjacent sub-images in the same sub-image sequence to be analyzed, subtract the actual correlation value from the corresponding second correlation threshold to obtain the actual correlation value difference, and perform mean processing on the mean processing result to obtain the second correction coefficient; The second correlation degree of the corresponding sub-image sequence to be analyzed is corrected according to the second correction coefficient to obtain the third correlation degree; Each sub-image in the sub-image sequence with the highest third correlation is set as the main sub-image of each associated sub-text; The main sub-image is compared and analyzed with other sub-images in the sub-image sequence of the associated sub-text to obtain the image difference degree. Sub-images with an image difference degree greater than a preset difference degree threshold are set as secondary sub-images of the corresponding sub-text.
[0051] In this embodiment, the logical relationships between adjacent sub-texts include causality, progression, parallelism, and reference. Different logical relationships correspond to corresponding first association thresholds. The above correspondences are pre-set. When there is a logical relationship between adjacent sub-texts, the sub-images associated with the sub-texts will also have corresponding association strength requirements.
[0052] In this embodiment, the first correction coefficient is calculated and converted based on the first correlation degree between adjacent sub-images corresponding to adjacent sub-texts. For example, if a and b are adjacent sub-texts, and a1 and b1 are adjacent sub-images corresponding to adjacent sub-texts, and the first correlation degree between a and a1 is 0.6, and the first correlation degree between b and b1 is 0.7, then the first correction coefficient = (0.6 + 0.7) / 2 = 0.65. If the logical relationship between a and b is "progressive", and the first correlation threshold is 0.75, then the corrected second correlation threshold = 0.75 × 0.65 = 0.4875.
[0053] In this embodiment, the actual association value is obtained by calculating the weighted comprehensive value of image features (such as edge contour matching degree, texture similarity, color distance, etc.) of adjacent sub-images. The difference of the actual association value = actual association value - second association threshold. When the mean is negative and smaller, the second correction coefficient obtained by quantization is smaller. When the mean is positive and larger, the second correction coefficient obtained by quantization is larger. The value range of the second correction coefficient is (0.8, 1.2).
[0054] In this embodiment, the process of image difference degree calculation is similar to that of calculating actual association value, and will not be repeated here. The purpose is to determine local detail features or auxiliary semantic information that can be supplemented by the main sub-image. For example, in the sub-text association of "red sneakers", the main sub-image may be an image of sneakers showing the overall appearance, while the secondary sub-image may include sub-images such as close-up of the sole texture and details of the shoe label, thereby providing richer visual information dimensions for multimodal feature fusion.
[0055] In this embodiment, the preset difference threshold refers to the minimum difference that can represent a significant difference between different sub-images, and it is set in advance based on historical data.
[0056] In this embodiment, by randomly generating the sequence of sub-images to be analyzed multiple times and calculating the third correlation degree, the optimal matching sequence is finally selected. This sequence can satisfy the logical correlation requirements of the sub-text sequence and the matching strength between the sub-image and the sub-text to the greatest extent, ensuring that the correlation between the sub-text and the sub-image conforms to the text logic relationship and the image feature correlation. At the same time, by dividing the sub-images into primary and secondary sub-images, the core related image information is retained while supplementing details, providing hierarchical image association objects for subsequent multimodal feature fusion.
[0057] In some embodiments of this application, a text-image mapping table is obtained based on the analysis results, including: Construct a corresponding subtext-subimage mapping table for the main and secondary subimages of each subtext; The subtext-subimage mapping table includes several main feature information associated with the main subimage and its corresponding subtext, as well as secondary feature information associated with the secondary subimage and its corresponding subtext. Construct a text-image mapping table based on the subtext-subimage mapping table of all subtexts and the order in which the subtexts are arranged.
[0058] In this embodiment, the subtext-subimage mapping table uses the unique identifier of the subtext as an index to record the main feature information associated with the corresponding main subimage and the corresponding subtext (such as key texture parameters in texture features, main color distribution in color features, core category labels in semantic features, etc.) and the secondary feature information associated with the secondary subimage and the corresponding subtext (such as local texture details, auxiliary color components, scene association information, etc.), forming a structured feature association record.
[0059] In this embodiment, the text-image mapping table is based on the arrangement order of the sub-text sequence, integrating the contents of all sub-text-sub-image mapping tables. It sequentially associates the corresponding primary sub-images, secondary sub-images and their feature information according to the position of the sub-text in the sequence. At the same time, it records the logical relationship between sub-texts and the correlation strength of adjacent sub-images in the sub-image sequence. Finally, it forms a comprehensive mapping model that can intuitively reflect the multi-dimensional correlation between the subject to be identified in real-time text information and real-time image information, providing a clear structured data index for the output of multimodal content recognition results.
[0060] In some embodiments of this application, several single fusion results are generated, including: For each sub-text and the main sub-image, generate several first fusion feature information, and form a basic fusion vector based on the several first fusion feature information; For each sub-text and secondary sub-image, several secondary feature information is generated to form several second fusion feature information. After dynamically allocating weights through an attention mechanism, the second fusion is performed with the basic fusion vector to generate multimodal fusion features. Sequence fusion features are obtained by performing sequence modeling on all multimodal fusion features based on a bidirectional long short-term memory network; The sequence fusion features are reduced in dimensionality and classified based on a fully connected layer, and several single fusion results are output.
[0061] In this embodiment, the weight allocation of the attention mechanism is based on the first correlation between the secondary sub-image and the sub-text and the information entropy of the features of the secondary sub-image. The secondary sub-image with the higher the first correlation and the greater the information entropy (i.e. the richer the feature diversity), the greater the weight of its features in the secondary fusion.
[0062] In this embodiment, the hidden layer state vector of BiLSTM integrates the semantic information of forward and backward multimodal fusion features, effectively solving the problem of long sequence dependencies. The fully connected layer uses the Softmax activation function to output the class probability distribution, and at the same time optimizes the network parameters through gradient backpropagation to ensure the classification accuracy of the single fusion result.
[0063] In some embodiments of this application, a comprehensive fusion result is generated based on all individual fusion results, and a multimodal content recognition result is output, including: The weighted summation of the individual fusion results corresponding to all sub-texts is used to obtain the comprehensive fusion vector; Principal component analysis is performed on the integrated vector to extract principal component features; The principal component features are input into a pre-trained multimodal classification model to generate the final integrated fusion result; The integrated results are structured and organized according to a preset output format, and the output is a multimodal content recognition result.
[0064] In this embodiment, the weighting coefficients for the weighted summation are determined based on the importance of the subtext in the text sequence and the confidence level of a single fusion result.
[0065] In this embodiment, the comprehensive fusion result includes the comprehensive category determination of the subject to be identified, multi-dimensional feature description, and the contribution of each sub-modal information.
[0066] In this embodiment, principal component analysis is used to reduce the dimensionality of the fused vector, remove redundant features, retain key information, and improve the processing efficiency and recognition accuracy of subsequent classification models.
[0067] In this embodiment, the pre-trained multimodal classification model refers to the model trained on a dataset constructed from massive image-text pairs (covering multiple scenarios and domains). This model is used to achieve deep interaction between image features and text features and to optimize parameters to adapt to the recognition needs of specific application scenarios.
[0068] In this embodiment, the structured output format can be customized according to actual application needs, or a recognition report with natural language description can be generated to clearly present the comprehensive recognition results of multimodal content.
[0069] In some embodiments of this application, a multimodal content recognition system based on the fusion of image and text information is also included: The acquisition module is used to acquire real-time image information and real-time text information of the subject to be identified, generate image evaluation value based on real-time image information, generate text evaluation value based on real-time text information, and set a first preprocessing strategy and a second preprocessing strategy respectively. The preprocessing module is used to preprocess real-time image information according to the first preprocessing strategy to obtain several real-time sub-images and corresponding image features, and to preprocess real-time text information according to the second preprocessing strategy to obtain several real-time sub-texts and corresponding text features. The association module is used to perform association analysis on several real-time sub-images and several real-time sub-texts, obtain a text-image mapping table based on the analysis results, and generate several single fusion results. The generation module is used to generate a comprehensive fusion result based on all individual fusion results and output the multimodal content recognition result.
[0070] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and substitutions can be made without departing from the technical principles of this application, and these improvements and substitutions should also be considered within the scope of protection of this application.
Claims
1. A multi-modal content recognition method based on image and text information fusion, characterized in that, The method comprises the following steps: acquiring real-time image information and real-time text information of a subject to be identified, generating an image evaluation value based on the real-time image information, generating a text evaluation value based on the real-time text information, and setting a first preprocessing strategy and a second preprocessing strategy respectively; preprocessing the real-time image information according to the first preprocessing strategy to obtain a plurality of real-time sub-images and corresponding image features, and preprocessing the real-time text information according to the second preprocessing strategy to obtain a plurality of real-time sub-texts and corresponding text features; performing correlation analysis on the plurality of real-time sub-images and the plurality of real-time sub-texts, obtaining a text-image mapping table based on the analysis result, and generating a plurality of single fusion results; generating a comprehensive fusion result based on all the single fusion results, and outputting a multi-modal content recognition result. 2.The multi-modal content recognition method based on image and text information fusion of claim 1, wherein, The method for generating the image evaluation value based on the real-time image information and the text evaluation value based on the real-time text information comprises the following steps: pre-setting a plurality of image evaluation indexes; generating an image sub-evaluation value of each image evaluation index based on the real-time image information, and obtaining the image evaluation value by combining the weight coefficient of the image evaluation index; pre-setting a plurality of text evaluation indexes; generating a text sub-evaluation value of each text evaluation index based on the real-time text information, and obtaining the text evaluation value by combining the weight coefficient of the text evaluation index. 3.The multi-modal content recognition method based on image and text information fusion of claim 2, wherein, The method for setting the first preprocessing strategy and the second preprocessing strategy respectively comprises the following steps: pre-setting a first preset image evaluation value interval, a second preset image evaluation value interval, a third preset image evaluation value interval, and a fourth preset image evaluation value interval; when the image evaluation value is in the first preset image evaluation value interval, selecting the fourth preset first preprocessing strategy as the first preprocessing strategy for the real-time image information; when the image evaluation value is in the second preset image evaluation value interval, selecting the third preset first preprocessing strategy as the first preprocessing strategy for the real-time image information; when the image evaluation value is in the third preset image evaluation value interval, selecting the second preset first preprocessing strategy as the first preprocessing strategy for the real-time image information; when the image evaluation value is in the fourth preset image evaluation value interval, selecting the first preset first preprocessing strategy as the first preprocessing strategy for the real-time image information; pre-setting a first preset text evaluation value interval, a second preset text evaluation value interval, a third preset text evaluation value interval, and a fourth preset text evaluation value interval; when the text evaluation value is in the first preset text evaluation value interval, selecting the fourth preset second preprocessing strategy as the second preprocessing strategy for the real-time text information; when the text evaluation value is in the second preset text evaluation value interval, selecting the third preset second preprocessing strategy as the second preprocessing strategy for the real-time text information; when the text evaluation value is in the third preset text evaluation value interval, selecting the second preset second preprocessing strategy as the second preprocessing strategy for the real-time text information; when the text evaluation value is in the fourth preset text evaluation value interval, selecting the first preset second preprocessing strategy as the second preprocessing strategy for the real-time text information. 4.The multi-modal content recognition method based on image and text information fusion of claim 3, wherein, The real-time image information is preprocessed according to the first preprocessing strategy to obtain a plurality of real-time sub-images and corresponding image features, and the real-time text information is preprocessed according to the second preprocessing strategy to obtain a plurality of real-time sub-texts and corresponding text features, comprising: The real-time image information is preprocessed according to the first preprocessing strategy to obtain a plurality of real-time sub-images and corresponding image features, and the real-time text information is preprocessed according to the second preprocessing strategy to obtain a plurality of real-time sub-texts and corresponding text features, comprising: The real-time image information is preprocessed according to the first preprocessing strategy to obtain a plurality of real-time sub-images and corresponding image features, and the real-time text information is preprocessed according to the second preprocessing strategy to obtain a plurality of real-time sub-texts and corresponding text features, comprising: The real-time image information is preprocessed according to the first preprocessing strategy to obtain a plurality of real-time sub-images and corresponding image features, and the real-time text information is preprocessed according to the second preprocessing strategy to obtain a plurality of real-time sub-texts and corresponding text features, comprising: The real-time image information is preprocessed according to the first preprocessing strategy to obtain a plurality of real-time sub-images and corresponding image features, and the real-time text information is preprocessed according to the second preprocessing strategy to obtain a plurality of real-time sub-texts and corresponding text features, comprising: 5.The multi-modal content recognition method based on image and text information fusion of claim 4, wherein, Before the correlation analysis of the plurality of real-time sub-images and the plurality of real-time sub-texts, further comprising: Correlate the image features of adjacent real-time sub-images to obtain a plurality of first correlation coefficients; Generate a first comprehensive correlation coefficient of adjacent real-time sub-images according to the plurality of first correlation coefficients and the weight coefficients of the corresponding image features; Splice the real-time sub-images whose first comprehensive correlation coefficients are greater than a preset first comprehensive correlation coefficient threshold, and fuse a plurality of image features of the same dimension of the spliced real-time sub-images to obtain a spliced sub-image and corresponding fused image features; Construct a sub-image sequence according to the spliced sub-image and the real-time sub-images that are not spliced; Correlate the text features of adjacent real-time sub-texts to obtain a plurality of second correlation coefficients; Generate a second comprehensive correlation coefficient of adjacent real-time sub-texts according to the plurality of second correlation coefficients and the weight coefficients of the corresponding text features; Splice the real-time sub-texts whose second comprehensive correlation coefficients are greater than a preset second comprehensive correlation coefficient threshold, and fuse a plurality of text features of the same dimension of the spliced real-time sub-texts to obtain a spliced sub-text and corresponding fused text features; Construct a sub-text sequence according to the spliced sub-text and the real-time sub-texts that are not spliced; Extract the keywords of each sub-text in the sub-text sequence, and generate a query condition corresponding to the sub-text, the query condition comprising whether there is an association relationship with the keywords; Based on the query condition of each sub-text, traverse each sub-image in the sub-image sequence and the corresponding image features, filter out the sub-images that satisfy the query condition of each sub-text, and calculate the corresponding first correlation degree, and construct a sub-image sub-sequence associated with each sub-text according to the first correlation degree; Wherein, the first correlation degree is calculated according to the number of keywords with an association relationship in the same sub-image, the corresponding initial correlation degree and the weight coefficient of the corresponding keyword. 6.The multi-modal content recognition method based on image and text information fusion of claim 5, wherein, The correlation analysis of the plurality of real-time sub-images and the plurality of real-time sub-texts comprises: Randomly select one sub-image in the sub-image sub-sequence associated with each sub-text, and sort the selected sub-image according to the arrangement order of the sub-texts in the sub-text sequence to obtain a to-be-analyzed sub-image sequence; Generate a plurality of to-be-analyzed sub-image sequences; Generate a second correlation degree of the corresponding to-be-analyzed sub-image sequence according to the first correlation degrees of all sub-images in the same to-be-analyzed sub-image sequence and the associated sub-texts and the weight coefficients of the corresponding sub-texts; According to the logical relationship of adjacent subtexts in the subtext sequence, a first correlation threshold between corresponding adjacent subimages in each to-be-analyzed subimage sequence is set, and a first correction coefficient is generated according to a first correlation degree of a subimage in each to-be-analyzed subimage sequence and a corresponding subtext; The first correlation threshold is corrected according to the first correction coefficient to obtain a second correlation threshold; An actual correlation value between adjacent subimages in the same to-be-analyzed subimage sequence is calculated, and the actual correlation value is subtracted from the corresponding second correlation threshold to obtain an actual correlation value difference, which is then subjected to mean value processing, quantization, and second correction coefficient generation; The second correlation degree of the corresponding to-be-analyzed subimage sequence is corrected according to the second correction coefficient to obtain a third correlation degree; Each subimage in the to-be-analyzed subimage sequence with the largest third correlation degree is set as a main subimage corresponding to each subtext to which the subimage is correlated; The main subimage is compared and analyzed with other subimages in the subimage sequence of the correlated subtext to obtain an image difference degree, and a subimage with an image difference degree greater than a preset difference threshold is set as a secondary subimage corresponding to the subtext. 7.The multi-modal content recognition method based on image and text information fusion of claim 6, wherein, A text-image mapping table is obtained according to the analysis results, including: A subtext-subimage mapping table is constructed for each subtext and main subimage and secondary subimage; The subtext-subimage mapping table includes main subimage and corresponding subtext correlated main feature information and secondary subimage and corresponding subtext correlated secondary feature information; A text-image mapping table is constructed according to the subtext-subimage mapping table of all subtexts and the arrangement order of the subtexts. 8.The multi-modal content recognition method based on image and text information fusion of claim 7, wherein, A plurality of single fusion results are generated, including: A plurality of first fusion feature information is generated for each subtext and main subimage main feature information, and a basic fusion vector is formed according to the plurality of first fusion feature information; A plurality of second fusion feature information is generated for each subtext and secondary subimage secondary feature information, and after dynamically allocating weights through an attention mechanism, the plurality of second fusion feature information is secondarily fused with the basic fusion vector to generate multi-modal fusion features; Sequence modeling is performed on all multi-modal fusion features based on a bidirectional long short-term memory network to obtain sequence fusion features; Dimension reduction and classification mapping are performed on the sequence fusion features based on a fully connected layer to output a plurality of single fusion results. 9.The multi-modal content recognition method based on image and text information fusion of claim 8, wherein, A comprehensive fusion result is generated according to all single fusion results, and a multi-modal content recognition result is output, including: A comprehensive fusion vector is obtained by weighted summation of all single fusion results corresponding to the subtexts; Principal component features are extracted by principal component analysis of the comprehensive fusion vector; The principal component features are input into a pre-trained multi-modal classification model to generate a final comprehensive fusion result; The comprehensive fusion result is structured and sorted according to a preset output format, and the sorted result is output as a multi-modal content recognition result.
10. A multi-modal content recognition system based on fusion of image and text information, characterized in that, The method includes: An acquisition module is configured to acquire real-time image information and real-time text information of a to-be-identified subject, generate an image evaluation value based on the real-time image information, generate a text evaluation value based on the real-time text information, and set a first preprocessing strategy and a second preprocessing strategy, respectively. The preprocessing module is configured to preprocess the real-time image information according to a first preprocessing strategy to obtain a plurality of real-time sub-images and corresponding image features, and preprocess the real-time text information according to a second preprocessing strategy to obtain a plurality of real-time sub-texts and corresponding text features. The association module is configured to perform association analysis on the plurality of real-time sub-images and the plurality of real-time sub-texts, obtain a text-image mapping table according to an analysis result, and generate a plurality of single fusion results. The generation module is configured to generate a comprehensive fusion result according to all the single fusion results, and output a multi-modal content recognition result.