Multi-modal data identification and analysis system based on deep learning

Through a multimodal data identification and analysis system based on deep learning, the problem of insufficient independence and semantic consistency of multimodal data processing in the prior art is solved, efficient feature fusion and alignment are achieved, and the accuracy and stability of data processing are improved.

CN120105259AInactive Publication Date: 2025-06-06SHENZHEN STANDARD & POORS CLOUD TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510172932.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology has relatively independent processing modes for heterogeneous data, lacking a unified mechanism for feature fusion between multimodals, resulting in insufficient semantic consistency of cross-modal data, low modal alignment accuracy, lack of dynamic adjustment capabilities of logical verification, insufficient adaptability and comprehensiveness of classification models, simple methods for handling conflict fields, which affect the stability and overall accuracy of data output.

Method used

The multimodal data recognition and analysis system based on deep learning is adopted to extract and align images and text features through the multimodal data preprocessing module to generate preliminary multimodal feature data; the multimodal alignment feature generation module generates aligned multimodal feature distribution values ​​through feature cross-combination and nonlinear mapping; the knowledge logic correction module performs regular checksum context updates; the classification field structure prediction module constructs a new field mapping relationship; the dynamic result optimization generation module optimizes conflict field processing.

Benefits of technology

It improves the ability to fusion between modes, enhances the accuracy of data alignment, improves the correlation and verification adaptability between fields, expands the classification capabilities, ensures the semantic integrity and consistency of the output data, and improves the accuracy, robustness and comprehensiveness of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105259A_ABST
    Figure CN120105259A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of neural networks, in particular to a deep learning-based multi-modal data recognition and analysis system, which comprises a multi-modal data preprocessing module, a multi-modal alignment feature generation module, a knowledge logic correction module, a classification field structure prediction module and a dynamic result optimization generation module. According to the method, through feature point extraction and region segmentation of the image and serialized coding processing of the text, unified representation of heterogeneous data is achieved, the capability of feature fusion between modals is improved, the semantic deficiency problem of mismatched features between the modals is solved based on feature cross combination and fitting operation, the accuracy of data alignment is enhanced, and the accuracy of data alignment is improved. Logic verification is combined with a context updating mechanism, the relevance between fields and verification adaptability are improved, the error rate caused by a single rule in the verification process is reduced, a new field mapping relation is constructed for features which are not covered in classification, the classification capacity is expanded, and the data classification meticulous performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of neural network technology, and in particular to a multimodal data recognition and analysis system based on deep learning. Background Art

[0002] The field of neural network technology includes computational models and methods used in artificial intelligence to simulate human brain neural activities. Its core content is to extract features and recognize patterns from input data through a multi-layer network structure. Neural network technology includes various model types such as convolutional neural networks, recursive neural networks, and generative adversarial networks, and is widely used in image recognition, natural language processing, speech recognition, and recommendation systems. The focus of this technical field is to use mathematical models and algorithms to represent, classify, and analyze complex data, specifically through parameter optimization and training methods in a hierarchical structure, so as to achieve modeling and solving nonlinear problems.

[0003] Among them, the multimodal data recognition and analysis system refers to the unified processing and analysis of data from different sources or types based on neural network technology. This type of system is usually completed by multimodal feature extraction, inter-modal correlation learning, and data fusion modeling for the fusion recognition problem of heterogeneous data such as images, texts, and voices. Specifically, the system constructs a cross-modal feature sharing network and a modal correlation learning network, uses a deep neural network to extract the representation features of different modal data, and combines the modal alignment method to achieve the consistency of semantic information between data. Intelligent multimodal enhanced recognition (AMMR, Augmented Multimodal Recognition) technology combines multiple recognition modes, such as vision, sound, touch, etc., to improve the accuracy and robustness of recognition. This technology can effectively make up for the limitations of a single modality in a specific environment when processing heterogeneous data by fusing the features of different modalities, thereby improving the overall recognition performance. The advantage of AMMR is that it can solve the challenges of insufficient single data source or noise interference in traditional models through the mutual complementation between multiple modalities, thereby improving the stability and accuracy of the system.

[0004] The existing technologies have relatively independent processing modes for heterogeneous data and lack a unified mechanism for multi-modal feature fusion, resulting in insufficient semantic consistency of cross-modal data. In the process of modal alignment, a single matching strategy is difficult to handle mismatched features between modalities, which weakens the alignment accuracy. The lack of dynamic adjustment capabilities in logical verification can easily lead to verification failures or erroneous updates due to the limited scope of application of rules. In the classification stage, there is a lack of expansion mechanism for uncovered features, which reduces the adaptability and comprehensiveness of the classification model. In the result generation, the conflict field processing method is simple, and the field interaction characteristics cannot be fully optimized, which affects the stability and overall accuracy of data output. Summary of the invention

[0005] The purpose of the present invention is to solve the shortcomings existing in the prior art and propose a multimodal data recognition and analysis system based on deep learning.

[0006] In order to achieve the above object, the present invention adopts the following technical solution: A multimodal data recognition and analysis system based on deep learning includes:

[0007] The multimodal data preprocessing module extracts feature points and performs region segmentation operations on the pixel matrix in the product image based on the input product image and the text information it contains, performs serialization encoding processing on the continuous character area in the text information, extracts features from the image and text respectively, and aligns the joint embedding space so that the two modal data can be associated in the same representation space to establish preliminary multimodal feature data;

[0008] The multimodal alignment feature generation module matches the extracted local area features of the image with the text feature vector based on the preliminary multimodal feature data, calculates the feature similarity, cross-combines the features for the mismatched parts through the joint embedding representation and performs nonlinear mapping operations, and generates an aligned multimodal feature distribution value after screening and discriminating the generated data;

[0009] The knowledge logic correction module performs rule verification and adjusts the range of the logic nodes in the field distribution value of the specification parameter based on the aligned multimodal feature distribution value, ensures that the numerical field in the specification parameter meets the preset logical constraints, replaces the node value that fails the verification according to the association rule, performs field update and syntax verification in combination with the context conditions, and generates a semantic logic adjustment value;

[0010] The classification field structure prediction module extracts the field logic relationship and groups it based on the semantic logic adjustment value, performs item-by-item comparison on the key feature values ​​screened in the commodity classification field, constructs a new field mapping relationship for the feature values ​​not covered by the classification standard, and generates a feature field classification prediction value;

[0011] The dynamic result optimization generation module performs priority assignment operations on conflicting field values ​​in the recommended tags based on the classification prediction values ​​of the feature fields, builds mapping relationships after rearranging the fields, and dynamically updates and merges the field interaction characteristics to generate optimized structured output data.

[0012] The multimodal feature data specifically includes image feature point distribution values, image area segmentation values, and text serialization encoding values; the multimodal alignment feature distribution values ​​specifically include image character area alignment values, text sequence matching values, and feature cross-combination values; the semantic logic adjustment values ​​specifically include field logic check values, field update values, and syntax check values; the feature field classification prediction values ​​specifically include grouping field logic relationship values, new field mapping relationship values, and classification key feature values; the optimized structured output data includes field priority assignment values, field interaction characteristic update values, and field rearrangement mapping relationship values.

[0013] As a further solution of the present invention, the step of acquiring the preliminary multimodal feature data is specifically as follows:

[0014] Extract a pixel matrix from the input image data, calculate the absolute value of the gradient change of each pixel point by performing point-by-point operation and comparison on the gradient change value of each pixel point in the pixel matrix, compare the value with the set threshold, mark the pixel points whose absolute value of the change is greater than the threshold, summarize the marked data after marking, and generate a feature point distribution matrix;

[0015] Based on the feature point distribution matrix, the ratio of the local mean of the grayscale value of the pixel matrix to the variance of the pixel value is calculated, the area where the ratio value changes significantly is screened, the boundary of each area that satisfies the ratio change is divided, and the identification classification after the area division is generated. After the classification, the boundary features and grayscale value parameters of each area are counted to generate a set of image area feature parameters;

[0016] Perform serialization encoding processing on the continuous character area of ​​the text data, calculate the recurrence frequency of the character segment, sort each segment according to the frequency and length pattern, group the sorted character segments, perform operations on each group of character segments, record the sequence pattern and feature parameters of each group, and generate a text sequence feature parameter set;

[0017] The image region feature parameter set and the text sequence feature parameter set are called to calculate the numerical relationship of the cross-comparison of the two sets of fields, and the formula is adopted by adjusting the field key weight:

[0018]

[0019] Calculate preliminary multimodal feature data D m ;

[0020] Among them, w p Represents the critical weight of the image field, P g,p Represents the image region feature parameter value, P t,q represents the characteristic parameter value of the text sequence, p represents the pth characteristic field, N g Represents the total number of feature fields in the image region.

[0021] As a further solution of the present invention, the step of obtaining the aligned multimodal feature distribution value is specifically as follows:

[0022] Extracting image character regions and text sequence values ​​from the preliminary multimodal feature data, matching image fields and text fields pair by pair according to the grouping rule, calculating the fit value between each pair of fields, screening and recording all field pairs whose fit values ​​are lower than a set threshold, and generating a set of unmatched field pairs;

[0023] Based on the set of unmatched field pairs, cross-combination is performed for the corresponding characteristic parameters in each pair of fields, the weight coefficient of each parameter after combination is adjusted according to the field criticality weight, the weighted fitting value of the combined parameter is calculated, the difference between the fitting value and the original field parameter is corrected, and a set of field characteristic cross-combination fitting values ​​is generated;

[0024] The field feature cross-combination fitting value set is called, and the field screening results are judged in combination with the field criticality weight and matching degree, and the field feature difference and dynamic adjustment threshold are used for screening, using the formula:

[0025]

[0026] Calculate the aligned multimodal feature distribution value F d ;

[0027] Among them, w p Represents the criticality weight of the field, C g,p Represents the characteristic value of the image field, C t,q Represents the characteristic value of the text field, p represents the pth field, N f Represents the total number of feature fields.

[0028] As a further solution of the present invention, the step of obtaining the semantic logic adjustment value is specifically as follows:

[0029] Extracting logical nodes in the field distribution value from the aligned multimodal feature distribution value, for each field value of the logical node, gradually performing a determination of a numerical range according to a rule check, marking and recording logical nodes that fail to pass the check during the determination process, and generating a set of logical node values ​​that fail to pass the check;

[0030] Based on the set of logical node values ​​that have not passed the verification, calling the association rule corresponding to each of the logical nodes, calculating the difference parameters between the logical node value and the target value of the association rule item by item, performing operations on the difference parameters to obtain replacement values, and updating the original logical node value with the replacement value to generate a set of field logical node correction values;

[0031] The field logic node correction value set is called, the field value of the dependent logic node is dynamically updated in combination with the context conditions, and the context logic conditions that the field depends on are checked one by one, using the formula:

[0032]

[0033] Calculate the semantic logic adjustment value L s ;

[0034] Among them, w s Represents the critical weight of the logical node, R s represents the target value of the association rule, T s Represents the current value of the logical node, D s Represents the context dependency parameter of field update, S s Represents the syntax dependency parameter of the field value, s represents the sth logical node, N l Represents the total number of logical nodes.

[0035] As a further solution of the present invention, the step of obtaining the classification prediction value of the feature field is specifically as follows:

[0036] Extracting the field logical relationship from the semantic logic adjustment value, gradually determining the logical relevance between the fields according to the associated characteristic parameters between the fields, classifying and grouping the fields whose logical relevance meets the same characteristic parameters, and generating a field logical grouping result;

[0037] Based on the field logic grouping result, for the key feature values ​​selected in the grouping field, the deviation parameter of each feature value and the feature mean in the group is calculated item by item, the field feature values ​​whose deviation parameters exceed the set range are recorded, and the results are summarized to form an uncovered feature value set;

[0038] The uncovered feature value set is called, the field mapping relationship is rebuilt according to the logical grouping and deviation parameters, and the field feature value is dynamically updated on the mapping relationship, using the formula:

[0039]

[0040] Calculate the classification prediction value F of the feature field c ;

[0041] Among them, w r Represents the key weight of the field, T r Represents the current value of the feature field, M r Represents the target value of the feature field mapping, P r Represents the logical grouping deviation factor, Q r represents the eigenvalue correlation factor, r represents the rth field, N f Represents the total number of fields.

[0042] As a further solution of the present invention, the step of obtaining the optimized structured output data is specifically as follows:

[0043] Filtering conflicting field values ​​from the characteristic field classification prediction values, performing priority parameter determination for each conflicting field value, determining whether to retain or replace the field value by comparing the priority parameter size of the field value, recording the processed field value, and generating a priority allocation field result;

[0044] Based on the priority allocation field result, the processed field values ​​are rearranged, and the field arrangement order is adjusted and an updated field mapping relationship is constructed in combination with the logical association between field features and the grouping mapping rules, so as to generate a field mapping relationship result;

[0045] The field mapping relationship result is called, the field interaction characteristics are dynamically updated, and the interaction parameters between the fields are determined item by item. The optimization value is constructed based on the collaborative characteristics between the fields, using the formula:

[0046]

[0047] Compute optimized structured output data O s ;

[0048] Among them, w r,s Represents the field interaction weight parameter, I r,s Represents the actual value of the interaction between fields, A r,s represents the interaction target value between fields, R r,s represents the field coordination adjustment ratio, C r,s Represents the field associated logical parameter, E r,s represents the field dynamic deviation factor, P r,s represents the interaction parameter between fields, Q r,s represents the field interaction feature correction parameter, r and s represent the field number, N f Represents the total number of fields.

[0049] Compared with the prior art, the advantages and positive effects of the present invention are:

[0050] In the present invention, a unified representation of heterogeneous data is achieved by extracting feature points and segmenting regions of images, and serializing and encoding text, thereby improving the ability to fuse features between modalities. Based on feature cross-combination and fitting operations, the semantic loss problem of mismatched features between modalities is solved, and the accuracy of data alignment is enhanced. Logical verification is combined with a context update mechanism to improve the correlation between fields and the adaptability of verification, and reduce the error rate caused by a single rule during the verification process. For features not covered in the classification, a new field mapping relationship is constructed to expand the classification capability and improve the meticulousness of data classification. The dynamic update mechanism optimizes the processing of conflicting fields to ensure the semantic integrity and consistency of the output data structure. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 is a system flow chart of the present invention;

[0052] Figure 2 This is a flow chart of the steps for obtaining preliminary multimodal feature data of the present invention;

[0053] Figure 3 A flowchart of the steps for obtaining the distribution value of the aligned multimodal features of the present invention;

[0054] Figure 4 A flowchart of the steps for obtaining the semantic logic adjustment value of the present invention;

[0055] Figure 5 A flowchart of the steps for obtaining the classification prediction value of the feature field of the present invention;

[0056] Figure 6 Flow chart of the steps for obtaining structured output data optimized for the present invention. DETAILED DESCRIPTION

[0057] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0058] In the description of the present invention, it should be understood that the terms "length", "width", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, in the description of the present invention, "multiple" means two or more, unless otherwise clearly and specifically defined.

[0059] Embodiment 1

[0060] See also Figure 1 , the multimodal data recognition and analysis system based on deep learning includes:

[0061] The multimodal data preprocessing module extracts feature points and performs region segmentation operations on the pixel matrix in the product image based on the input product image and the text information it contains, performs serialization encoding processing on the continuous character area in the text information, extracts features from the image and text respectively, and aligns the joint embedding space so that the two modal data can be associated in the same representation space to establish preliminary multimodal feature data;

[0062] This module can efficiently extract key information from images and texts by performing feature point extraction and region segmentation operations on the pixel matrix of product images, as well as serialization and encoding processing of character regions in text information. This provides accurate preliminary data support for subsequent multimodal data fusion. The role of AMMR in this link is reflected in its ability to fuse features between images and texts. It can intelligently identify the correlation between product images and the text information they contain, and improve the accuracy of preliminary feature extraction.

[0063] The multimodal alignment feature generation module matches the extracted local area features of the image with the text feature vector based on the preliminary multimodal feature data, calculates the feature similarity, cross-combines the features for the mismatched parts through the joint embedding representation and performs nonlinear mapping operations, and generates an aligned multimodal feature distribution value after screening and discriminating the generated data;

[0064] By comparing the character area and text sequence values ​​of the image-text combination area, and performing feature cross-combination and fitting on the mismatched parts, high-precision alignment between the product image and the text information it contains can be effectively achieved. The successful application of this module helps to generate more consistent and complete multimodal feature data. The role of AMMR technology at this stage is to enhance the robustness of feature alignment and improve the mismatch handling capabilities between modalities, thereby improving alignment accuracy and consistency of cross-modal data.

[0065] The knowledge logic correction module performs rule verification and adjusts the range of the logic nodes in the field distribution value of the specification parameter based on the aligned multimodal feature distribution value, ensures that the numerical field in the specification parameter meets the preset logical constraints, replaces the node value that fails the verification according to the association rule, performs field update and syntax verification in combination with the context conditions, and generates a semantic logic adjustment value;

[0066] This module can ensure the rationality of data logic by performing logical verification and adjustment on the aligned feature data, and perform necessary grammatical verification according to contextual conditions to generate logical adjustment values ​​that conform to semantics. Through this correction, the logical consistency and accuracy of the data are significantly improved. AMMR plays a key role in this module. By dynamically adapting to different data sources, it ensures that the logical relationship between features can be effectively corrected across modalities, especially in the adjustment of semantic associations between images and texts.

[0067] The classification field structure prediction module extracts the field logic relationship and groups it based on the semantic logic adjustment value, performs item-by-item comparison on the key feature values ​​screened in the commodity classification field, constructs a new field mapping relationship for the feature values ​​not covered by the classification standard, and generates a feature field classification prediction value;

[0068] By extracting semantic logic adjustment values ​​and grouping field logical relationships, the feature values ​​in the product classification fields are compared item by item, which can not only accurately screen out the most critical features, but also deal with feature values ​​not covered by the classification standards. AMMR technology improves the ability to process multimodal features in classification prediction, especially when data sources such as images and texts are not synchronized. It can establish new field mapping relationships through intelligent recognition, improving the comprehensiveness and accuracy of classification predictions.

[0069] The dynamic result optimization generation module performs priority assignment operations on conflicting field values ​​in the recommended tags based on the classification prediction values ​​of the feature fields, builds mapping relationships after rearranging the fields, and dynamically updates and merges the field interaction characteristics to generate optimized structured output data.

[0070] This module prioritizes and dynamically updates conflicting field values ​​in recommendation tags, and generates more stable and accurate structured output data by optimizing field interaction characteristics. The optimization effect of this stage helps to improve the overall response efficiency of the system and the stability of output data. The application of AMMR further enhances the cross-modal feature fusion in the result optimization process, especially when dealing with conflicting fields, which can better analyze the interaction relationship between data of each modality, thereby achieving dynamic update and optimization, and improving the robustness and stability of the final result.

[0071] In the entire multimodal data recognition and analysis system based on deep learning, the introduction of AMMR (intelligent multimodal enhanced recognition) technology has significantly improved the system's performance in various stages, including the fusion of product images and the text information they contain, inter-modal alignment, data correction, classification prediction, and result optimization. Especially when processing heterogeneous data, AMMR can solve the problems of multimodal data inconsistency, information loss, and low processing accuracy faced by traditional methods through intelligent enhancement and complementarity between multimodal features, thereby greatly improving the accuracy, robustness, and comprehensiveness of the system.

[0072] Multimodal feature data specifically include image feature point distribution values, image area segmentation values, and text serialization encoding values. Multimodal alignment feature distribution values ​​specifically include image character area alignment values, text sequence matching values, and feature cross-combination values. Semantic logic adjustment values ​​specifically include field logic check values, field update values, and syntax check values. Feature field classification prediction values ​​specifically include grouping field logical relationship values, new field mapping relationship values, and classification key feature values. Optimized structured output data includes field priority assignment values, field interaction feature update values, and field rearrangement mapping relationship values.

[0073] See also Figure 2 ,The specific steps for obtaining preliminary multimodal feature data are as follows:

[0074] Extract the pixel matrix from the input image data, calculate the absolute value of the gradient change of each pixel point by performing point-by-point operations and comparisons on the gradient change value of each pixel point in the pixel matrix, compare the value with the set threshold, mark the pixels whose absolute value of the change is greater than the threshold, and after marking, summarize the marked data to generate a feature point distribution matrix;

[0075] To extract the pixel matrix from the input image data, it is necessary to first parse the format of the input image, such as JPEG, PNG or BMP, and obtain the data matrix of the RGB channel or grayscale channel of the image through the decoding algorithm. Then, traverse all the pixels of the matrix, and for each pixel, calculate the gradient change value between it and the adjacent pixels. Use the Sobel operator or Prewitt operator to calculate the gradient components in the horizontal and vertical directions, which are denoted as G respectively. x and G y , by calculating the gradient amplitude Get the absolute value of the gradient change of each pixel, set the threshold T to compare the gradient amplitude of each pixel, if the G of a pixel is greater than T, mark the pixel as an edge pixel to form a binary marking matrix. After marking, count the coordinate information of all marked pixels and summarize them to form a feature point distribution matrix. For example, assuming that the input image is a grayscale image of 512×512 pixels, randomly select a region and the pixel value is:

[0076] Table 1 Example of image pixel gradient calculation

[0077] Pixels Gray value <![CDATA[G x ]]> <![CDATA[G y ]]> Gradient amplitude G Whether to mark (10,10) 120 10 15 18.03 yes (10,11) 125 8 12 14.42 no (10,12) 130 15 20 25.00 yes

[0078] As shown in Table 1, assuming that the threshold T=15 is set, the gradient amplitudes of the pixels (10,10) and (10,12) are greater than 15, so they are marked, and the final feature point distribution matrix is ​​used for subsequent analysis.

[0079] Based on the feature point distribution matrix, the ratio of the local mean of the grayscale value of the pixel matrix to the pixel value variance is calculated, and the areas with significant ratio value changes are screened. The boundaries of each area that meets the ratio change are divided, and the identification classification after the area division is generated. After classification, the boundary features and grayscale value parameters of each area are counted to generate a set of image area feature parameters;

[0080] Based on the feature point distribution matrix, the ratio of the local mean of the grayscale value of the pixel matrix to the pixel value variance is calculated. The pixel matrix needs to be divided into multiple regions. The size of each region can be set to n×n pixels. The grayscale mean μ and variance σ of all pixels in each region are calculated. 2 , the mean calculation formula is as follows:

[0081]

[0082] The variance calculation formula is as follows:

[0083]

[0084] Then calculate the ratio:

[0085]

[0086] For R greater than a certain threshold T R The area is screened and the boundaries are divided to form independent areas. When n=4, the pixel data in a certain area is as follows:

[0087] Table 2 Image region grayscale feature calculation

[0088] Pixels Gray value (1,1) 120 (1,2) 130 (2,1) 140 (2,2) 150

[0089] Calculated:

[0090]

[0091] ratio:

[0092]

[0093] If T is set R =0.9, the region is selected as the partition region and a set of regional feature parameters is generated.

[0094] Perform serialization encoding processing on the continuous character area of ​​the text data, calculate the recurrence frequency of the character segment, sort each segment according to the frequency and length pattern, group the sorted character segments, perform operations on each group of character segments, record the sequence pattern and feature parameters of each group, and generate a text sequence feature parameter set;

[0095] Perform serialization encoding processing on the continuous character area of ​​the text data. First, extract the character sequence in the text data and store it according to the character encoding. For example, for the text data "AABBBCCCAA", according to the character frequency statistics, obtain the character segments "AA", "BBB", "CCC", "AA", record the number of repetitions and length patterns of each character segment, and calculate the character segment frequency:

[0096] Assuming the total length of the text is 10, calculate the frequency of each character segment:

[0097] Table 3 Text character frequency calculation

[0098] Character segment length frequency AA 2 0.2 BBB 3 0.3 CCC 3 0.3 AA 2 0.2

[0099] After sorting by frequency, each group of character segments is operated, the character sequence pattern is recorded, and a set of text sequence feature parameters is generated.

[0100] Call the image region feature parameter set and the text sequence feature parameter set, calculate the numerical relationship of the cross-comparison of the two set fields, and adjust the field criticality weight using the formula:

[0101]

[0102] Calculate preliminary multimodal feature data D m ;

[0103] Among them, w p Represents the critical weight of the image field, P g,p Represents the image region feature parameter value, P t,q represents the characteristic parameter value of the text sequence, p represents the pth characteristic field, N g Represents the total number of feature fields in the image region.

[0104] Call the image region feature parameter set and the text sequence feature parameter set, calculate the numerical relationship between the cross-comparison of the two set fields, and set the feature field weight w i are 0.5, 0.3, and 0.2 respectively, corresponding to the eigenvalue P gi and P ti as follows:

[0105] Table 4 Image and text feature parameters

[0106] Feature fields <![CDATA[w p ]]> <![CDATA[P g,p ]]> <![CDATA[P t,q ]]> F1 0.5 0.8 0.7 F2 0.3 0.6 0.5 F3 0.2 0.7 0.6

[0107] Substitute into the calculation formula:

[0108]

[0109] calculate:

[0110]

[0111] Bring in:

[0112]

[0113] This result shows that the calculated multimodal feature data is 0.95 and can be used for subsequent analysis.

[0114] See also Figure 3 , the steps for obtaining the distribution value of aligned multimodal features are as follows:

[0115] Extracting image character regions and text sequence values ​​from preliminary multimodal feature data, matching image fields and text fields pair by pair according to grouping rules, calculating the fit value between each pair of fields, screening and recording all field pairs with fit values ​​lower than a set threshold, and generating a set of unmatched field pairs;

[0116] To extract image character regions and text sequence values ​​from preliminary multimodal feature data, it is necessary to first parse the image field and text field information stored in the multimodal data. During the image field extraction process, key parameters such as grayscale value, edge features, and contour shape of a specific area are obtained and classified and stored in the corresponding feature matrix. The text field is parsed through character encoding to extract parameters such as character sequence, repetitive pattern, and character spacing, and a text feature matrix is ​​constructed. Subsequently, the image field and the text field are matched according to the grouping rules. The grouping rules can be set according to the feature category, such as mapping and matching according to the character shape, size, and arrangement. The fit value is calculated for each pair of matching fields. During the fit calculation process, the similarity score between the two fields is first calculated. The score can be calculated by the Euclidean distance of the field numerical parameters. For example, the image field feature value vector G is set to [g 1 ,g 2 ,...,g n ] and text field feature value vector T =

[0117] [t 1 ,t 2 ,...,t n ], calculate the Euclidean distance between the two:

[0118]

[0119] Then all calculated fitness values ​​and the set threshold T f Compare, if the fit value of a field pair is lower than T f , it is determined as an unmatched field pair, and its field number and corresponding feature value are recorded, and finally a set of unmatched field pairs is formed. For example, set T f =0.8, if the calculated fit values ​​of some field pairs are 0.72, 0.65, and 0.79 respectively, these field pairs are marked as mismatched field pairs, as shown in Table 5.

[0120] Table 5 Examples of mismatched field pairs

[0121] Field Number Image field eigenvalue G Text field feature value T Goodness of fit F1 0.85 0.70 0.72 F2 0.60 0.45 0.65 F3 0.90 0.80 0.79

[0122] As shown in Table 5, all field pairs with a fit lower than 0.8 were screened and recorded to generate a set of unmatched field pairs.

[0123] Based on the set of unmatched field pairs, a cross combination is performed for the corresponding characteristic parameters in each pair of fields. The weight coefficient of each parameter after the combination is adjusted according to the field criticality weight, and the weighted fitting value of the combined parameter is calculated. The difference between the fitting value and the original field parameter is corrected to generate a set of field feature cross combination fitting values.

[0124] Based on the set of unmatched field pairs, the corresponding feature parameters in each pair of fields are cross-combined. First, for each pair of unmatched fields, the key feature parameters of the image field and the text field are selected and cross-combined. During the cross-combination process, each image field parameter can be combined with multiple text field parameters. For example, assuming that the image field feature parameter contains the shape factor G s , Contrast G c and edge strength G e , while the text field characteristic parameters include the character spacing T d , stroke complexity T b and character width T w , then the combination is as follows: (G s ,T d ),(G c ,T b ),(G e ,T w );

[0125] Then, the weight coefficient of each parameter is adjusted according to the key weight of the field. For example, the key weights are set to w s =0.4, w c =0.35, w e =0.25, then calculate the weighted combination value: C i =ws ·|G s -T d |+w c ·|G c -T b |+w e ·|G e -T w |;

[0126] After calculating the weighted fitting value of each combination, the difference between it and the original field parameter is corrected. For example, if the calculated combined fitting value is 0.75, and the original field parameter fitting value is 0.60, it needs to be corrected and adjusted to reduce the difference, and finally a set of field feature cross-combination fitting values ​​is generated, as shown in Table 6.

[0127] Table 6 Cross-combination fitting value calculation

[0128] Combination Fields Original Fit Value Weighted Fitted Values Difference (G_s,T_d) 0.60 0.75 0.15 (G_c,T_b) 0.55 0.72 0.17 (G_e,T_w) 0.58 0.70 0.12

[0129] As shown in Table 6, the calculated cross-combination fitting value is improved compared with the original fitting value, and the difference value is recorded for subsequent matching screening.

[0130] Call the field feature cross-combination fitting value set, combine the field critical weight and matching degree to judge the field screening results, and screen according to the field feature difference and dynamic adjustment threshold, using the formula:

[0131]

[0132] Calculate the aligned multimodal feature distribution value F d ;

[0133] Among them, w p Represents the criticality weight of the field, C g,p Represents the characteristic value of the image field, C t,q Represents the characteristic value of the text field, p represents the pth field, N f Represents the total number of feature fields.

[0134] Call the field feature cross-combination fitting value set, and judge the field screening results by combining the field critical weight and matching degree. First, compare the corrected fitting value and field critical weight to screen the field differences. For example, set the dynamic adjustment threshold T d Used to filter field combinations with large feature differences. d =0.1, then all field combinations with difference values ​​greater than 0.1 need to be corrected and screened, and then the final matching value F is calculated. d :

[0135]

[0136] Among them, w p Represents the criticality weight of the field, C g,p Represents the characteristic value of the image field, C t,q

[0137] Represents the characteristic value of the text field, p represents the pth field, N f Represents the total number of feature fields. The calculation example is as follows:

[0138] C g1 =0.75,C t1 =0.65,w 1 =0.4;

[0139] C g2 =0.72,C t2 =0.60,w 2 =0.35;

[0140] C g3 =0.70,C t3 =0.58,w 3 =0.25;

[0141] Substitute into the calculation:

[0142]

[0143] Table 7 Matching calculation results

[0144] Fields <![CDATA[Image eigenvalue C g,p > <![CDATA[Text feature value C t,q > <![CDATA[Critical weight w p > <![CDATA[Match degree calculation F d > F1 0.75 0.65 0.4 0.126 F2 0.72 0.60 0.35 0.121 F3 0.70 0.58 0.25 0.087

[0145] The final calculated aligned multimodal feature distribution value F d =0.334, which is used for matching judgment and further optimization and adjustment.

[0146] See also Figure 4 , the specific steps for obtaining the semantic logic adjustment value are:

[0147] Extracting logical nodes in the field distribution value from the aligned multimodal feature distribution value, for the field value of each logical node, gradually performing the determination of the value range according to the rule verification, marking and recording the logical nodes that failed the verification during the determination process, and generating a set of logical node values ​​that failed the verification;

[0148] To extract the logical nodes in the field distribution value from the aligned multimodal feature distribution value, you first need to obtain the original multimodal feature data, such as structured and unstructured information from sources such as text, images, audio or video. These data are converted into a standardized distribution value set through a specific format. Then, you need to perform logical node extraction on the data set. The logical node is a key data point used to represent the correspondence between field values ​​in different modes. The field value of each logical node can be represented numerically. For example, text data can be converted through word vectors, and image data can be represented by pixel features or CNN extraction vectors. For each logical node field value, its rationality needs to be verified one by one according to the rules. Here, multiple verification criteria can be set. For example, the field value should meet a specific range (such as a normalized value between 0 and 1) or conform to a certain distribution characteristic (such as normal distribution). During the verification process, if a field value does not meet the set standard, it needs to be marked and the relevant information of the logical node is recorded to form a set of logical node values ​​that have not passed the verification. For example, if a logical node represents a word vector of text data, and the value of a certain dimension exceeds the set standard deviation threshold, the logical node needs to be marked and added to the set that has not passed the verification to ensure that the data quality is controllable. Finally, all logical node values ​​that have not passed the verification will form a set for adjustment and optimization in subsequent steps.

[0149] Based on the set of logical node values ​​that have not passed the verification, the association rule corresponding to each logical node is called, the difference parameters between the logical node value and the target value of the association rule are calculated item by item, the difference parameters are operated to obtain the replacement value, and the original logical node value is updated with the replacement value to generate a set of field logical node correction values;

[0150] Based on the set of logical node values ​​that failed the verification, the association rules corresponding to each logical node are called, and the difference parameters are calculated for the field values ​​that failed the verification. The difference parameters here can be calculated using absolute difference, normalized error or weighted deviation. For each logical node field value T i and its association rule target value R i , calculate its difference parameter |R i -T i |, and then perform operations on the difference parameters to obtain the corrected value. For example, for a text data field, assuming its word vector value T i =0.75, and the association rule target value R i = 0.5, then the difference parameter is |0.75-0.5| = 0.25. If the weighted adjustment method is used for correction, a weight factor w can be set i And calculate the new field correction value, which can be specifically calculated using the correction formula: If you set w i =0.6, then the corrected Finally, a set of field logic node correction values ​​is obtained, which will be used for dynamic adjustment of subsequent logic node dependencies.

[0151] Call the field logic node correction value set, dynamically update the field value of the dependent logic node in combination with the context conditions, and check the contextual logic conditions that the field depends on one by one, using the formula:

[0152]

[0153] Calculate the semantic logic adjustment value L s ;

[0154] Among them, w s Represents the critical weight of the logical node, R s represents the target value of the association rule, T s Represents the current value of the logical node, D s Represents the context dependency parameter of field update, S s Represents the syntax dependency parameter of the field value, s represents the sth logical node, N l Represents the total number of logical nodes.

[0155] Call the field logic node to correct the value set, and dynamically update the field value of the dependent logic node in combination with the context conditions. First, determine the dependency of the logic node. The dependency can be determined by the association structure between the fields. For example, the adjustment of a field value may affect the calculation of another field. The contextual logic conditions that the field depends on are checked one by one. Here, a weighted update method can be used. The formula is as follows:

[0156]

[0157] Among them, L s Represents the semantic logic adjustment value, w i is the critical weight of the logical node, R i is the target value of the association rule, T i is the current value of the logical node, D i is the context dependency parameter of the field update, S i It is the syntax dependency parameter of the field value, i represents the i-th logical node, and n represents the total number of logical nodes.

[0158] Formula parameter calculation: 1. Set the logical node weight w i , assuming the weight value is: w = [0.5, 0.7, 0.6, 0.8];

[0159] 2. Set the field value T i :T=[0.75,0.82,0.68,0.59];

[0160] 3. Set the association rule target value R i : R = [0.5, 0.78, 0.7, 0.65];

[0161] 4. Calculate the difference parameter |R i -T i |: |RT|=[0.25,0.04,0.02,0.06];

[0162] 5. Set the context dependency of the field D i :D=[0.4,0.5,0.6,0.7];

[0163] 6. Set the syntax dependency S of the field i : S = [0.3, 0.2, 0.4, 0.5];

[0164] 7. Calculate the sum of squares in the formula:

[0165] 8. Calculate the molecular part:

[0166]

[0167] 9. Calculate the denominator:

[0168] ∑w i =0.5+0.7+0.6+0.8=2.6;

[0169]

[0170] Formula analysis: The innovation of this formula is that it combines the context dependency D of the logic node i and grammatical dependency S i To calculate the weighted adjustment value, compared with the traditional mean correction method, this formula can adaptively adjust the influence between different logical nodes, making the overall adjustment more targeted, and ensure that the data is updated within a given range during the calculation process, and finally obtain the semantic logic adjustment value L s =0.627.

[0171] See also Figure 5 , the specific steps for obtaining the classification prediction value of the feature field are:

[0172] Extract the field logical relationship from the semantic logic adjustment value, gradually determine the logical correlation between the fields according to the associated characteristic parameters between the fields, classify and group the fields whose logical correlation meets the same characteristic parameters, and generate the field logical grouping result;

[0173] To extract the field logical relationship from the semantic logic adjustment value, we first need to obtain the feature values ​​of each field. These feature values ​​may come from data types such as text, numerical values, and images, and are converted into structured data through specific data processing methods. For example, text data can be processed through word vectorization, numerical data can be used directly, and image data can be represented by vectors extracted through pixel features or convolutional neural networks. After obtaining the structured feature values, calculations are performed based on the associated feature parameters between fields. The correlation between fields needs to be evaluated during calculations. Pearson correlation coefficient, mutual information, or covariance analysis can usually be used. The Pearson correlation coefficient is calculated as follows:

[0174]

[0175] in, and are the means of field A and field B respectively. After calculation, if the correlation coefficient r>0.6, it means that there is a strong positive correlation between the two fields. If r<-0.6, it means that the negative correlation is strong. If r is between -0.3 and 0.3, the correlation is weak. Based on the calculated field correlation, the fields with logical correlations that meet the same characteristic parameters are classified and grouped. For example, in the financial data analysis scenario, sales, profits, and taxes may belong to the same logical grouping, while the number of employees and average salary may belong to another logical grouping. After all fields are grouped, the field logical grouping results are finally generated.

[0176] Based on the field logic grouping results, for the key feature values ​​selected in the grouping fields, the deviation parameters of each feature value and the feature mean in the group are calculated item by item, and the field feature values ​​whose deviation parameters exceed the set range are recorded and summarized to form an uncovered feature value set;

[0177] Based on the field logic grouping results, the deviation parameters of each eigenvalue and the eigenvalue mean in the grouping field are calculated item by item for the key eigenvalues ​​in the grouping field. The eigenvalue mean of each group is first obtained during the calculation:

[0178]

[0179] Among them, M g represents the mean of group g, T j represents the characteristic value of the jth field in the group, m is the total number of fields in the group, and after calculating the mean, the deviation parameter of each field is calculated:

[0180] D i =|T i -M g |;

[0181] If the deviation parameter D iIf the value exceeds the set threshold, the field value is recorded as an uncovered feature value. The threshold can be set to 1.5 times the standard deviation of the mean:

[0182] Threshold=M g +1.5×σ;

[0183] Among them, σ is the standard deviation, assuming that the group mean M g is 100 and the standard deviation σ is 15, then the threshold is calculated as follows:

[0184] Threshold=100+1.5×15=122.5;

[0185] If a field value T i is 130, then the deviation D i =130-100=30, if it exceeds the threshold, the field is classified into the uncovered feature value set. Finally, all uncovered feature values ​​are aggregated to form an uncovered feature value set. Table 8 lists an example data set containing the calculated deviation parameters and standard deviation thresholds:

[0186] Table 8 Field characteristic value deviation calculation table

[0187] Field Name <![CDATA[Eigenvalue T i > <![CDATA[Within-group mean M g > <![CDATA[Deviation D i > Standard deviation Threshold A 130 100 30 15 122.5 B 95 100 5 15 122.5 C 110 100 10 15 122.5 D 80 100 20 15 122.5

[0188] As shown in Table 8, the deviations of field A and field D exceed the set threshold and are therefore classified into the uncovered feature value set.

[0189] Call the uncovered feature value set, rebuild the field mapping relationship based on the logical grouping and deviation parameters, and dynamically update the field feature values ​​of the mapping relationship using the formula:

[0190]

[0191] Calculate the classification prediction value F of the feature field c ;

[0192] Among them, w r Represents the key weight of the field, T r Represents the current value of the feature field, M r Represents the target value of the feature field mapping, P r Represents the logical grouping deviation factor, Q r represents the eigenvalue correlation factor, r represents the rth field, N f Represents the total number of fields.

[0193] Call the uncovered feature value set, rebuild the field mapping relationship based on the logical grouping and deviation parameters, and dynamically update the field mapping relationship by calculating the field deviation and critical weight. The calculation formula is as follows:

[0194]

[0195] Among them, w r Represents the key weight of the field, T r Represents the current value of the feature field, M r

[0196] Represents the target value of the feature field mapping, P r Represents the logical grouping deviation factor, Q r

[0197] represents the eigenvalue correlation factor, r represents the rth field, N f Represents the total number of fields. Table 9 lists the example fields and their weight values:

[0198] Table 9 Key field weight distribution table

[0199] Field Name <![CDATA[Weight w i > A 0.6 B 0.8 C 0.7 D 0.9

[0200] As shown in Table 9, different fields have different weights, indicating that they have different influences in the prediction value calculation. After setting the critical weight, the mapping target value M of the field is calculated. i , which can be obtained by using group mean or trend prediction. For example, assuming that the current field feature value T r is: T = [130, 95, 110, 80];

[0201] Mapping target value M r M = [100, 100, 100, 100];

[0202] Logical grouping deviation factor P r is: P = [10, 8, 12, 9];

[0203] Eigenvalue correlation factor Q r Q = [5, 6, 7, 4];

[0204] Enter the formula to calculate:

[0205]

[0206] Calculate the square root:

[0207]

[0208] Calculate the numerator:

[0209] 0.6×30=18, 0.8×5=4, 0.7×10=7, 0.9×20=18;

[0210] Final calculation:

[0211]

[0212] The final calculated F c The value is 12.32, which represents the updated field mapping value and serves as the feature field classification prediction value.

[0213] See also Figure 6 , the specific steps for obtaining optimized structured output data are:

[0214] Filter conflicting field values ​​from the feature field classification prediction values, perform priority parameter determination for each conflicting field value, determine whether to retain or replace the field value by comparing the priority parameter size of the field value, record the processed field value, and generate a priority allocation field result;

[0215] To filter out conflicting field values ​​from the classification prediction values ​​of feature fields, we first need to clarify the definition of conflicting fields, that is, field values ​​from different data sources under the same category are inconsistent or have large deviations. We need to identify conflicting fields by comparing field values. Field value comparison can be calculated based on absolute difference, normalized error, or statistical distribution deviation. For example, for fields A and B from the same data source, calculate their difference parameters.

[0216] |AB|, if the difference exceeds the set threshold, the field is determined to be a conflict field. When setting the threshold, it can be calculated based on historical data or statistical methods, such as using standard deviation to set the threshold in is the mean, σ is the standard deviation, k is the adjustment coefficient. If the difference in field values ​​exceeds the threshold θ, the field is determined to be a conflicting field. Priority parameter determination is performed for each conflicting field value. The priority parameter can be calculated based on the reliability, timestamp or historical weight of the data source. For example, assuming that data sources S1 and S2 provide field values ​​100 and 120 respectively, and the weight of S1 is 0.8 and the weight of S2 is 0.6, then the weighted field value V = (100×0.8+120×0.6) / (0.8+0.6) = 107.14 is calculated. If the deviation between the field value of the data source with a higher weight and the calculated value is within the set range, the field value is retained. Otherwise, it is replaced with the calculated weighted value. The processed field value is recorded to finally generate the priority allocation field result.

[0217] Based on the priority allocation field result, the processed field values ​​are rearranged, and the field arrangement order is adjusted and the updated field mapping relationship is constructed in combination with the logical association between field features and the group mapping rules to generate the field mapping relationship result;

[0218] Based on the priority assignment field results, the processed field values ​​are rearranged. The field arrangement is adjusted based on the logical relationship and data dependency between the fields. First, the logical association between the fields is determined. The logical association can be calculated by the correlation, dependency or causal relationship between the fields. For example, the correlation between field A and field B can be calculated by the Pearson correlation coefficient:

[0219]

[0220] If |r|>0.7, it is considered that fields A and B are strongly associated, and the fields need to be arranged in close relative positions. If |r|<0.3, the association between the fields is weak and they can be arranged separately. Based on the correlation calculation results, the fields are logically grouped and mapped. The grouping mapping rules can be determined based on historical data or domain knowledge. For example, revenue, profit, and cost in financial data can be grouped into one group, while click-through rate, browsing time, and bounce rate in user behavior data can be grouped into another group. After all fields are grouped, the field arrangement order is adjusted to keep logically related fields in adjacent positions, and an updated field mapping relationship is constructed to finally generate the field mapping relationship result. Table 10 shows the example fields and their logical association calculation values.

[0221] Table 10 Field logical association calculation table

[0222] Field Name Related fields Correlation coefficient r Logical Grouping A B 0.82 Group 1 B C 0.75 Group 1 D E 0.91 Group 2 F G 0.28 Group 3

[0223] As shown in Table 10, fields A, B, and C belong to the same logical grouping, fields D and E belong to another group, and fields F and G have no obvious association and can be arranged separately.

[0224] Call the field mapping relationship results, perform dynamic updates on the field interaction characteristics, determine each item through the field interaction parameters, and build an optimization value based on the collaborative characteristics between the fields. The formula is:

[0225]

[0226] Compute optimized structured output data O s ;

[0227] Among them, w r,s Represents the field interaction weight parameter, I r,s Represents the actual value of the interaction between fields, A r,s represents the interaction target value between fields, R r,s represents the field coordination adjustment ratio, C r,s Represents the field associated logical parameter, E r,s represents the field dynamic deviation factor, P r,s represents the interaction parameter between fields, Q r,srepresents the field interaction feature correction parameter, r and s represent the field number, N f Represents the total number of fields.

[0228] Call the field mapping relationship result to dynamically update the field interaction characteristics. The field interaction characteristics are determined item by item based on the interaction parameters between fields. The interaction parameters can be calculated through the coordinated changes of field values. For example, if the change trends of field A and field B are consistent, their synergy characteristics are strong, and the coordinated adjustment ratio R ij It can be calculated by the mean square error:

[0229]

[0230] If R r,s <0.05, the field collaborative adjustment ratio is high, and joint optimization can be performed. The optimization calculation uses the formula:

[0231]

[0232] Among them, w r,s Represents the field interaction weight parameter, I r,s Represents the actual value of the interaction between fields, A r,s represents the interaction target value between fields, R r,s represents the field coordination adjustment ratio, C r,s Represents the field associated logical parameter, E r,s represents the field dynamic deviation factor, P r,s represents the interaction parameter between fields, Q r,s represents the field interaction feature correction parameter, r and s represent the field number, N f Represents the total number of fields.

[0233] Table 11 Field interaction parameter table

[0234]

[0235] As shown in Table 11, different field pairs have different interaction weights and adjustment ratios, which are calculated by the formula:

[0236]

[0237] Calculate the square root:

[0238]

[0239] Calculate the numerator part:

[0240]

[0241] Final calculation:

[0242]

[0243] The final calculated O s The value is 1.54, which represents optimized structured output data.

[0244] The above are only preferred embodiments of the present invention and are not intended to limit the present invention in other forms. Any technician familiar with the profession may use the technical contents disclosed above to change or modify them into equivalent embodiments with equivalent changes and apply them to other fields. However, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution of the present invention still falls within the protection scope of the technical solution of the present invention.

Claims

1. A multimodal data recognition and analysis system based on deep learning, characterized in that: The system comprises: The multimodal data preprocessing module extracts feature points and performs region segmentation operations on the pixel matrix in the product image based on the input product image and the text information it contains, performs serialization encoding processing on the continuous character area in the text information, extracts features from the image and text respectively, and aligns the joint embedding space so that the two modal data can be associated in the same representation space to establish preliminary multimodal feature data; The multimodal alignment feature generation module matches the extracted local area features of the image with the text feature vector based on the preliminary multimodal feature data, calculates the feature similarity, cross-combines the features for the mismatched parts through the joint embedding representation and performs nonlinear mapping operations, and generates an aligned multimodal feature distribution value after screening and discriminating the generated data; The knowledge logic correction module performs rule verification and adjusts the range of the logic nodes in the field distribution value of the specification parameter based on the aligned multimodal feature distribution value, ensures that the numerical field in the specification parameter meets the preset logical constraints, replaces the node value that fails the verification according to the association rule, performs field update and syntax verification in combination with the context conditions, and generates a semantic logic adjustment value; The classification field structure prediction module extracts the field logic relationship and groups it based on the semantic logic adjustment value, performs item-by-item comparison on the key feature values ​​screened in the commodity classification field, constructs a new field mapping relationship for the feature values ​​not covered by the classification standard, and generates a feature field classification prediction value; The dynamic result optimization generation module performs priority assignment operations on conflicting field values ​​in the recommended tags based on the classification prediction values ​​of the feature fields, builds mapping relationships after rearranging the fields, and dynamically updates and merges the field interaction characteristics to generate optimized structured output data.

2. The multimodal data recognition and analysis system based on deep learning according to claim 1, characterized in that: The multimodal feature data specifically includes image feature point distribution values, image area segmentation values, and text serialization encoding values; the multimodal alignment feature distribution values ​​specifically include image character area alignment values, text sequence matching values, and feature cross-combination values; the semantic logic adjustment values ​​specifically include field logic check values, field update values, and syntax check values; the feature field classification prediction values ​​specifically include grouping field logic relationship values, new field mapping relationship values, and classification key feature values; the optimized structured output data includes field priority assignment values, field interaction characteristic update values, and field rearrangement mapping relationship values.

3. The multimodal data recognition and analysis system based on deep learning according to claim 2 is characterized in that: The steps for obtaining the preliminary multimodal feature data are specifically as follows: Extract a pixel matrix from the input image data, calculate the absolute value of the gradient change of each pixel point by performing point-by-point operation and comparison on the gradient change value of each pixel point in the pixel matrix, compare the value with the set threshold, mark the pixel points whose absolute value of the change is greater than the threshold, summarize the marked data after marking, and generate a feature point distribution matrix; Based on the feature point distribution matrix, the ratio of the local mean of the grayscale value of the pixel matrix to the variance of the pixel value is calculated, the area where the ratio value changes significantly is screened, the boundary of each area that satisfies the ratio change is divided, and the identification classification after the area division is generated. After the classification, the boundary features and grayscale value parameters of each area are counted to generate a set of image area feature parameters; Perform serialization encoding processing on the continuous character area of ​​the text data, calculate the recurrence frequency of the character segment, sort each segment according to the frequency and length pattern, group the sorted character segments, perform operations on each group of character segments, record the sequence pattern and feature parameters of each group, and generate a text sequence feature parameter set; The image region feature parameter set and the text sequence feature parameter set are called to calculate the numerical relationship of the cross-comparison of the two sets of fields, and the formula is adopted by adjusting the field key weight: Calculate preliminary multimodal feature data D m ; Among them, w p Represents the critical weight of the image field, P g,p Represents the image region feature parameter value, P t,q represents the characteristic parameter value of the text sequence, p represents the pth characteristic field, N g Represents the total number of feature fields in the image region.

4. The multimodal data recognition and analysis system based on deep learning according to claim 3 is characterized in that: The steps of obtaining the distribution value of the aligned multimodal features are specifically as follows: Extracting image character regions and text sequence values ​​from the preliminary multimodal feature data, matching image fields and text fields pair by pair according to the grouping rule, calculating the fit value between each pair of fields, screening and recording all field pairs whose fit values ​​are lower than a set threshold, and generating a set of unmatched field pairs; Based on the set of unmatched field pairs, cross-combination is performed for the corresponding characteristic parameters in each pair of fields, the weight coefficient of each parameter after combination is adjusted according to the field criticality weight, the weighted fitting value of the combined parameter is calculated, the difference between the fitting value and the original field parameter is corrected, and a set of field characteristic cross-combination fitting values ​​is generated; The field feature cross-combination fitting value set is called, and the field screening results are judged in combination with the field criticality weight and matching degree, and the field feature difference and dynamic adjustment threshold are used for screening, using the formula: Calculate the aligned multimodal feature distribution value F d ; Among them, w p Represents the criticality weight of the field, C g,p Represents the characteristic value of the image field, C t,q Represents the characteristic value of the text field, p represents the pth field, N f Represents the total number of feature fields.

5. The multimodal data recognition and analysis system based on deep learning according to claim 4 is characterized in that: The steps for obtaining the semantic logic adjustment value are specifically as follows: Extracting logical nodes in the field distribution value from the aligned multimodal feature distribution value, for each field value of the logical node, gradually performing a determination of a numerical range according to a rule check, marking and recording logical nodes that fail to pass the check during the determination process, and generating a set of logical node values ​​that fail to pass the check; Based on the set of logical node values ​​that have not passed the verification, calling the association rule corresponding to each of the logical nodes, calculating the difference parameters between the logical node value and the target value of the association rule item by item, performing operations on the difference parameters to obtain replacement values, and updating the original logical node value with the replacement value to generate a set of field logical node correction values; The field logic node correction value set is called, the field value of the dependent logic node is dynamically updated in combination with the context conditions, and the context logic conditions that the field depends on are checked one by one, using the formula: Calculate the semantic logic adjustment value L s ; Among them, w s Represents the critical weight of the logical node, R s represents the target value of the association rule, T s Represents the current value of the logical node, D s Represents the context dependency parameter of field update, S s Represents the syntax dependency parameter of the field value, s represents the sth logical node, N l Represents the total number of logical nodes.

6. The multimodal data recognition and analysis system based on deep learning according to claim 5, characterized in that: The steps for obtaining the classification prediction value of the feature field are specifically as follows: Extracting the field logical relationship from the semantic logic adjustment value, gradually determining the logical relevance between the fields according to the associated characteristic parameters between the fields, classifying and grouping the fields whose logical relevance meets the same characteristic parameters, and generating a field logical grouping result; Based on the field logic grouping result, for the key feature values ​​selected in the grouping field, the deviation parameter of each feature value and the feature mean in the group is calculated item by item, the field feature values ​​whose deviation parameters exceed the set range are recorded, and the results are summarized to form an uncovered feature value set; The uncovered feature value set is called, the field mapping relationship is rebuilt according to the logical grouping and deviation parameters, and the field feature value is dynamically updated on the mapping relationship, using the formula: Calculate the classification prediction value F of the feature field c ; Among them, w r Represents the key weight of the field, T r Represents the current value of the feature field, M r Represents the target value of the feature field mapping, P r Represents the logical grouping deviation factor, Q r represents the eigenvalue correlation factor, r represents the rth field, N f Represents the total number of fields.

7. The multimodal data recognition and analysis system based on deep learning according to claim 6, characterized in that: The steps of obtaining the optimized structured output data are specifically as follows: Filtering conflicting field values ​​from the characteristic field classification prediction values, performing priority parameter determination for each conflicting field value, determining whether to retain or replace the field value by comparing the priority parameter size of the field value, recording the processed field value, and generating a priority allocation field result; Based on the priority allocation field result, the processed field values ​​are rearranged, and the field arrangement order is adjusted and an updated field mapping relationship is constructed in combination with the logical association between field features and the grouping mapping rules, so as to generate a field mapping relationship result; The field mapping relationship result is called, the field interaction characteristics are dynamically updated, and the interaction parameters between the fields are determined item by item. The optimization value is constructed based on the collaborative characteristics between the fields, using the formula: Compute optimized structured output data O s ; Among them, w r,s Represents the field interaction weight parameter, I r,s Represents the actual value of the interaction between fields, A r,s represents the interaction target value between fields, R r,s represents the field coordination adjustment ratio, C r,s Represents the field associated logical parameter, E r,s represents the field dynamic deviation factor, P r,s represents the interaction parameter between fields, Q r,s represents the field interaction feature correction parameter, r and s represent the field number, N f Represents the total number of fields.

Citation Information

Cited By

  • Freezing circle multi-source heterogeneous data semantic fusion method and system

    CN120744836A