Multimodal data feature processing method, system and medium based on deep learning
Through the multimodal data feature processing method based on deep learning, the problem that traditional technology cannot recognize multiple types of data images is solved, effective processing and feature extraction of multiple types of data images is realized, and more complete data representation is provided.
Patent Information
- Application Number
- CN202510026855.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-08
AI Technical Summary
Traditional multimodal data processing technology cannot identify multiple types of data images, resulting in the inability to provide multiple types of data images when users look up data, resulting in information loss or misidentification, insufficient feature expression, and degradation of processing results.
The multimodal data feature processing method based on deep learning is adopted. By acquiring different types of data images, the corresponding image pre-processing algorithm is determined, the data images are converted into structured data, label data is added, feature data is extracted, alignment factors are selected for alignment, dimensionality reduction processing, and the deep learning model is trained to obtain the data images corresponding to the tag data.
It realizes effective processing and feature extraction of multiple types of data images, ensuring that different modal data can correspond to the same entity or event, providing more complete and consistent data representation, and solving the problem that users cannot provide multiple types of data images when looking up data.
Smart Images

Figure CN119418142B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image feature processing, and in particular to a multimodal data feature processing method, system and medium based on deep learning. Background Art
[0002] When users search for resources on the system, since most of the materials collected by the system are unstructured data, the general query function can only match the file name according to keywords, and cannot search according to the file content, especially image format. As a result, users always cannot find all or the wrong information when searching, resulting in a large amount of data in the system not being fully utilized.
[0003] When using traditional multimodal data processing technology to collect data in image format files, traditional multimodal data processing technology is limited to processing a certain type of image, such as processing map images or animal images. However, in reality, although the collected data are all images, there are often multiple types of images, some are photos of a table content, some are photos of a document letter, some are scanned copies, and some are maps. Although they are all in image format, they are heterogeneous data. When a multimodal data processing technology is used to process such heterogeneous data, it is impossible to fully capture and represent the information in the image, which may lead to information loss or misidentification, resulting in insufficient feature expression and a decrease in the quality of the final processing results.
[0004] Therefore, there is an urgent need for a multimodal data feature processing method, system and medium based on deep learning to solve the problem that traditional multimodal data processing technology cannot recognize multiple types of data images and thus cannot provide multiple types of data images when users search for data. Summary of the invention
[0005] In view of the above-mentioned deficiencies in the prior art, the present application provides a multimodal data feature processing method, system and medium based on deep learning to solve the problem that traditional multimodal data processing technology cannot recognize multiple types of data images and thus cannot provide multiple types of data images when users search for data.
[0006] In a first aspect, the present application provides a method for processing multimodal data features based on deep learning, the method comprising:
[0007] Acquire data images of various types, determine the corresponding image preprocessing algorithm based on the specific type of the data image, and convert the data image into structured data; wherein the types of data images include at least text data images, table data images, and map data images; add label data to the structured data; determine the corresponding feature extraction algorithm according to the image type corresponding to the structured data, and obtain the feature data set corresponding to each image type; wherein the feature data set includes label data; obtain an alignment factor, and then obtain several feature data sets belonging to the same alignment factor according to the alignment factor; wherein the alignment factor includes at least one or more of the following: time, unique identifier, preset content keyword, event name, object name, geographic coordinates, etc. The invention relates to a method for obtaining a dimensionally reduced data matrix and a plurality of feature data sets belonging to the same alignment factor. The method comprises the following steps: first, a principal component space is obtained by calculating the covariance matrix, and the normalized data matrix is projected into the principal component space to obtain a dimensionally reduced data matrix; the preset deep learning model is trained by the dimensionally reduced data matrix and the corresponding label data to obtain a trained preset deep learning model; when a user retrieves information, the corresponding alignment factor is determined according to the user retrieval information, and then the corresponding dimensionally reduced data matrix is determined; the dimensionally reduced data matrix is used as the input of the trained preset deep learning model to obtain the label data; and then the data images of various types corresponding to the label data are obtained.
[0008] The multimodal data feature processing method provided in the embodiment of the present application provides an image preprocessing algorithm for processing different types of data images, which can convert different types of data images into structured data, and can process data images such as text data images, table data images, and map data images; in addition, corresponding feature extraction algorithms are set for different types of data images. After extracting the features of each modality, different alignment factors can be selected according to actual needs for alignment, and they are aligned together to ensure that data of different modalities can correspond to the same entity or event, so as to ensure that the features after subsequent fusion are meaningful. The purpose of alignment is to ensure that information of different modalities can complement and support each other, thereby providing a more complete and consistent data representation. Through the trained preset deep learning model, the reduced dimensionality data matrix corresponding to the actual needs of subsequent users is used as input to obtain the label data corresponding to the actual needs of the user, and then obtain the data images of each type corresponding to the label data, which solves the problem that multiple types of data images cannot be provided when users search for data.
[0009] In one implementation of the present application, various types of data images are obtained, and corresponding image preprocessing algorithms are determined based on the specific types of the data images to convert the data images into structured data, specifically including:
[0010] When the data image is a text-type data image, the text-type data image is converted into structured data in an editable text format through OCR technology; when the data image is a table-type data image, the table-type data image is grayed and binarized through image preprocessing, and the color table-type data image is converted into a black and white table-type data image; the median filtering technology is used to remove noise in the table-type data image and correct the tilt of the table-type data image; the edge detection algorithm is used to find the lines in the table-type data image, the Hough line detection algorithm is used to detect the straight lines in the lines, and the contour detection function is used to find all closed contours in the straight lines; each closed contour is cut out, and the OCR technology is used to identify the text in each closed contour, and then the structured data corresponding to the table-type data image is obtained; when the data image is a map-type data image, the geographic information system technology is used in combination with the OCR technology to extract the geographic information in the map-type data image as the structured data of the current data image; wherein the geographic information at least includes longitude, latitude, street name, and street number.
[0011] In one implementation of the present application, adding label data to structured data specifically includes: obtaining label data corresponding to each structured data through a preset data upload terminal; and identifying preset annotation keywords corresponding to the structured data through a preset semantic recognition algorithm, and determining that the label data corresponding to the preset annotation keywords is the label data corresponding to the current structured data.
[0012] In one implementation of the present application, according to the image type corresponding to the structured data, a corresponding feature extraction algorithm is determined to obtain a feature data set corresponding to each image type, specifically including:
[0013] When the structured data corresponds to a text-type data image, a text feature extraction algorithm is called to extract vector feature data from the structured data, and a feature data set corresponding to the current text-type data image is generated; when the structured data corresponds to a table-type data image, the data type of each column is identified; wherein the data type includes at least: numerical type and categorical type; the numerical structured data is processed into feature data using normalization processing; the categorical structured data is processed into binary feature data using one-hot encoding to obtain a feature data set corresponding to the current table-type data image; when the structured data corresponds to a map-type data image, geographic coordinate features and regional attribute features are extracted from the structured data; a preset clustering label corresponding to the current geographic coordinate is determined by a preset geographic coordinate feature set containing the current geographic coordinate feature and a K-Means clustering algorithm; a preset number of geographic coordinate features closest to the current geographic coordinate feature are obtained as spatial relationship features by a preset geographic coordinate feature set containing the current geographic coordinate feature and a haversine method; the geographic coordinate features, regional attribute features, preset clustering labels and spatial relationship features are added to the feature data set corresponding to the current map-type data image.
[0014] In one implementation of the present application, after obtaining the alignment factor, and then obtaining several feature data sets belonging to the same alignment factor according to the alignment factor, the method includes:
[0015] Confirm whether the data formats in several feature data sets belonging to the same alignment factor are consistent; when there is inconsistency in the data formats, modify the inconsistent data formats to a preset unified format.
[0016] In one implementation of the present application, several feature data sets belonging to the same alignment factor are spliced into a feature vector to obtain a normalized data matrix and a covariance matrix of the feature vector; the principal component space is calculated through the covariance matrix, and the normalized data matrix is projected into the principal component space to obtain a data matrix after dimensionality reduction, specifically including:
[0017] The keywords in several feature data sets with the same alignment factor are concatenated together to generate a total feature set, and the total feature set is converted into a feature vector; wherein the feature vector is a data matrix of m features of n feature data sets; the feature vector is normalized using MinMaxScaler to obtain a normalized data matrix;
[0018] By formula:
[0019] , calculate the covariance matrix Y; where, represents the normalized data matrix, n represents the total number of feature data sets belonging to the same alignment factor;
[0020] The numpy.linalg.eig function in the python function library is used to calculate the one-dimensional eigenvalue array of the covariance matrix; the first K values in the one-dimensional eigenvalue array are taken as the principal component space;
[0021] By formula:
[0022] , project the normalized data matrix into the principal component space to obtain the reduced-dimensional data matrix Z; where, represents the principal component space.
[0023] In one implementation of the present application, when performing user retrieval information, the corresponding alignment factor is determined according to the user retrieval information, and then the corresponding reduced-dimensional data matrix is determined, specifically including:
[0024] The user alignment factor is extracted from the user retrieval information, and several feature data sets belonging to the user alignment factor are obtained and spliced into a feature vector to obtain the normalized data matrix and covariance matrix of the feature vector; the principal component space is calculated through the covariance matrix, and the normalized data matrix is projected into the principal component space to obtain the data matrix after dimensionality reduction.
[0025] In one implementation of the present application, after acquiring each type of data image corresponding to the label data, the method further includes:
[0026] Each type of data image is vectorized into a total matrix; the Pearson correlation coefficient between each two data image vectorizations in the total matrix is calculated through the np.corrcoef function; and an association relationship graph of the data images is generated, wherein the association relationship graph is composed of nodes and edges, and the nodes represent the data images and the edges represent the correlation coefficient between the two data images.
[0027] The multimodal data feature processing method provided in the embodiment of the present application generates a correlation diagram of the data image through multimodal data correlation analysis, can identify the correlation between different modalities, can solve the problem that a single modality data cannot present the correlation, and helps to build a more comprehensive data representation.
[0028] In a second aspect, the present application provides a multimodal data feature processing system based on deep learning, the system comprising:
[0029] A conversion module is used to obtain data images of various types, determine the corresponding image preprocessing algorithm based on the specific type of the data image, and convert the data image into structured data; wherein the types of data images include at least text data images, table data images and map data images; add label data to the structured data; a feature acquisition module is used to determine the corresponding feature extraction algorithm according to the image type corresponding to the structured data, and obtain the feature data set corresponding to each image type; wherein the feature data set includes label data; a matrix acquisition module is used to obtain an alignment factor, and then obtain a number of feature data sets belonging to the same alignment factor according to the alignment factor; wherein the alignment factor includes at least any one or more of the following: time, unique identifier, preset content keyword, event name, The invention relates to a method for obtaining a plurality of feature data sets belonging to the same alignment factor, and a normalized data matrix and a covariance matrix of the feature vector are obtained; the principal component space is obtained by calculating the covariance matrix, and the normalized data matrix is projected into the principal component space to obtain a data matrix after dimensionality reduction; a data fusion display module is used to train a preset deep learning model through the data matrix after dimensionality reduction and the corresponding label data to obtain a trained preset deep learning model; when a user retrieves information, the corresponding alignment factor is determined according to the user retrieval information, and then the corresponding data matrix after dimensionality reduction is determined; the data matrix after dimensionality reduction is used as the input of the trained preset deep learning model to obtain label data; and then various types of data images corresponding to the label data are obtained.
[0030] In a third aspect, the present application provides a non-volatile computer storage medium having computer instructions stored thereon, which when executed implement a multimodal data feature processing method based on deep learning as any of the above items.
[0031] Those skilled in the art can understand that the present application has at least the following beneficial effects:
[0032] The present application provides an image preprocessing algorithm for processing different types of data images, which can convert different types of data images into structured data, and can process data images such as text data images, table data images, and map data images; in addition, corresponding feature extraction algorithms are set for different types of data images. After extracting the features of each modality, different alignment factors can be selected according to actual needs for alignment, and they are aligned together to ensure that data of different modalities can correspond to the same entity or event, so as to ensure that the features after subsequent fusion are meaningful. The purpose of alignment is to ensure that information of different modalities can complement and support each other, thereby providing a more complete and consistent data representation. Through the trained preset deep learning model, the reduced dimensionality data matrix corresponding to the actual needs of subsequent users is used as input to obtain the label data corresponding to the actual needs of the user, and then obtain the various types of data images corresponding to the label data, which solves the problem of multiple types of data images that cannot be provided when users search for data.
[0033] In addition, the present application generates a correlation diagram of data images through multimodal data correlation analysis, which can identify the correlation between different modalities, solve the problem that single modality data cannot present correlation, and help build a more comprehensive data representation. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Some embodiments of the present disclosure are described below with reference to the accompanying drawings, in which:
[0035] Figure 1 This is a flow chart of a multimodal data feature processing method based on deep learning provided in an embodiment of the present application.
[0036] Figure 2 It is a schematic diagram of the internal structure of a multimodal data feature processing system based on deep learning provided in an embodiment of the present application. DETAILED DESCRIPTION
[0037] It should be understood by those skilled in the art that the embodiments described below are only preferred embodiments of the present disclosure, and do not mean that the present disclosure can only be implemented through the preferred embodiments. The preferred embodiments are only used to explain the technical principles of the present disclosure, and are not used to limit the protection scope of the present disclosure. Based on the preferred embodiments provided by the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work should still fall within the protection scope of the present disclosure.
[0038] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0039] The technical solution proposed in the embodiments of the present application is described in detail below with reference to the accompanying drawings.
[0040] The embodiment provides a multimodal data feature processing method based on deep learning, such as Figure 1 As shown, the method provided in the embodiment of the present application mainly includes the following steps:
[0041] Step 110: Acquire data images of various types, determine corresponding image preprocessing algorithms based on specific types of the data images, convert the data images into structured data; and add label data to the structured data.
[0042] It should be noted that the types of data images include at least text data images, table data images and map data images.
[0043] In the step, various types of data images are obtained, and the corresponding image preprocessing algorithm is determined based on the specific type of the data image to convert the data image into structured data, which can be specifically:
[0044] When the data image is a text data image, the text data image is converted into structured data in an editable text format through OCR (Optical Character Recognition) technology.
[0045] When the data image is a table-type data image, the table-type data image is grayed and binarized through image preprocessing to convert the color table-type data image into a black and white table-type data image; the median filtering technology is used to remove the noise in the table-type data image and correct the tilt of the table-type data image; the edge detection algorithm is used to find the lines in the table-type data image, the Hough line detection algorithm is used to detect the straight lines in the lines, and the contour detection function is used to find all closed contours in the straight lines; each closed contour is cropped, and the OCR technology is used to recognize the text in each closed contour, so as to obtain the structured data corresponding to the table-type data image.
[0046] It should be noted that the solution for correcting the tilt of the table-type data image here is an existing solution and this application does not limit it.
[0047] When the data image is a map-type data image, the geographic information system technology is used in combination with the OCR technology to extract the geographic information in the map-type data image as the structured data of the current data image; wherein the geographic information at least includes longitude, latitude, street name, street number, and other detailed addresses.
[0048] The tag data is added to the structured data, specifically:
[0049] The label data corresponding to each structured data is obtained through the preset data upload terminal; and the preset annotation keywords corresponding to the structured data are identified through the preset semantic recognition algorithm, and the label data corresponding to the preset annotation keywords are determined to be the label data corresponding to the current structured data.
[0050] It should be noted that the preset data upload terminal can be a device for the labeling personnel (corresponding to the preset data upload terminal) to read the text and then label the text with event category labels such as "economy", "people's livelihood", "sports", etc. according to the content of the text.
[0051] In addition, label data can be added to structured data according to pre-designed annotation rules, such as dividing age into children, youth, adults, and the elderly. For geographic information data, longitude and latitude can be divided as regional markers according to the preset regional division range.
[0052] Step 120: Determine a corresponding feature extraction algorithm according to the image type corresponding to the structured data, and obtain a feature data set corresponding to each image type.
[0053] It should be noted that the feature data set includes label data.
[0054] This step can be specifically as follows:
[0055] When the structured data corresponds to a text-type data image, a text feature extraction algorithm is called to extract vector feature data in the structured data and generate a feature data set corresponding to the current text-type data image.
[0056] It should be noted that before calling the text feature extraction algorithm, the present application may perform text cleaning on the text. The text feature extraction algorithm may specifically be:
[0057] 1. Use a tokenization tool to break the text into lexical units. 2. Stop word removal: Remove common stop words (such as "de", "shi", "zai", etc.) to reduce noise. 3. Part-of-speech tagging: Tag the part of speech for each word, such as "wo (pronoun)". 4. Word sense disambiguation: Determine the specific meaning of a polysemous word in a particular context based on the context. 5. Word frequency statistics: Calculate the frequency of each word occurrence to provide a basis for subsequent feature selection. 6. Word frequency calculation: Use TF-IDF to calculate the term frequency-inverse document frequency to highlight important words in the document. 7. Categorical features (categorical features in structured data): Use one-hot encoding to convert categorical features into binary vectors.
[0058] When the structured data corresponds to a table-like data image, identify the data types of each column; among them, the data types include at least: numerical type, categorical type; use normalization processing to process the numerical structured data into feature data; use one-hot encoding to process the categorical structured data into binary feature data to obtain the feature data set corresponding to the current table-like data image.
[0059] It should be noted that before identifying the data types of each column, data cleaning can be performed on the identified column data to clean null values, outliers, etc. The specific method of normalization processing can be to use MinMaxScaler to normalize numerical features. In addition, when missing values occur during the processing of this application, for missing data, if it is date or sorting data, it is filled in by forward filling, and if it is data representing a status, it is filled in according to the actual situation.
[0060] When the structured data corresponds to a map-like data image, extract geographical coordinate features and regional attribute features from the structured data; through a preset geographical coordinate feature set containing the current geographical coordinate features and the K-Means clustering algorithm, determine the preset clustering label corresponding to the current geographical coordinate; through the preset geographical coordinate feature set containing the current geographical coordinate features and the haversine method, obtain the preset number of geographical coordinate features closest to the current geographical coordinate feature as spatial relationship features; add the geographical coordinate features, regional attribute features, preset clustering label, and spatial relationship features to the feature data set corresponding to the current map-like data image.
[0061] It should be noted that determining the preset clustering label corresponding to the current geographical coordinate through the preset geographical coordinate feature set containing the current geographical coordinate features and the K-Means clustering algorithm can be specifically:
[0062] Suppose there is a series of geographic coordinate data (including the current geographic coordinate features), including longitude and latitude values, and now we want to divide it into 10 clusters, then K=10, use the KMeans algorithm in python to generate cluster labels, KMeans will use the 10 initial nodes as the center points, divide the data into 10 categories, and get cluster labels (preset cluster labels).
[0063] Step 130: obtain an alignment factor, and then obtain several feature data sets belonging to the same alignment factor according to the alignment factor; concatenate several feature data sets belonging to the same alignment factor into a feature vector to obtain a normalized data matrix and a covariance matrix of the feature vector; calculate the principal component space through the covariance matrix, project the normalized data matrix into the principal component space, and obtain the data matrix after dimensionality reduction.
[0064] It should be noted that the alignment factor includes at least one or more of the following: time, unique identifier (each piece of data has a unique identifier that runs through the entire data processing process), preset content keywords, event name, object name, geographic coordinates, and area range.
[0065] After obtaining the alignment factor and then obtaining several feature data sets belonging to the same alignment factor according to the alignment factor, the method further includes:
[0066] Confirm whether the data formats in several feature data sets belonging to the same alignment factor are consistent; when there is inconsistency in the data formats, modify the inconsistent data formats to a preset unified format.
[0067] For example, date format, multi-source data should maintain a unified format.
[0068] Among them, according to the alignment factor, several feature data sets belonging to the same alignment factor are obtained, which can be exemplified as follows:
[0069] The text data records the details of the event, the time of occurrence, the location, etc. The table data contains the date, event name, responsible person, etc. The geographic information data records the specific address information. Date alignment: Align the date in the text data with the date in the table data to summarize the events that occurred on a certain day; coordinate alignment: summarize all the events that occurred at a certain location. Assign a unique representation to the event,
[0070] Those skilled in the art will appreciate that, by aligning the features of the alignment factors, it is possible to ensure that there is a correlation between multi-source data, and that data of different modalities can be associated.
[0071] In the step, several feature data sets belonging to the same alignment factor are concatenated into a feature vector to obtain the normalized data matrix and covariance matrix of the feature vector; the principal component space is calculated through the covariance matrix, and the normalized data matrix is projected into the principal component space to obtain the data matrix after dimensionality reduction, which can be specifically:
[0072] The keywords in several feature data sets belonging to the same alignment factor are concatenated together to generate a total feature set, and the total feature set is converted into a feature vector; wherein the feature vector is a data matrix of m features of n feature data sets.
[0073] Use MinMaxScaler to normalize the feature vector to obtain the normalized data matrix.
[0074] By formula:
[0075] , calculate the covariance matrix Y; where, represents the normalized data matrix, n represents the total number of feature data sets belonging to the same alignment factor; the numpy.linalg.eig function in the python function library is used to calculate the one-dimensional eigenvalue array of the covariance matrix; the first K values in the one-dimensional eigenvalue array are taken as the principal component space.
[0076] By formula:
[0077] , project the normalized data matrix into the principal component space to obtain the reduced-dimensional data matrix Z; where, represents the principal component space.
[0078] Step 140: train a preset deep learning model using the reduced-dimensionality data matrix and the corresponding label data to obtain a trained preset deep learning model; when performing user retrieval information, determine the corresponding alignment factor based on the user retrieval information, and then determine the corresponding reduced-dimensionality data matrix; use the reduced-dimensionality data matrix as the input of the trained preset deep learning model to obtain label data; and then obtain data images of various types corresponding to the label data.
[0079] In the step, when the user retrieves information, the corresponding alignment factor is determined according to the user retrieval information, and then the corresponding data matrix after dimensionality reduction is determined, which can be specifically:
[0080] The user alignment factor is extracted from the user retrieval information, and several feature data sets belonging to the user alignment factor are obtained and spliced into a feature vector to obtain the normalized data matrix and covariance matrix of the feature vector; the principal component space is calculated through the covariance matrix, and the normalized data matrix is projected into the principal component space to obtain the data matrix after dimensionality reduction.
[0081] As an example, the above process can be specifically as follows: the user reports that parking charges are unreasonable on a certain street (user retrieves information), the system first pre-processes the user input (extracts the user alignment factor from the user retrieval information), and then performs feature fusion (several feature data sets belonging to the user alignment factor are spliced into a feature vector to obtain the normalized data matrix and covariance matrix of the feature vector; the principal component space is calculated through the covariance matrix, and the normalized data matrix is projected into the principal component space to obtain the reduced-dimensional data matrix), the fused features (the reduced-dimensional data matrix) are input into the trained preset deep learning model to obtain the prediction results, and finally, according to the label data in the prediction results, the corresponding original data (data image) is searched and displayed to the user.
[0082] In addition, the present application can also identify the association relationship between different modalities, solve the problem that single modality data cannot present the association relationship, and help to build a more comprehensive data representation. The specific process can be:
[0083] After obtaining each type of data image corresponding to the label data, each type of data image is vectorized into a total matrix; the Pearson correlation coefficient between each two data image vectorizations in the total matrix is calculated through the np.corrcoef function; and an association relationship graph of the data images is generated, wherein the association relationship graph is composed of nodes and edges, and the nodes represent the data images and the edges represent the correlation coefficient between the two data images.
[0084] The above process can be specifically described as follows:
[0085] The prediction results are vectorized to obtain the data matrix X (total matrix); the Pearson correlation coefficient of each two vectorized representations in the overall vector is calculated through the following np.corrcoef function. The value range is -1 to 1, and the larger the absolute value, the stronger the correlation.
[0086] correlation_matrix = np.corrcoef(XT),
[0087] Here, np.corrcoef is a function in Python's NumPy library, and XT is the transpose of the data matrix X.
[0088] Based on the above description, it can be understood by those skilled in the art. The present invention provides a method, system and medium for processing multimodal data features based on deep learning, a method for processing the same type of data format but unstructured data with different contents, such as images in jpg format, but the image contents include text data images, table data images, map data images, etc. The effect is that data of different modalities provide complementary information. For example, data images can provide visual information, while text data can provide contextual or descriptive information. This complementarity helps to improve the accuracy of classification tasks. In addition to jpg format, the present application can also process data in png, excel, PDF and other formats. In addition, the present application can generate an association relationship diagram of data images, and then can identify the association relationship between different modalities, solve the problem that single modality data cannot present association relationships, and help to build a more comprehensive data representation.
[0089] In addition, this application Figure 2 A multimodal data feature processing system based on deep learning is provided in the embodiment of the present application. Figure 2 As shown, the system provided in the embodiment of the present application mainly includes:
[0090] The conversion module 210 is used to obtain data images of various types, determine the corresponding image preprocessing algorithm based on the specific type of the data image, and convert the data image into structured data; wherein the types of data images include at least text data images, table data images and map data images; and add label data to the structured data.
[0091] The feature acquisition module 220 is used to determine the corresponding feature extraction algorithm according to the image type corresponding to the structured data, and obtain the feature data set corresponding to each image type; wherein the feature data set includes label data.
[0092] The matrix acquisition module 230 is used to obtain the alignment factor, and then obtain several feature data sets belonging to the same alignment factor according to the alignment factor; wherein the alignment factor includes at least one or more of the following: time, unique identifier, preset content keyword, event name, object name, geographic coordinates, and area range; several feature data sets belonging to the same alignment factor are spliced into a feature vector to obtain the normalized data matrix and covariance matrix of the feature vector; the principal component space is calculated through the covariance matrix, and the normalized data matrix is projected into the principal component space to obtain the data matrix after dimensionality reduction.
[0093] The data fusion display module 240 is used to train a preset deep learning model through the reduced-dimensional data matrix and the corresponding label data to obtain a trained preset deep learning model; when a user retrieves information, the corresponding alignment factor is determined according to the user retrieval information, and then the corresponding reduced-dimensional data matrix is determined; the reduced-dimensional data matrix is used as the input of the trained preset deep learning model to obtain the label data; and then the data images of various types corresponding to the label data are obtained.
[0094] In addition, an embodiment of the present application further provides a non-volatile computer storage medium on which executable instructions are stored. When the executable instructions are executed, a multimodal data feature processing method based on deep learning as described above is implemented.
[0095] So far, the technical solutions of the present disclosure have been described in combination with the above multiple embodiments, but it is easy for those skilled in the art to understand that the protection scope of the present disclosure is not limited to these specific embodiments. Without departing from the technical principles of the present disclosure, those skilled in the art can split and combine the technical solutions in the above-mentioned various embodiments, and can also make equivalent changes or replacements to the relevant technical features. Any changes, equivalent replacements, improvements, etc. made within the technical concept and / or technical principle of the present disclosure will fall within the protection scope of the present disclosure.
Claims
1. A multimodal data feature processing method based on deep learning, characterized in that: The method comprises: Acquire various types of data images, determine corresponding image preprocessing algorithms based on specific types of the data images, and convert the data images into structured data; wherein the types of the data images at least include text data images, table data images, and map data images; add label data to the structured data; According to the image type corresponding to the structured data, a corresponding feature extraction algorithm is determined to obtain a feature data set corresponding to each image type; wherein the feature data set includes label data; Acquire an alignment factor, and then acquire a plurality of feature data sets belonging to the same alignment factor according to the alignment factor; wherein the alignment factor includes at least one or more of the following: time, unique identifier, preset content keyword, event name, object name, geographic coordinates, and area range; Several feature data sets belonging to the same alignment factor are concatenated into a feature vector to obtain the normalized data matrix and covariance matrix of the feature vector; the principal component space is calculated through the covariance matrix, and the normalized data matrix is projected into the principal component space to obtain the data matrix after dimensionality reduction; The preset deep learning model is trained by the data matrix after dimensionality reduction and the corresponding label data to obtain the trained preset deep learning model; when performing user retrieval information, the corresponding alignment factor is determined according to the user retrieval information, and then the corresponding data matrix after dimensionality reduction is determined; specifically, the method includes: extracting the user alignment factor from the user retrieval information, obtaining a number of feature data sets belonging to the user alignment factor, splicing them into a feature vector, and obtaining a normalized data matrix and a covariance matrix of the feature vector; calculating and obtaining the principal component space through the covariance matrix, projecting the normalized data matrix into the principal component space, and obtaining the data matrix after dimensionality reduction; The reduced-dimensional data matrix is used as the input of a trained preset deep learning model to obtain label data; and then various types of data images corresponding to the label data are obtained.
2. The multimodal data feature processing method based on deep learning according to claim 1, characterized in that: Obtain various types of data images, determine the corresponding image preprocessing algorithm based on the specific type of the data image, and convert the data image into structured data, including: When the data image is a text data image, the text data image is converted into structured data in an editable text format by using OCR technology; When the data image is a table-type data image, the table-type data image is grayed and binarized through image preprocessing, and the color table-type data image is converted into a black and white table-type data image; the median filtering technology is used to remove the noise in the table-type data image, and the tilt of the table-type data image is corrected; the edge detection algorithm is used to find the lines in the table-type data image, the Hough line detection algorithm is used to detect the straight lines in the lines, and the contour detection function is used to find all the closed contours in the straight lines; each closed contour is cut out, and the OCR technology is used to recognize the text in each closed contour, so as to obtain the structured data corresponding to the table-type data image; When the data image is a map-type data image, the geographic information in the map-type data image is extracted using geographic information system technology combined with OCR technology as structured data of the current data image; wherein the geographic information includes at least longitude, latitude, street name, and street number.
3. The multimodal data feature processing method based on deep learning according to claim 1, characterized in that: Add label data to structured data, including: Obtain label data corresponding to each structured data through the preset data upload terminal; And by using a preset semantic recognition algorithm, a preset annotation keyword corresponding to the structured data is identified, and the label data corresponding to the preset annotation keyword is determined to be the label data corresponding to the current structured data.
4. The multimodal data feature processing method based on deep learning according to claim 1, characterized in that: According to the image type corresponding to the structured data, the corresponding feature extraction algorithm is determined to obtain the feature data set corresponding to each image type, specifically including: When the structured data corresponds to a text-type data image, a text feature extraction algorithm is called to extract vector feature data in the structured data to generate a feature data set corresponding to the current text-type data image; When the structured data corresponds to a table-type data image, identify the data type of each column; wherein the data type includes at least: numerical type and categorical type; use normalization to process the numerical structured data into feature data; use one-hot encoding to process the categorical structured data into binary feature data, and obtain a feature data set corresponding to the current table-type data image; When the structured data corresponds to a map-type data image, geographic coordinate features and regional attribute features are extracted from the structured data; the preset clustering label corresponding to the current geographic coordinate is determined by a preset geographic coordinate feature set including the current geographic coordinate feature and a K-Means clustering algorithm; a preset number of geographic coordinate features closest to the current geographic coordinate feature are obtained as spatial relationship features by a preset geographic coordinate feature set including the current geographic coordinate feature and a haversine method; the geographic coordinate features, regional attribute features, preset clustering labels and spatial relationship features are added to the feature data set corresponding to the current map-type data image.
5. The multimodal data feature processing method based on deep learning according to claim 1, characterized in that: After obtaining the alignment factor, and then obtaining several feature data sets belonging to the same alignment factor according to the alignment factor, the method includes: Confirm whether the data formats in several feature data sets belonging to the same alignment factor are consistent; When there is inconsistency in the data format, the inconsistent data format is modified to a preset unified format.
6. The multimodal data feature processing method based on deep learning according to claim 1, characterized in that: Several feature data sets belonging to the same alignment factor are concatenated into a feature vector to obtain the normalized data matrix and covariance matrix of the feature vector; The principal component space is calculated through the covariance matrix, and the normalized data matrix is projected into the principal component space to obtain the data matrix after dimensionality reduction, which includes: The keywords in several feature data sets belonging to the same alignment factor are concatenated together to generate a total feature set, and the total feature set is converted into a feature vector; wherein the feature vector is a data matrix of m features of n feature data sets; Use MinMaxScaler to normalize the feature vector to obtain the normalized data matrix; By formula: , calculate and obtain the covariance matrix Y; in, represents the normalized data matrix, n represents the total number of feature data sets belonging to the same alignment factor; Use the numpy.linalg.eig function in the python function library to calculate the one-dimensional eigenvalue array of the covariance matrix; Take the first K values in the one-dimensional eigenvalue array as the principal component space; By formula: , project the normalized data matrix into the principal component space to obtain the reduced-dimensional data matrix Z; in, represents the principal component space.
7. The multimodal data feature processing method based on deep learning according to claim 1, characterized in that: After acquiring each type of data image corresponding to the label data, the method further includes: Vectorize each type of data image into one total matrix; The Pearson correlation coefficient between each two data image vectors in the total matrix is calculated using the np.corrcoef function; Generate an association relationship graph of the data images, wherein the association relationship graph consists of nodes and edges, and the nodes represent the data images and the edges represent the correlation coefficients between the two data images.
8. A multimodal data feature processing system based on deep learning, characterized in that: The system comprises: A conversion module is used to obtain various types of data images, determine the corresponding image preprocessing algorithm based on the specific type of the data image, and convert the data image into structured data; wherein the types of data images at least include text data images, table data images and map data images; add label data to the structured data; A feature acquisition module is used to determine a corresponding feature extraction algorithm according to the image type corresponding to the structured data, and obtain a feature data set corresponding to each image type; wherein the feature data set includes label data; A matrix acquisition module is used to obtain an alignment factor, and then obtain several feature data sets belonging to the same alignment factor according to the alignment factor; wherein the alignment factor includes at least one or more of the following: time, unique identifier, preset content keyword, event name, object name, geographic coordinates, and area range; several feature data sets belonging to the same alignment factor are spliced into a feature vector to obtain a normalized data matrix and a covariance matrix of the feature vector; a principal component space is calculated through the covariance matrix, and the normalized data matrix is projected into the principal component space to obtain a data matrix after dimensionality reduction; The data fusion display module is used to train a preset deep learning model through a data matrix after dimensionality reduction and corresponding label data to obtain a trained preset deep learning model; when performing user retrieval information, the corresponding alignment factor is determined according to the user retrieval information, and then the corresponding data matrix after dimensionality reduction is determined; specifically, the module includes: extracting the user alignment factor from the user retrieval information, obtaining a number of feature data sets belonging to the user alignment factor, splicing them into a feature vector, and obtaining a normalized data matrix and a covariance matrix of the feature vector; calculating and obtaining the principal component space through the covariance matrix, projecting the normalized data matrix into the principal component space, and obtaining the data matrix after dimensionality reduction; using the data matrix after dimensionality reduction as the input of the trained preset deep learning model to obtain label data; and then obtaining various types of data images corresponding to the label data.
9. A non-volatile computer storage medium, characterized in that: Computer instructions are stored thereon, and when the computer instructions are executed, they implement a multimodal data feature processing method based on deep learning as described in any one of claims 1-7.
Citation Information
Patent Citations
Multi-modal named entity recognition method and device, equipment and storage medium
CN116151263A
Multi-modal data processing method and device based on unified representation model, equipment and medium
CN119272229A