Multi-modal data processing method and device, electronic equipment and storage medium

By performing hierarchical clustering and synonym/antonym expansion on text data to generate a text relationship graph, the problem of low efficiency in forming sample pairs of text and non-text data in existing technologies is solved, achieving more efficient and accurate multimodal data processing.

CN120974176APending Publication Date: 2025-11-18ZHEJIANG DAHUA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510968621.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing technologies directly pair text data with non-text data for comparative learning, which limits efficiency and effectiveness. In particular, it is difficult to achieve accurate modality alignment in open-set texts, and there is a lack of structured learning.

Method used

By performing hierarchical clustering on multiple text data, a text relationship graph is generated, the synonyms and antonyms of each node are determined, positive and negative text data are expanded, more balanced sample pairs are generated, and the multimodal model is optimized through multiple iterations of training.

Benefits of technology

It improves the efficiency and accuracy of contrastive learning, enhances the model's learning of text structure, generates more balanced positive and negative sample pairs, and improves the effect of multimodal data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120974176A_ABST
    Figure CN120974176A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a multi-modal data processing method and device, electronic equipment and a storage medium, and the method comprises the steps: carrying out the hierarchical clustering of a plurality of pieces of text data, and obtaining a text relation graph; the text relation graph comprises nodes corresponding to a plurality of pieces of text data; for any node, determining a first text which is the same as the text data semantics corresponding to the node, and determining a second text which is opposite to the text data semantics corresponding to the node; for non-text data corresponding to any node, combining the non-text data with each forward text and each second text corresponding to the node to obtain a sample pair; wherein the forward text of the node comprises the text data corresponding to the node and the first text, and further comprises the text data corresponding to the superior node of the node and the first text. According to the embodiment, not only are more forward texts obtained, but also the more targeted second text is obtained, and the efficiency and effect of comparative learning are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a multimodal data processing method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the rapid development of artificial intelligence and information technology, multimodal technology has emerged and gradually become a research hotspot. Information in the real world exhibits multimodal characteristics. Humans perceive the world through multiple senses, such as vision, hearing, and touch, in a coordinated manner. Single-modal data processing is insufficient to meet the demands for comprehensive and accurate understanding of information in complex scenarios. Multimodal technology aims to integrate data from different modalities, such as images, text, audio, and video, to fully explore the correlations and complementary information between these modalities, thereby improving the performance and cognitive capabilities of intelligent systems.

[0003] In related technologies, text data and non-text data are directly paired to form sample pairs for comparative learning. During the training process of comparative learning, data of different modalities are matched and matched in terms of feature space and semantic space, which is to say, modality alignment is performed.

[0004] However, due to the diversity and relevance of texts, directly pairing text data with non-text data may affect the efficiency and effectiveness of contrastive learning. Summary of the Invention

[0005] This application provides a multimodal data processing method, apparatus, electronic device, and storage medium to improve the efficiency and effectiveness of comparative learning.

[0006] In a first aspect, embodiments of this application provide a first multimodal data processing method, the method comprising:

[0007] Hierarchical clustering is performed on multiple text data to obtain a text relationship graph; wherein, the text relationship graph contains the nodes corresponding to each of the multiple text data;

[0008] For any given node, determine a first text that has the same semantic meaning as the text data corresponding to the node, and determine a second text that has the opposite semantic meaning to the text data corresponding to the node;

[0009] For any non-text data corresponding to a node, the non-text data is combined with each positive text and each second text corresponding to the node to obtain a sample pair; wherein, the positive text of the node includes the text data corresponding to the node and the first text, and also includes the text data corresponding to the parent node of the node and the first text.

[0010] In some optional implementations, determining a second text with semantics opposite to the text data corresponding to the node includes:

[0011] Based on a preset correspondence, a second text with the opposite semantics to the text data corresponding to the node is determined; wherein, the preset correspondence includes the negative text corresponding to each of the multiple preset texts.

[0012] In some optional implementations, based on a preset correspondence, a second text with semantically opposite meaning to the text data corresponding to the node is determined, including:

[0013] Determine whether the node corresponds to text data in the plurality of preset texts;

[0014] If so, then the negative text corresponding to the text data in the preset correspondence relationship shall be taken as the second text;

[0015] Otherwise, select the second text from the text data corresponding to the sibling nodes of the node.

[0016] In some alternative implementations, hierarchical clustering is performed on multiple text data to obtain a text relationship graph, including:

[0017] Feature extraction is performed on multiple text data to obtain the text features corresponding to each of the multiple text data;

[0018] Hierarchical clustering is performed on the obtained text features to obtain the text relationship graph.

[0019] In some optional implementations, the non-text data is combined with each of the forward texts and each of the second texts corresponding to the node to obtain sample pairs, including:

[0020] The non-text data is combined with the corresponding positive text of each node to obtain positive sample pairs; and

[0021] The non-text data is combined with the second text corresponding to each node to obtain negative sample pairs.

[0022] In some optional implementations, after combining the non-text data with each forward text and each second text corresponding to the node to obtain sample pairs, the method further includes:

[0023] Based on the selected sample pairs, the multimodal model is trained iteratively multiple times; each iteration includes:

[0024] A positive loss is obtained based on the positive feature pairs corresponding to the selected positive sample pairs; and a negative loss is obtained based on the negative feature pairs corresponding to the selected negative sample pairs.

[0025] The multimodal model is tuned based on all the obtained positive and negative losses.

[0026] In some optional implementations, the multimodal model is hyperparameterized based on all obtained positive and negative losses, including:

[0027] A first loss is obtained based on the positive loss and the positive weights; and a second loss is obtained based on the negative loss and the negative weights corresponding to the types of the negative sample pairs.

[0028] A target loss is determined based on the first loss and the second loss, and the multimodal model is tuned based on the target loss.

[0029] Secondly, embodiments of this application provide a first type of multimodal data processing apparatus, the apparatus comprising:

[0030] A relationship building module is used to perform hierarchical clustering on multiple text data to obtain a text relationship graph; wherein, the text relationship graph contains the nodes corresponding to each of the multiple text data;

[0031] The text expansion module is used to determine, for any given node, a first text with the same semantic meaning as the text data corresponding to the node, and a second text with the opposite semantic meaning to the text data corresponding to the node;

[0032] The sampling module is used to combine the non-text data corresponding to any node with each positive text and each second text corresponding to the node to obtain sample pairs; wherein, the positive text of the node includes the text data corresponding to the node and the first text, and also includes the text data corresponding to the parent node of the node and the first text.

[0033] In some alternative implementations, the text expansion module is specifically used for:

[0034] Based on a preset correspondence, a second text with the opposite semantics to the text data corresponding to the node is determined; wherein, the preset correspondence includes the negative text corresponding to each of the multiple preset texts.

[0035] In some alternative implementations, the text expansion module is specifically used for:

[0036] Determine whether the node corresponds to text data in the plurality of preset texts;

[0037] If so, then the negative text corresponding to the text data in the preset correspondence relationship shall be taken as the second text;

[0038] Otherwise, select the second text from the text data corresponding to the sibling nodes of the node.

[0039] In some optional implementations, the relationship building module is specifically used for:

[0040] Feature extraction is performed on multiple text data to obtain the text features corresponding to each of the multiple text data;

[0041] Hierarchical clustering is performed on the obtained text features to obtain the text relationship graph.

[0042] In some alternative implementations, the sampling module is specifically used for:

[0043] The non-text data is combined with the corresponding positive text of each node to obtain positive sample pairs; and

[0044] The non-text data is combined with the second text corresponding to each node to obtain negative sample pairs.

[0045] In some optional implementations, a training module is also included, which, after the sampling module combines the non-textual data with each forward text and each second text corresponding to the node to obtain sample pairs, is used for:

[0046] Based on the selected sample pairs, the multimodal model is trained iteratively multiple times; each iteration includes:

[0047] A positive loss is obtained based on the positive feature pairs corresponding to the selected positive sample pairs; and a negative loss is obtained based on the negative feature pairs corresponding to the selected negative sample pairs.

[0048] The multimodal model is tuned based on all the obtained positive and negative losses.

[0049] In some alternative implementations, the training module is specifically used for:

[0050] A first loss is obtained based on the positive loss and the positive weights; and a second loss is obtained based on the negative loss and the negative weights corresponding to the types of the negative sample pairs.

[0051] A target loss is determined based on the first loss and the second loss, and the multimodal model is tuned based on the target loss.

[0052] Thirdly, embodiments of this application provide an electronic device, including at least one processor and at least one memory, wherein the memory stores a computer program, and when the program is executed by the processor, the processor performs the multimodal data processing method described in any of the first aspects above.

[0053] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program executable by a processor, which, when run on the processor, causes the processor to perform the multimodal data processing method described in any of the first aspects above.

[0054] In this embodiment, considering the correlation of text, a text relationship graph representing the correlation between multiple text data is obtained by hierarchically clustering multiple text data. Based on this text relationship graph, the correlation between nodes can be known, such as superior-subordinate nodes, sibling nodes, etc. Furthermore, each node is expanded with synonyms to obtain synonyms (first text) for each text data, and each node is expanded with antonyms instead of directly selecting text data from other nodes as antonyms to obtain more targeted antonyms (second text) for the text data. When generating sample pairs, not only are the text data of each node and the first text used as the positive text of that node, but also the text data of its superior node and the first text are used as the positive text of that node. This increases the number of positive texts and makes the relationship between positive and negative texts more balanced, improving the efficiency of subsequent comparative learning. In addition, since more targeted antonyms (second text) for the text data are generated, the second text is used as the negative text of that node, improving the accuracy of comparative learning. Attached Figure Description

[0055] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0056] Figure 1 A flowchart illustrating the first multimodal data processing method provided in this application embodiment;

[0057] Figure 2 This is a schematic diagram of text relationship provided for an embodiment of this application;

[0058] Figure 3 A flowchart illustrating the second multimodal data processing method provided in this application embodiment;

[0059] Figure 4 This is a schematic diagram of the sample pair distribution under relevant technologies;

[0060] Figure 5 This is a schematic diagram of the sample pair distribution provided in the embodiments of this application;

[0061] Figure 6 A flowchart illustrating the third multimodal data processing method provided in this application embodiment;

[0062] Figure 7 This is a schematic diagram of the structure of the multimodal data processing device provided in the embodiments of this application;

[0063] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0064] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0065] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "a plurality of" means two or more.

[0066] In the description of this application, it should be noted that, unless otherwise expressly specified and limited, the term "connection" should be interpreted broadly. For example, it can refer to a direct connection, an indirect connection through an intermediate medium, or a connection within two devices. Those skilled in the art can understand the specific meaning of the above term in this application based on the specific circumstances.

[0067] With the rapid development of artificial intelligence and information technology, multimodal technology has emerged and gradually become a research hotspot. Information in the real world exhibits multimodal characteristics. Humans perceive the world through multiple senses, such as vision, hearing, and touch, in a coordinated manner. Single-modal data processing is insufficient to meet the demands for comprehensive and accurate understanding of information in complex scenarios. Multimodal technology aims to integrate data from different modalities, such as images, text, audio, and video, to fully explore the correlations and complementary information between these modalities, thereby improving the performance and cognitive capabilities of intelligent systems.

[0068] In the medical field, integrating medical images and electronic medical records can assist in disease diagnosis and treatment planning. In intelligent driving, vehicles need to integrate multimodal data such as camera images, radar point clouds, and in-vehicle voice commands to achieve precise environmental perception and decision control. In the education field, the multimodal fusion of multimedia teaching resources can provide richer materials and interactive methods for personalized learning. In intelligent security, combining multimodal data such as video surveillance images, audio information, and personnel identity information can improve the efficiency and accuracy of security monitoring.

[0069] In related technologies, text data and non-text data are directly paired to form sample pairs for comparative learning. During the training process of comparative learning, data of different modalities are matched and matched in terms of feature space and semantic space, which is to say, modality alignment is performed.

[0070] However, due to the diversity and relevance of text, directly pairing textual and non-textual data may affect the efficiency and effectiveness of contrastive learning. This is especially true for open-set texts, where achieving accurate modality alignment is difficult.

[0071] In some embodiments, the text data is clustered directly, making the clustering classification more accurate.

[0072] However, this approach lacks understanding of hierarchical relationships within text. In multimodal learning, information can only be obtained at the semantic cluster level through feature associations between targets, but not at the hierarchical information between clusters, thus lacking structured learning.

[0073] In some embodiments, a dual-tower network model is used to extract features from multimodal data (text data and non-text data), and the generated multimodal fusion features are spliced ​​together before comparative learning is performed.

[0074] However, the above method, for a non-sample data set, only has one positive text data (only one positive sample pair) and lacks targeted negative text data.

[0075] In view of this, embodiments of this application propose a multimodal data processing method, apparatus, electronic device, and storage medium to improve the efficiency and effectiveness of contrastive learning. The method includes: performing hierarchical clustering on multiple text data to obtain a text relationship graph; wherein the text relationship graph includes nodes corresponding to each of the multiple text data; for any node, determining a first text with the same semantics as the text data corresponding to the node, and determining a second text with the opposite semantics to the text data corresponding to the node; for any non-text data corresponding to a node, combining the non-text data with each positive text and each second text corresponding to the node to obtain sample pairs; wherein the positive text of the node includes the text data corresponding to the node and the first text, and also includes the text data corresponding to the parent node of the node and the first text.

[0076] The technical solution of this application and how it solves the above-mentioned technical problems will be described in detail below with reference to the accompanying drawings and specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0077] Figure 1 This is a flowchart illustrating the first multimodal data processing method provided in the embodiments of this application, applied to electronic devices, such as... Figure 1 As shown, it includes the following steps:

[0078] Step S101: Perform hierarchical clustering on multiple text data to obtain a text relationship graph; wherein the text relationship graph contains the nodes corresponding to each of the multiple text data.

[0079] In practice, if ordinary clustering is used, only multiple semantic clusters can be obtained, and the hierarchical information between clusters cannot be obtained; while texts are usually related, such as people and men, which have a hierarchical relationship, and men are included in the scope of people.

[0080] Based on this, this embodiment obtains a text relationship graph that represents the relationship between multiple text data by performing hierarchical clustering on multiple text data. Based on this text relationship graph, the relationship between nodes can be known, such as superior-subordinate nodes, sibling nodes, etc.

[0081] See Figure 2 As shown, hierarchical clustering forms a multi-branch tree structure, which not only contains different semantic clusters (cluster 1 to cluster N), but also contains hierarchical nodes within each semantic cluster.

[0082] For example, the top level of cluster 1 is the node corresponding to a person, the next level has two nodes corresponding to different genders (men and women), and the next level has nodes corresponding to various professions (such as salesperson, teacher, etc.) for each gender.

[0083] By encoding the nodes, information identifying each node is obtained.

[0084] Step S102: For any node, determine a first text with the same semantic meaning as the text data corresponding to the node, and determine a second text with the opposite semantic meaning to the text data corresponding to the node.

[0085] In practice, if the text data and corresponding non-text data of each node are directly used as positive sample pairs, the number of positive sample pairs will be small. The imbalance of positive and negative sample pairs will further affect the efficiency and effect of contrastive learning. Furthermore, due to the diversity of texts, contrastive learning will also suffer from insufficient learning of similar parts of speech.

[0086] Based on this, this embodiment expands each node with synonyms to obtain the synonyms (first text) of each text data.

[0087] In some optional implementations, after obtaining the first text, further filtering operations such as deduplication can be performed on the first text to obtain the filtered first text.

[0088] Furthermore, for a given node, directly using the text data of other nodes as the negative text of the non-text data of that node will affect the model's learning. For example, if "teacher" is the positive text, using "person" and "man" as negative text will interfere with the model's learning, because "teacher" and "person" and "man" are not opposites.

[0089] Based on this, this embodiment does not directly use the text data of other nodes as the negative text of the non-text data of the node. Instead, it expands each node with antonyms to obtain antonyms (second text) that are more targeted to the text data.

[0090] In practice, if the length of the text corresponding to a node (the original text data, the generated first text, and the second text) exceeds the preset token length, token truncation can be performed.

[0091] Step S103: For any non-text data corresponding to a node, combine the non-text data with each positive text and each second text corresponding to the node to obtain a sample pair; wherein, the positive text of the node includes the text data and the first text corresponding to the node, and also includes the text data and the first text corresponding to the parent node of the node.

[0092] This embodiment generates a text relationship graph representing the connections between multiple text data. Based on this text relationship graph, the relationships between nodes can be determined, such as hierarchical nodes, sibling nodes, etc. Since the text of a lower-level node is contained within the text of its higher-level node (e.g., "man" is contained within "person"), not only can the text data corresponding to a node and its first text be considered as positive text, but also the text data corresponding to its higher-level node and its first text can be considered as positive text. This not only further expands the number of positive texts but also strengthens the subsequent model's learning of text structure.

[0093] The above scheme, considering the correlation of text, performs hierarchical clustering on multiple text data to obtain a text relationship graph representing the correlation between multiple text data. Based on this text relationship graph, the correlation between nodes can be identified, such as superior-subordinate nodes, sibling nodes, etc. Furthermore, each node is expanded with synonyms to obtain synonyms for each text data (first text), and each node is expanded with antonyms instead of directly selecting text data from other nodes as antonyms to obtain more targeted antonyms for the text data (second text). When generating sample pairs, not only are the text data of each node and the first text used as positive text for that node, but also the text data of its superior node and the first text are used as positive text for that node. This increases the number of positive texts and makes the relationship between positive and negative texts more balanced, improving the efficiency of subsequent comparative learning. In addition, since more targeted antonyms (second text) are generated for the text data, the second text is used as the negative text for that node, improving the accuracy of comparative learning.

[0094] In some optional implementations, step S101 above can be implemented in, but is not limited to, the following ways:

[0095] Feature extraction is performed on multiple text data to obtain the text features corresponding to each of the multiple text data;

[0096] Hierarchical clustering is performed on the obtained text features to obtain the text relationship graph.

[0097] Since the core of hierarchical clustering is to calculate the distance or similarity between data points, it is difficult to directly calculate the distance or similarity of raw text data;

[0098] Based on this, this embodiment first extracts features from multiple text data to obtain the text features corresponding to each text data; further, it performs hierarchical clustering on the obtained multiple text features to obtain a text relationship graph.

[0099] This embodiment does not specifically limit the above feature extraction method. For example, a bidirectional Encoder Representation from Transformers model can be used to extract features from text data.

[0100] In some optional implementations, determining the second text in step S102 above can be achieved in, but is not limited to, the following ways:

[0101] Based on a preset correspondence, a second text with the opposite semantics to the text data corresponding to the node is determined; wherein, the preset correspondence includes the negative text corresponding to each of the multiple preset texts.

[0102] In practice, some text data (proprietary text) has specific antonyms, such as the antonym of "man" being "woman," or the antonym of "white T-shirt" being "T-shirt of other colors," making it difficult to accurately select the antonyms of this text data directly from other nodes; while other text data (general text) does not have specific antonyms, such as "person," "animal," and "window," making it difficult to summarize the antonyms of this text data.

[0103] Based on this, this embodiment uses proprietary text as preset text and establishes a preset correspondence relationship including the negative texts corresponding to each of the multiple preset texts. In this way, by querying the preset correspondence relationship, the specific antonyms of the proprietary text can be found.

[0104] During implementation, the preset correspondence can be dynamically adjusted based on the changes in the proprietary text.

[0105] In some optional implementations, the second text can be determined from the preset correspondence in the following ways:

[0106] Determine whether the node corresponds to text data in the plurality of preset texts;

[0107] If so, then the negative text corresponding to the text data in the preset correspondence relationship shall be taken as the second text;

[0108] Otherwise, select the second text from the text data corresponding to the sibling nodes of the node.

[0109] In practice, by comparing the text data corresponding to the node with multiple preset texts in the preset correspondence, if multiple preset texts contain the above text data, it indicates that the text data is proprietary text. By querying the negative text corresponding to the text data in the preset correspondence, the antonym (second text) specific to the text data can be determined.

[0110] Conversely, if multiple preset texts do not contain the text data corresponding to the node, it means that the text data is not proprietary text, but general text, and it is difficult to summarize the antonyms of these text data. Since the sibling node is a node at the same level as the node and has no subordinate relationship, selecting the second text of the node from the text data corresponding to the sibling node will not cause trouble in the training process.

[0111] Figure 3 A flowchart illustrating the second multimodal data processing method provided in this application embodiment is shown below. Figure 3 As shown, it includes the following steps:

[0112] Step S301: Perform hierarchical clustering on multiple text data to obtain a text relationship graph; wherein the text relationship graph contains the nodes corresponding to each of the multiple text data.

[0113] Step S302: For any node, determine a first text with the same semantic meaning as the text data corresponding to the node, and determine a second text with the opposite semantic meaning to the text data corresponding to the node.

[0114] The specific implementation of steps S301 to S302 can be referred to the above embodiments, and will not be repeated here.

[0115] Step S303: For any non-text data corresponding to a node, combine the non-text data with each positive text corresponding to the node to obtain a positive sample pair; and combine the non-text data with each second text corresponding to the node to obtain a negative sample pair.

[0116] In this embodiment, by expanding the positive text and combining the non-text data with each positive text, a positive sample pair is obtained. In this way, multiple positive sample pairs for the non-text data are obtained. By combining the non-text data with more targeted antonyms, more accurate negative sample pairs are obtained.

[0117] See Figure 4 As shown, taking m sets of text data and non-text data as an example, before expanding the forward text, for each non-text data I... i There is only one positive text (the original text data T). i Therefore, only one positive sample pair I can be obtained. i T i(Only the diagonal pairs are positive sample pairs, and the rest are negative sample pairs), resulting in a severe imbalance between positive and negative sample pairs. Furthermore, for a non-textual data point, combining all other textual data with that non-textual data to obtain a negative sample pair can cause problems in the training process. For example, combining "person" with "man" to obtain a negative sample pair, or combining "salesperson" with "man" to obtain a negative sample pair, is a subordinate relationship, not a completely opposing one. This method of obtaining negative sample pairs is unreasonable.

[0118] See Figure 5 As shown, by expanding the positive text, T i Expand to obtain T i1 ~T iN These N positive texts, and t i1 ~t in These n second texts, for each non-text data I i Since there are N positive texts, N positive sample pairs can be obtained. Furthermore, in this embodiment, for a non-text data, instead of combining all other text data with the non-text data to obtain a negative sample pair, a second text from a sibling node is selected and combined with the non-text data to obtain a negative sample pair.

[0119] Figure 6 A flowchart illustrating the third multimodal data processing method provided in this application embodiment is shown below. Figure 6 As shown, it includes the following steps:

[0120] Step S601: Perform hierarchical clustering on multiple text data to obtain a text relationship graph; wherein the text relationship graph contains the nodes corresponding to each of the multiple text data.

[0121] Step S602: For any node, determine a first text with the same semantic meaning as the text data corresponding to the node, and determine a second text with the opposite semantic meaning to the text data corresponding to the node.

[0122] Step S603: For any non-text data corresponding to a node, combine the non-text data with each positive text corresponding to the node to obtain a positive sample pair; and combine the non-text data with each second text corresponding to the node to obtain a negative sample pair.

[0123] The specific implementation of steps S601 to S602 can be referred to the above embodiments, and will not be repeated here.

[0124] Step S604: Based on the selected sample pairs, perform multiple iterations of training for the multimodal model.

[0125] Each iteration includes:

[0126] A positive loss is obtained based on the positive feature pairs corresponding to the selected positive sample pairs; and a negative loss is obtained based on the negative feature pairs corresponding to the selected negative sample pairs.

[0127] The multimodal model is tuned based on all the obtained positive and negative losses.

[0128] In practice, after obtaining positive and negative sample pairs, positive and negative sample pairs can be selected to train the multimodal model.

[0129] For example, for each selected positive sample pair, the text features of the positive text (positive text features) and the non-text features of the non-text data in the positive sample pair are determined. Based on the positive text features and non-text features, the positive loss is determined. Among them, the similarity between the positive text features and non-text features is positively correlated with the positive loss.

[0130] For each selected negative sample pair, the text features (negative text features) of the second text in the negative sample pair and the non-text features of the non-text data are determined. Based on the negative text features and non-text features, the negative loss is determined. The similarity between the negative text features and non-text features is negatively correlated with the negative loss.

[0131] Furthermore, the multimodal model is tuned based on all positive and negative losses.

[0132] In practice, the above parameter tuning process can be achieved through, but is not limited to, the following methods:

[0133] A first loss is obtained based on the positive loss and the positive weights; and a second loss is obtained based on the negative loss and the negative weights corresponding to the types of the negative sample pairs.

[0134] A target loss is determined based on the first loss and the second loss, and the multimodal model is tuned based on the target loss.

[0135] In implementation, since positive and negative sample pairs have different impacts on training, it is necessary to set positive weights for positive sample pairs and negative weights for negative sample pairs. Furthermore, different methods can be used to determine the second text, generating different types of negative sample pairs. These different types of negative sample pairs also have different impacts on training; therefore, negative weights need to be divided into negative weights corresponding to different types of negative sample pairs. For example, based on the presence or absence of specific antonyms, text data can be divided into proprietary text and general text. Therefore, there will be negative sample pairs corresponding to proprietary text (proprietary negative sample pairs) and negative sample pairs corresponding to general text (general negative sample pairs). The negative weights include a first negative weight corresponding to the proprietary negative sample pair and a second negative weight corresponding to the general negative sample pair.

[0136] Furthermore, a first loss is obtained based on the positive loss and positive weights; a second loss is obtained based on the negative loss and the negative weights corresponding to the types of negative sample pairs; and the target loss is determined based on the first loss and the second loss.

[0137] For example, the target weight LOSS = α*loss1 + β1*loss21 + β2*loss22; where α is the positive weight, loss1 is the positive loss of all positive sample pairs, β1 is the negative weight corresponding to the negative sample pairs of type 1, loss21 is the negative loss of all negative samples of type 1, β2 is the negative weight corresponding to the negative sample pairs of type 2, and loss22 is the negative loss of all negative sample pairs of type 2.

[0138] The above example uses two types of negative sample pairs. In practice, different dimensions may be used to obtain the second text, and there may be more or fewer types of negative sample pairs. This embodiment does not make specific limitations on this.

[0139] like Figure 7 As shown, this application embodiment provides a multimodal data processing device 700, which includes:

[0140] The relationship construction module 701 is used to perform hierarchical clustering on multiple text data to obtain a text relationship graph; wherein, the text relationship graph contains the nodes corresponding to each of the multiple text data;

[0141] The text expansion module 702 is used to determine, for any node, a first text with the same semantic meaning as the text data corresponding to the node, and a second text with the opposite semantic meaning to the text data corresponding to the node;

[0142] The sampling module 703 is used to combine the non-text data corresponding to any node with each positive text and each second text corresponding to the node to obtain sample pairs; wherein, the positive text of the node includes the text data corresponding to the node and the first text, and also includes the text data corresponding to the parent node of the node and the first text.

[0143] In some optional implementations, the text expansion module 702 is specifically used for:

[0144] Based on a preset correspondence, a second text with the opposite semantics to the text data corresponding to the node is determined; wherein, the preset correspondence includes the negative text corresponding to each of the multiple preset texts.

[0145] In some optional implementations, the text expansion module 702 is specifically used for:

[0146] Determine whether the node corresponds to text data in the plurality of preset texts;

[0147] If so, then the negative text corresponding to the text data in the preset correspondence relationship shall be taken as the second text;

[0148] Otherwise, select the second text from the text data corresponding to the sibling nodes of the node.

[0149] In some optional implementations, the relationship building module 701 is specifically used for:

[0150] Feature extraction is performed on multiple text data to obtain the text features corresponding to each of the multiple text data;

[0151] Hierarchical clustering is performed on the obtained text features to obtain the text relationship graph.

[0152] In some optional implementations, the sampling module 703 is specifically used for:

[0153] The non-text data is combined with the corresponding positive text of each node to obtain positive sample pairs; and

[0154] The non-text data is combined with the second text corresponding to each node to obtain negative sample pairs.

[0155] In some optional implementations, a training module 704 is also included, which, after the sampling module 703 combines the non-text data with each forward text and each second text corresponding to the node to obtain sample pairs, is used for:

[0156] Based on the selected sample pairs, the multimodal model is trained iteratively multiple times; each iteration includes:

[0157] A positive loss is obtained based on the positive feature pairs corresponding to the selected positive sample pairs; and a negative loss is obtained based on the negative feature pairs corresponding to the selected negative sample pairs.

[0158] The multimodal model is tuned based on all the obtained positive and negative losses.

[0159] In some alternative implementations, the training module 704 is specifically used for:

[0160] A first loss is obtained based on the positive loss and the positive weights; and a second loss is obtained based on the negative loss and the negative weights corresponding to the types of the negative sample pairs.

[0161] A target loss is determined based on the first loss and the second loss, and the multimodal model is tuned based on the target loss.

[0162] Since this device is the same as the device in the method of this application embodiment, and the principle of the device in solving the problem is similar to that of the method, the implementation of the device can be referred to the implementation of the method, and the repeated parts will not be described again.

[0163] Based on the same technical concept, this application also provides an electronic device 800, such as... Figure 8 As shown, it includes at least one processor 801 and a memory 802 connected to at least one processor. In this embodiment, the specific connection medium between the processor 801 and the memory 802 is not limited. Figure 8 Taking the connection between the processor 801 and the memory 802 via bus 803 as an example, the bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, Figure 8 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0164] The processor 801 is the control center of the electronic device. It can connect to various parts of the electronic device through various interfaces and lines, and performs data processing by running or executing instructions stored in the memory 802 and calling data stored in the memory 802. Optionally, the processor 801 may include one or more processing units. The processor 801 may integrate an application processor and a modem processor. The application processor mainly handles the operating system, user interface, and application programs, while the modem processor mainly handles issuing instructions. It is understood that the modem processor may not be integrated into the processor 801. In some embodiments, the processor 801 and the memory 802 may be implemented on the same chip; in some embodiments, they may also be implemented on separate chips.

[0165] The processor 801 can be a general-purpose processor, such as a CPU, digital signal processor, application-specific integrated circuit (ASIC), field-programmable gate array or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of the multimodal data processing method can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.

[0166] Memory 802, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Memory 802 may include at least one type of storage medium, such as flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic storage, magnetic disk, optical disk, etc. Memory 802 can be any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but is not limited thereto. In the embodiments of this application, memory 802 can also be a circuit or any other device capable of implementing storage functions for storing program instructions and / or data.

[0167] In this embodiment, the memory 802 stores a computer program, which, when executed by the processor 801, causes the processor 801 to perform the following:

[0168] Hierarchical clustering is performed on multiple text data to obtain a text relationship graph; wherein, the text relationship graph contains the nodes corresponding to each of the multiple text data;

[0169] For any given node, determine a first text that has the same semantic meaning as the text data corresponding to the node, and determine a second text that has the opposite semantic meaning to the text data corresponding to the node;

[0170] For any non-text data corresponding to a node, the non-text data is combined with each positive text and each second text corresponding to the node to obtain a sample pair; wherein, the positive text of the node includes the text data corresponding to the node and the first text, and also includes the text data corresponding to the parent node of the node and the first text.

[0171] In some alternative implementations, processor 801 specifically performs:

[0172] Based on a preset correspondence, a second text with the opposite semantics to the text data corresponding to the node is determined; wherein, the preset correspondence includes the negative text corresponding to each of the multiple preset texts.

[0173] In some alternative implementations, processor 801 specifically performs:

[0174] Determine whether the node corresponds to text data in the plurality of preset texts;

[0175] If so, then the negative text corresponding to the text data in the preset correspondence relationship shall be taken as the second text;

[0176] Otherwise, select the second text from the text data corresponding to the sibling nodes of the node.

[0177] In some alternative implementations, processor 801 specifically performs:

[0178] Feature extraction is performed on multiple text data to obtain the text features corresponding to each of the multiple text data;

[0179] Hierarchical clustering is performed on the obtained text features to obtain the text relationship graph.

[0180] In some alternative implementations, processor 801 specifically performs:

[0181] The non-text data is combined with the corresponding positive text of each node to obtain positive sample pairs; and

[0182] The non-text data is combined with the second text corresponding to each node to obtain negative sample pairs.

[0183] In some optional implementations, after combining the non-text data with each forward text and each second text corresponding to the node to obtain sample pairs, the processor 801 further executes:

[0184] Based on the selected sample pairs, the multimodal model is trained iteratively multiple times; each iteration includes:

[0185] A positive loss is obtained based on the positive feature pairs corresponding to the selected positive sample pairs; and a negative loss is obtained based on the negative feature pairs corresponding to the selected negative sample pairs.

[0186] The multimodal model is tuned based on all the obtained positive and negative losses.

[0187] In some alternative implementations, processor 801 specifically performs:

[0188] A first loss is obtained based on the positive loss and the positive weights; and a second loss is obtained based on the negative loss and the negative weights corresponding to the types of the negative sample pairs.

[0189] A target loss is determined based on the first loss and the second loss, and the multimodal model is tuned based on the target loss.

[0190] Since the electronic device is the same as the electronic device in the method of this application embodiment, and the principle of the electronic device in solving the problem is similar to that of the method, the implementation of the electronic device can refer to the implementation of the method, and the repeated parts will not be described again.

[0191] Based on the same technical concept, embodiments of this application also provide a computer-readable storage medium storing a computer program executable by a processor, which, when run on the processor, causes the processor to perform the steps of the above-described multimodal data processing method.

[0192] In some alternative implementations, various aspects of the multimodal data processing method provided in this application may also be implemented as a program product containing computer-executable instructions. When the program product is run on a computer device, the computer-executable instructions are used to cause the computer device to perform the steps of the multimodal data processing method according to the various exemplary embodiments of this application described above.

[0193] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0194] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0195] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0196] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0197] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0198] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A multimodal data processing method, characterized in that, The method includes: Hierarchical clustering is performed on multiple text data to obtain a text relationship graph; wherein, the text relationship graph contains the nodes corresponding to each of the multiple text data; For any given node, determine a first text that has the same semantic meaning as the text data corresponding to the node, and determine a second text that has the opposite semantic meaning to the text data corresponding to the node; For any non-text data corresponding to a node, the non-text data is combined with each positive text and each second text corresponding to the node to obtain a sample pair; wherein, the positive text of the node includes the text data corresponding to the node and the first text, and also includes the text data corresponding to the parent node of the node and the first text.

2. The method as described in claim 1, characterized in that, Determining a second text whose semantics are opposite to the text data corresponding to the node includes: Based on a preset correspondence, a second text with the opposite semantics to the text data corresponding to the node is determined; wherein, the preset correspondence includes the negative text corresponding to each of the multiple preset texts.

3. The method as described in claim 2, characterized in that, Based on a preset correspondence, a second text with semantically opposite meaning to the text data corresponding to the node is determined, including: Determine whether the node corresponds to text data in the plurality of preset texts; If so, then the negative text corresponding to the text data in the preset correspondence relationship shall be taken as the second text; Otherwise, select the second text from the text data corresponding to the sibling nodes of the node.

4. The method as described in claim 1, characterized in that, Hierarchical clustering of multiple text datasets yields a text relationship graph, including: Feature extraction is performed on multiple text data to obtain the text features corresponding to each of the multiple text data; Hierarchical clustering is performed on the obtained text features to obtain the text relationship graph.

5. The method as described in claim 1, characterized in that, The non-text data is combined with each of the forward texts and each of the second texts corresponding to the node to obtain sample pairs, including: The non-text data is combined with the corresponding positive text of each node to obtain positive sample pairs; and The non-text data is combined with the second text corresponding to each node to obtain negative sample pairs.

6. The method as described in claim 5, characterized in that, After combining the non-text data with each positive text and each second text corresponding to the node to obtain sample pairs, the method further includes: Based on the selected sample pairs, the multimodal model is trained iteratively multiple times; each iteration includes: A positive loss is obtained based on the positive feature pairs corresponding to the selected positive sample pairs; and a negative loss is obtained based on the negative feature pairs corresponding to the selected negative sample pairs. The multimodal model is tuned based on all the obtained positive and negative losses.

7. The method as described in claim 6, characterized in that, The multimodal model is tuned based on all obtained positive and negative losses, including: A first loss is obtained based on the positive loss and the positive weights; and a second loss is obtained based on the negative loss and the negative weights corresponding to the types of the negative sample pairs. A target loss is determined based on the first loss and the second loss, and the multimodal model is tuned based on the target loss.

8. A multimodal data processing device, characterized in that, The device includes: A relationship building module is used to perform hierarchical clustering on multiple text data to obtain a text relationship graph; wherein, the text relationship graph contains the nodes corresponding to each of the multiple text data; The text expansion module is used to determine, for any given node, a first text with the same semantic meaning as the text data corresponding to the node, and a second text with the opposite semantic meaning to the text data corresponding to the node; The sampling module is used to combine the non-text data corresponding to any node with each positive text and each second text corresponding to the node to obtain sample pairs; wherein, the positive text of the node includes the text data corresponding to the node and the first text, and also includes the text data corresponding to the parent node of the node and the first text.

9. An electronic device, characterized in that, It includes at least one processor and at least one memory, wherein the memory stores a computer program that, when executed by the processor, causes the processor to perform the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, It stores a computer program executable by a computer, which, when run on the computer, causes the computer to perform the method as described in any one of claims 1 to 7.