Cross-domain long tail data augmentation method, apparatus, device and medium

CN120974183BActive Publication Date: 2026-08-21PING AN TECH (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511059531.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2026-08-21
Estimated Expiration
2045-07-29

AI Technical Summary

Technical Problem

[0006]鉴于以上内容,有必要提供一种跨领域长尾数据增强方法、装置、设备及介质,旨在解决无法跨领域对长尾数据进行增强的问题

Benefits of technology

[0025]As can be seen from the above technical solutions, this invention can construct a target domain knowledge graph and a source domain knowledge graph, and perform cross-domain knowledge alignment between the target domain knowledge graph and the source domain knowledge graph to obtain a cross-domain knowledge graph, unifying knowledge from different domains into a shared semantic space, promoting the association and reuse of knowledge from different domains; it uses density peak clustering algorithm to identify long-tail scene data in multimodal initial features, and can accurately identify long-tail data through data distribution analysis; it uses generative adversarial networks and cross-domain knowledge graphs to enhance long-tail scene data, and adds supplementary features to multimodal initial features to obtain multimodal cross-domain enhanced features, supplementing the feature information of long-tail data and expanding the coverage of training data, thereby improving the processing capability for low-frequency, cross-domain and complex long-tail scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120974183B_ABST
    Figure CN120974183B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of artificial intelligence, and provides a cross-domain long-tail data enhancement method and device, equipment and a medium, which can be applied to financial, medical health and pension business scenarios, can construct a target domain knowledge graph and a source domain knowledge graph, and can perform cross-domain knowledge alignment on the target domain knowledge graph and the source domain knowledge graph, unify different domain knowledge to a shared semantic space, and promote the association and reuse of different domain knowledge; a density peak value clustering algorithm is used to identify long-tail scene data in multi-modal initial features, the long-tail data can be accurately identified through data distribution analysis; a generative adversarial network and the cross-domain knowledge graph are used to enhance the long-tail scene data, and supplementary features are added to the multi-modal initial features to obtain multi-modal cross-domain enhanced features, the feature information of the long-tail data is supplemented, and the training data coverage range is expanded, so that the processing capability for low-frequency, cross-domain and complex long-tail scenes is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a cross-domain long-tail data augmentation method, apparatus, device, and medium. Background Technology

[0002] Real-world multimodal data contains a large amount of low-frequency, complex, long-tail data, and traditional models are usually trained only on data from specific domains.

[0003] Because the model is limited by the distribution of training data and it is difficult to reuse existing domain knowledge in new fields, it is prone to misjudgment or inability to make decisions when encountering long-tail scenarios such as handling rare faults or operating in special environments.

[0004] For example, when a robot used in a financial services hall switches to a medical surgery scenario, the presence of long-tail data and the inability to transfer knowledge of object handling to the task of delivering surgical instruments not only cause the model to need to be retrained, but also result in poor training performance.

[0005] Therefore, how to combine cross-domain knowledge to enhance long-tail data and make it more suitable for VLA (Vision-Language-Action) tasks has become an urgent problem to be solved. Summary of the Invention

[0006] In view of the above, it is necessary to provide a cross-domain long-tail data augmentation method, apparatus, device and medium, which aims to solve the problem of the inability to augment long-tail data across domains.

[0007] A cross-domain long-tail data augmentation method, the cross-domain long-tail data augmentation method comprising:

[0008] In response to the instruction to enhance long-tail data within the target domain, a target domain knowledge graph of the target domain is constructed, and a source domain knowledge graph of the source domain is constructed.

[0009] Cross-domain knowledge alignment is performed between the target domain knowledge graph and the source domain knowledge graph to obtain a cross-domain knowledge graph;

[0010] The input data is obtained according to the enhancement instructions, and multimodal domain features are extracted from the input data to obtain multimodal initial features;

[0011] Density peak clustering algorithm is used to identify long-tailed scene data in the initial features of the multimodal dataset;

[0012] The long-tail scene data is enhanced using generative adversarial networks and the cross-domain knowledge graph to obtain supplementary features;

[0013] The supplementary features are added to the initial multimodal features to obtain multimodal cross-domain enhanced features.

[0014] A cross-domain long-tail data augmentation device, the cross-domain long-tail data augmentation device comprising:

[0015] The construction unit is configured to construct a target domain knowledge graph of the target domain and a source domain knowledge graph of the source domain in response to an enhancement instruction for long-tail data in the target domain.

[0016] An alignment unit is used to perform cross-domain knowledge alignment between the target domain knowledge graph and the source domain knowledge graph to obtain a cross-domain knowledge graph.

[0017] An extraction unit is used to acquire input data according to the enhancement instructions and to extract multimodal domain features from the input data to obtain multimodal initial features;

[0018] The identification unit is used to identify long-tail scene data in the multimodal initial features using a density peak clustering algorithm;

[0019] An enhancement unit is used to enhance the long-tail scene data using a generative adversarial network and the cross-domain knowledge graph to obtain supplementary features;

[0020] An addition unit is used to add the supplementary features to the multimodal initial features to obtain multimodal cross-domain enhanced features.

[0021] A computer device, the computer device comprising:

[0022] Memory, storing at least one instruction; and

[0023] The processor executes instructions stored in the memory to implement the cross-domain long-tail data augmentation method.

[0024] A computer-readable storage medium storing at least one instruction that is executed by a processor in a computer device to implement the cross-domain long-tail data augmentation method.

[0025] As can be seen from the above technical solutions, this invention can construct a target domain knowledge graph and a source domain knowledge graph, and perform cross-domain knowledge alignment between the target domain knowledge graph and the source domain knowledge graph to obtain a cross-domain knowledge graph, unifying knowledge from different domains into a shared semantic space, promoting the association and reuse of knowledge from different domains; it uses density peak clustering algorithm to identify long-tail scene data in multimodal initial features, and can accurately identify long-tail data through data distribution analysis; it uses generative adversarial networks and cross-domain knowledge graphs to enhance long-tail scene data, and adds supplementary features to multimodal initial features to obtain multimodal cross-domain enhanced features, supplementing the feature information of long-tail data and expanding the coverage of training data, thereby improving the processing capability for low-frequency, cross-domain and complex long-tail scenes. Attached Figure Description

[0026] Figure 1 This is a flowchart of a preferred embodiment of the cross-domain long-tail data augmentation method of the present invention.

[0027] Figure 2 This is a functional block diagram of a preferred embodiment of the cross-domain long-tail data enhancement device of the present invention.

[0028] Figure 3 This is a schematic diagram of the structure of a computer device that implements a preferred embodiment of the cross-domain long-tail data augmentation method of the present invention. Detailed Implementation

[0029] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0030] like Figure 1 The diagram shown is a flowchart of a preferred embodiment of the cross-domain long-tail data augmentation method of the present invention. The order of the steps in this flowchart can be changed, and some steps can be omitted, depending on different requirements.

[0031] The cross-domain long-tail data augmentation method is applied to one or more computer devices. The computer device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0032] The computer device can be any electronic product that can interact with the user, such as a personal computer, tablet computer, smartphone, personal digital assistant (PDA), game console, interactive network television (IPTV), smart wearable device, etc.

[0033] The computer equipment may also include network equipment and / or user equipment. The network equipment includes, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of hosts or network servers.

[0034] The server can be a standalone server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0035] Artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0036] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0037] The network in which the computer device is located includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, and virtual private network (VPN).

[0038] S10, in response to the instruction to enhance long-tail data in the target domain, construct a target domain knowledge graph of the target domain and a source domain knowledge graph of the source domain.

[0039] In this embodiment, the target domain can be the domain in which a task needs to be performed, and the source domain can be the domain in which knowledge transfer needs to be performed.

[0040] For example, when it is necessary to switch the robot in the financial services hall to provide services in the medical operating room, the target field is the medical and health field, and the source field is the financial field.

[0041] In this embodiment, long-tail data refers to a distribution phenomenon in which a few mainstream categories account for the majority of the data volume, while a large number of non-mainstream categories account for only a small amount of data. This phenomenon is common in fields such as the Internet, business, and machine learning.

[0042] For example, in the financial sector, long-tail data can include rare financial fraud methods, low-frequency abnormal transaction behaviors, and risk events under special macroeconomic conditions; in the healthcare sector, long-tail data can include rare disease cases, adverse reactions of special individuals, and combinations of complex conditions.

[0043] In this embodiment, the enhanced command can be triggered by personnel in the relevant field according to actual needs.

[0044] In this embodiment, constructing the target domain knowledge graph includes:

[0045] Collect structured and unstructured knowledge within the target domain;

[0046] Identify entities and relationships between entities in the structured knowledge and the unstructured knowledge;

[0047] The entities are identified as nodes, and the nodes are connected according to the relationships between the entities to form edges, thereby obtaining the target domain knowledge graph.

[0048] For example, in the field of industrial assembly, structured knowledge such as industrial assembly manuals and user manuals for household items can be collected, as well as unstructured knowledge such as technical documents and user reviews. Entities can include "screwdriver," "dining table," "tightening operation," and "placement rules," while edges can include "tool-operation" and "item-location."

[0049] Similarly, the source domain knowledge graph can be constructed in a similar manner to the construction of the target domain knowledge graph, which will not be elaborated here.

[0050] S11, perform cross-domain knowledge alignment between the target domain knowledge graph and the source domain knowledge graph to obtain a cross-domain knowledge graph.

[0051] In this embodiment, the step of performing cross-domain knowledge alignment between the target domain knowledge graph and the source domain knowledge graph to obtain a cross-domain knowledge graph includes:

[0052] Identify nodes and edges in the target domain knowledge graph and the source domain knowledge graph that are semantically similar;

[0053] Construct a cross-domain mapping matrix from the source domain knowledge graph to the target domain knowledge graph based on the identified nodes and edges;

[0054] The target domain knowledge graph is aligned with the source domain knowledge graph based on the cross-domain mapping matrix to obtain the cross-domain knowledge graph.

[0055] Specifically, graph convolution operations can be performed on the target domain knowledge graph and the source domain knowledge graph, and the semantic similarity of cross-domain node pairs can be calculated.

[0056] The cross-domain node pair may include a pair of semantically similar nodes and edges.

[0057] The semantic similarity of the cross-domain node pairs can be calculated using a Softmax function with cosine similarity or a temperature coefficient.

[0058] Specifically, the embedding mapping from the source domain to the target domain can be learned through the optimal transfer or linear transformation matrix, thereby forming the cross-domain mapping matrix. The target domain knowledge graph and the source domain knowledge graph are then aligned according to the cross-domain mapping matrix to ensure the consistency of the semantic space.

[0059] S12, obtain input data according to the enhancement instruction, and perform multimodal domain feature extraction on the input data to obtain multimodal initial features.

[0060] In this embodiment, the step of extracting multimodal domain features from the input data to obtain initial multimodal features includes:

[0061] A domain-adaptive ResNeXt (Residual Networks with Next) network is used to extract features from the video frames in the input data to obtain domain visual features.

[0062] The language text in the input data is encoded using a pre-trained multilingual BERT (Bidirectional Encoder Representation from Transformers) model to obtain a first encoded feature; the domain label of the target domain is obtained; a target word embedding matrix is ​​selected from a pre-constructed set of word embedding matrices based on the domain label; the first encoded feature and the target word embedding matrix are fused to obtain the domain language feature;

[0063] The motion sensor data in the input data is encoded using a Transformer-based time series encoder to obtain a second encoded feature; domain knowledge of the target domain is acquired; and the second encoded feature is optimized using the domain knowledge to obtain domain motion features.

[0064] The domain visual features, domain language features, and domain action features are integrated to obtain the multimodal initial features.

[0065] In this process, the domain-adaptive ResNeXt network is used to extract domain visual features, which can reduce feature differences between domains while preserving common visual features.

[0066] In particular, combining a pre-trained multilingual BERT model with a domain-specific word embedding matrix to extract the domain language features can enhance the expressive power of language features for semantics in different domains.

[0067] Specifically, by combining a Transformer-based time-series encoder with domain knowledge to extract domain action features, the action features can be enhanced using domain knowledge. For example, in the medical field, surgical actions can be associated with human anatomy; in the field of financial service robots, action feature representations can be optimized by incorporating mechanical kinematics knowledge.

[0068] The above embodiments can effectively extract cross-domain features such as vision, language, and action from different domains, reduce feature differences between domains, and enhance the expressive power of features for semantics in different domains, providing high-quality initial features for subsequent processing.

[0069] S13, use density peak clustering algorithm to identify long-tail scene data in the multimodal initial features.

[0070] In this embodiment, the step of using density peak clustering algorithm to identify long-tail scene data in the multimodal initial features includes:

[0071] The density peak clustering algorithm is used to cluster the multimodal initial features to obtain multiple clusters;

[0072] Calculate the sample size for each cluster, and calculate the global sample size for the initial multimodal features;

[0073] The distribution percentage of each cluster is calculated based on the sample size of each cluster and the global sample size.

[0074] Calculate the average local density of each cluster and the global average density of the initial multimodal features;

[0075] Obtain the pre-configured percentage threshold and density threshold;

[0076] When a cluster is detected whose distribution proportion is less than the proportion threshold and whose average local density is less than the density threshold, the features within the detected cluster are identified as the long-tail scene data.

[0077] The percentage threshold and the density threshold can be optimal values ​​selected based on a large number of experiments.

[0078] The above embodiments enable accurate identification of long-tail data.

[0079] S14, the long-tail scene data is enhanced using generative adversarial networks and the cross-domain knowledge graph to obtain supplementary features.

[0080] In this embodiment, the augmentation of the long-tail scene data using generative adversarial networks and the cross-domain knowledge graph to obtain supplementary features includes:

[0081] Generative adversarial networks are used to enhance the long-tail scene data to obtain similar samples;

[0082] The long-tail scenario data is used to query the cross-domain knowledge graph to obtain triple data;

[0083] The supplementary features are obtained by concatenating the triplet data with the similar samples.

[0084] The generator of the generative adversarial network can take the long-tailed scene data and random noise as input.

[0085] The generator of the generative adversarial network can also incorporate scenario prior vectors extracted from the cross-domain knowledge graph, such as the subgraph embedding of "equipment failure-maintenance operation" extracted from the industrial knowledge graph.

[0086] The discriminator of the generative adversarial network can adopt a multi-branch discriminant structure to judge the text consistency, visual realism, and cross-modal correlation of the generated samples, so as to constrain the multimodal consistency of the generated data.

[0087] Specifically, subgraph walking algorithms can be used to extract relevant triples from the cross-domain knowledge graph. For example, in a scenario of abnormal financial transactions, the triple could be (abnormal transaction, associated risk, abnormal behavior characteristics).

[0088] In this process, the triplet data is converted into graph embedding vectors, concatenated to the similar samples, and fused using a multilayer perceptron to obtain the supplementary features, thereby supplementing the knowledge background of long-tail data, such as abnormal transaction knowledge in financial abnormal transaction pattern scenarios.

[0089] Through the above embodiments, effective data augmentation can be performed after accurately identifying long-tail data, supplementing the feature information of long-tail data, expanding the coverage of training data, and thus improving the processing capability for low-frequency, complex long-tail scenarios.

[0090] S15, add the supplementary features to the multimodal initial features to obtain multimodal cross-domain enhanced features.

[0091] In this embodiment, after obtaining the multimodal cross-domain enhancement features, the method further includes:

[0092] In response to the processing instruction for the target task, an initial model corresponding to the target task is obtained;

[0093] The initial model is trained using the multimodal cross-domain enhancement features to obtain the target model;

[0094] The target task is performed using the target model.

[0095] For example, when the target task is to assist in the diagnosis of diseases in the medical field, the multimodal cross-domain enhanced features can be used to train the classification model, so that the model can refer to the diagnostic experience of different departments and different cases to make inferences, thereby assisting doctors to make more accurate diagnostic decisions.

[0096] For example, when the target task is the object handling task of a service robot in a financial hall, the multimodal cross-domain enhanced features can be used to train the decision model, so that the model can refer to the handling knowledge of handling robots in the industrial field to make inferences, thereby assisting in more effective control of the service robot in the financial hall.

[0097] Through the above embodiments, the model can be trained and processed in a targeted manner by combining multimodal feature enhancement with deep integration of cross-domain knowledge graphs, enabling the model to learn from knowledge in different domains to solve new problems and improve the model's reasoning ability and decision-making accuracy in cross-domain scenarios.

[0098] In this embodiment, during the execution of the target task using the target model, the execution results feedback of the model decision can also be collected, including the accuracy of action execution, the rationality of language response, and the evaluation of cross-domain knowledge transfer effect. The decision performance in long-tail scenarios can be recorded, and the shortcomings of the model in low-frequency scenarios can be analyzed.

[0099] Furthermore, the results can be fed back for bidirectional optimization. On the one hand, the model parameters are updated through the backpropagation algorithm, and on the other hand, the alignment relationship and node features of the cross-domain knowledge graph are optimized, thereby adjusting the long-tail data augmentation strategy and feature supplementation method.

[0100] For example, in financial scenarios, feedback results from the model on the assessment of financial transaction risks can be collected. If a high misjudgment rate is found for a certain type of new risk, the model parameters can be updated on the one hand, and the relationship between relevant risk nodes in the financial knowledge graph can be optimized on the other hand, and the enhancement method of long-tail risk data can be adjusted, thereby improving the accuracy of subsequent risk assessments.

[0101] For example, in a medical setting, if there are deficiencies in the diagnosis of certain rare cases based on the model's feedback on case diagnosis, the model's feature extraction and inference network parameters can be updated. At the same time, the alignment of relevant disease nodes in the medical knowledge graph can be optimized, and the feature supplementation of long-tail case data can be improved, thereby enhancing the diagnostic effect.

[0102] Through the above embodiments, a continuous optimization cycle of the model and knowledge graph is formed, which can continuously improve the model performance and knowledge graph quality based on feedback from actual applications, and enhance the model's generalization ability and adaptability to cross-domain and long-tail scenarios.

[0103] As can be seen from the above technical solutions, this invention can construct a target domain knowledge graph and a source domain knowledge graph, and perform cross-domain knowledge alignment between the target domain knowledge graph and the source domain knowledge graph to obtain a cross-domain knowledge graph, unifying knowledge from different domains into a shared semantic space, promoting the association and reuse of knowledge from different domains; it uses density peak clustering algorithm to identify long-tail scene data in multimodal initial features, and can accurately identify long-tail data through data distribution analysis; it uses generative adversarial networks and cross-domain knowledge graphs to enhance long-tail scene data, and adds supplementary features to multimodal initial features to obtain multimodal cross-domain enhanced features, supplementing the feature information of long-tail data and expanding the coverage of training data, thereby improving the processing capability for low-frequency, cross-domain and complex long-tail scenes.

[0104] like Figure 2 The diagram shown is a functional block diagram of a preferred embodiment of the cross-domain long-tail data augmentation device of the present invention. The cross-domain long-tail data augmentation device 11 includes a construction unit 110, an alignment unit 111, an extraction unit 112, an identification unit 113, an augmentation unit 114, and an addition unit 115. The module / unit referred to in this invention refers to a series of computer program segments that can be executed by a processor and perform a fixed function, and are stored in memory. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.

[0105] The construction unit 110 is configured to construct a target domain knowledge graph of the target domain and a source domain knowledge graph of the source domain in response to an enhancement instruction for long-tail data in the target domain.

[0106] In this embodiment, the target domain can be the domain in which a task needs to be performed, and the source domain can be the domain in which knowledge transfer needs to be performed.

[0107] For example, when it is necessary to switch the robot in the financial services hall to provide services in the medical operating room, the target field is the medical and health field, and the source field is the financial field.

[0108] In this embodiment, long-tail data refers to a distribution phenomenon in which a few mainstream categories account for the majority of the data volume, while a large number of non-mainstream categories account for only a small amount of data. This phenomenon is common in fields such as the Internet, business, and machine learning.

[0109] For example, in the financial sector, long-tail data can include rare financial fraud methods, low-frequency abnormal transaction behaviors, and risk events under special macroeconomic conditions; in the healthcare sector, long-tail data can include rare disease cases, adverse reactions of special individuals, and combinations of complex conditions.

[0110] In this embodiment, the enhanced command can be triggered by personnel in the relevant field according to actual needs.

[0111] In this embodiment, the construction unit 110 constructs the target domain knowledge graph of the target domain, including:

[0112] Collect structured and unstructured knowledge within the target domain;

[0113] Identify entities and relationships between entities in the structured knowledge and the unstructured knowledge;

[0114] The entities are identified as nodes, and the nodes are connected according to the relationships between the entities to form edges, thereby obtaining the target domain knowledge graph.

[0115] For example, in the field of industrial assembly, structured knowledge such as industrial assembly manuals and user manuals for household items can be collected, as well as unstructured knowledge such as technical documents and user reviews. Entities can include "screwdriver," "dining table," "tightening operation," and "placement rules," while edges can include "tool-operation" and "item-location."

[0116] Similarly, the source domain knowledge graph can be constructed in a similar manner to the construction of the target domain knowledge graph, which will not be elaborated here.

[0117] The alignment unit 111 is used to perform cross-domain knowledge alignment between the target domain knowledge graph and the source domain knowledge graph to obtain a cross-domain knowledge graph.

[0118] In this embodiment, the alignment unit 111 performs cross-domain knowledge alignment between the target domain knowledge graph and the source domain knowledge graph to obtain a cross-domain knowledge graph, including:

[0119] Identify nodes and edges in the target domain knowledge graph and the source domain knowledge graph that are semantically similar;

[0120] Construct a cross-domain mapping matrix from the source domain knowledge graph to the target domain knowledge graph based on the identified nodes and edges;

[0121] The target domain knowledge graph is aligned with the source domain knowledge graph based on the cross-domain mapping matrix to obtain the cross-domain knowledge graph.

[0122] Specifically, graph convolution operations can be performed on the target domain knowledge graph and the source domain knowledge graph, and the semantic similarity of cross-domain node pairs can be calculated.

[0123] The cross-domain node pair may include a pair of semantically similar nodes and edges.

[0124] The semantic similarity of the cross-domain node pairs can be calculated using a Softmax function with cosine similarity or a temperature coefficient.

[0125] Specifically, the embedding mapping from the source domain to the target domain can be learned through the optimal transfer or linear transformation matrix, thereby forming the cross-domain mapping matrix. The target domain knowledge graph and the source domain knowledge graph are then aligned according to the cross-domain mapping matrix to ensure the consistency of the semantic space.

[0126] The extraction unit 112 is used to acquire input data according to the enhancement instruction, and to extract multimodal domain features from the input data to obtain multimodal initial features.

[0127] In this embodiment, the extraction unit 112 performs multimodal domain feature extraction on the input data to obtain multimodal initial features, including:

[0128] A domain-adaptive ResNeXt (Residual Networks with Next) network is used to extract features from the video frames in the input data to obtain domain visual features.

[0129] The language text in the input data is encoded using a pre-trained multilingual BERT (Bidirectional Encoder Representation from Transformers) model to obtain a first encoded feature; the domain label of the target domain is obtained; a target word embedding matrix is ​​selected from a pre-constructed set of word embedding matrices based on the domain label; the first encoded feature and the target word embedding matrix are fused to obtain the domain language feature;

[0130] The motion sensor data in the input data is encoded using a Transformer-based time series encoder to obtain a second encoded feature; domain knowledge of the target domain is acquired; and the second encoded feature is optimized using the domain knowledge to obtain domain motion features.

[0131] The domain visual features, domain language features, and domain action features are integrated to obtain the multimodal initial features.

[0132] In this process, the domain-adaptive ResNeXt network is used to extract domain visual features, which can reduce feature differences between domains while preserving common visual features.

[0133] In particular, combining a pre-trained multilingual BERT model with a domain-specific word embedding matrix to extract the domain language features can enhance the expressive power of language features for semantics in different domains.

[0134] Specifically, by combining a Transformer-based time-series encoder with domain knowledge to extract domain action features, the action features can be enhanced using domain knowledge. For example, in the medical field, surgical actions can be associated with human anatomy; in the field of financial service robots, action feature representations can be optimized by incorporating mechanical kinematics knowledge.

[0135] The above embodiments can effectively extract cross-domain features such as vision, language, and action from different domains, reduce feature differences between domains, and enhance the expressive power of features for semantics in different domains, providing high-quality initial features for subsequent processing.

[0136] The identification unit 113 is used to identify long-tail scene data in the multimodal initial features using the density peak clustering algorithm.

[0137] In this embodiment, the identification unit 113 uses the density peak clustering algorithm to identify long-tail scene data in the multimodal initial features, including:

[0138] The density peak clustering algorithm is used to cluster the multimodal initial features to obtain multiple clusters;

[0139] Calculate the sample size for each cluster, and calculate the global sample size for the initial multimodal features;

[0140] The distribution percentage of each cluster is calculated based on the sample size of each cluster and the global sample size.

[0141] Calculate the average local density of each cluster and the global average density of the initial multimodal features;

[0142] Obtain the pre-configured percentage threshold and density threshold;

[0143] When a cluster is detected whose distribution proportion is less than the proportion threshold and whose average local density is less than the density threshold, the features within the detected cluster are identified as the long-tail scene data.

[0144] The percentage threshold and the density threshold can be optimal values ​​selected based on a large number of experiments.

[0145] The above embodiments enable accurate identification of long-tail data.

[0146] The enhancement unit 114 is used to enhance the long-tail scene data using generative adversarial networks and the cross-domain knowledge graph to obtain supplementary features.

[0147] In this embodiment, the enhancement unit 114 uses a generative adversarial network and the cross-domain knowledge graph to enhance the long-tail scene data, obtaining supplementary features including:

[0148] Generative adversarial networks are used to enhance the long-tail scene data to obtain similar samples;

[0149] The long-tail scenario data is used to query the cross-domain knowledge graph to obtain triple data;

[0150] The supplementary features are obtained by concatenating the triplet data with the similar samples.

[0151] The generator of the generative adversarial network can take the long-tailed scene data and random noise as input.

[0152] The generator of the generative adversarial network can also incorporate scenario prior vectors extracted from the cross-domain knowledge graph, such as the subgraph embedding of "equipment failure-maintenance operation" extracted from the industrial knowledge graph.

[0153] The discriminator of the generative adversarial network can adopt a multi-branch discriminant structure to judge the text consistency, visual realism, and cross-modal correlation of the generated samples, so as to constrain the multimodal consistency of the generated data.

[0154] Specifically, subgraph walking algorithms can be used to extract relevant triples from the cross-domain knowledge graph. For example, in a scenario of abnormal financial transactions, the triple could be (abnormal transaction, associated risk, abnormal behavior characteristics).

[0155] In this process, the triplet data is converted into graph embedding vectors, concatenated to the similar samples, and fused using a multilayer perceptron to obtain the supplementary features, thereby supplementing the knowledge background of long-tail data, such as abnormal transaction knowledge in financial abnormal transaction pattern scenarios.

[0156] Through the above embodiments, effective data augmentation can be performed after accurately identifying long-tail data, supplementing the feature information of long-tail data, expanding the coverage of training data, and thus improving the processing capability for low-frequency, complex long-tail scenarios.

[0157] The adding unit 115 is used to add the supplementary features to the multimodal initial features to obtain multimodal cross-domain enhanced features.

[0158] In this embodiment, after obtaining the multimodal cross-domain enhancement features, an initial model corresponding to the target task is obtained in response to the processing instructions for the target task.

[0159] The initial model is trained using the multimodal cross-domain enhancement features to obtain the target model;

[0160] The target task is performed using the target model.

[0161] For example, when the target task is to assist in the diagnosis of diseases in the medical field, the multimodal cross-domain enhanced features can be used to train the classification model, so that the model can refer to the diagnostic experience of different departments and different cases to make inferences, thereby assisting doctors to make more accurate diagnostic decisions.

[0162] For example, when the target task is the object handling task of a service robot in a financial hall, the multimodal cross-domain enhanced features can be used to train the decision model, so that the model can refer to the handling knowledge of handling robots in the industrial field to make inferences, thereby assisting in more effective control of the service robot in the financial hall.

[0163] Through the above embodiments, the model can be trained and processed in a targeted manner by combining multimodal feature enhancement with deep integration of cross-domain knowledge graphs, enabling the model to learn from knowledge in different domains to solve new problems and improve the model's reasoning ability and decision-making accuracy in cross-domain scenarios.

[0164] In this embodiment, during the execution of the target task using the target model, the execution results feedback of the model decision can also be collected, including the accuracy of action execution, the rationality of language response, and the evaluation of cross-domain knowledge transfer effect. The decision performance in long-tail scenarios can be recorded, and the shortcomings of the model in low-frequency scenarios can be analyzed.

[0165] Furthermore, the results can be fed back for bidirectional optimization. On the one hand, the model parameters are updated through the backpropagation algorithm, and on the other hand, the alignment relationship and node features of the cross-domain knowledge graph are optimized, thereby adjusting the long-tail data augmentation strategy and feature supplementation method.

[0166] For example, in financial scenarios, feedback results from the model on the assessment of financial transaction risks can be collected. If a high misjudgment rate is found for a certain type of new risk, the model parameters can be updated on the one hand, and the relationship between relevant risk nodes in the financial knowledge graph can be optimized on the other hand, and the enhancement method of long-tail risk data can be adjusted, thereby improving the accuracy of subsequent risk assessments.

[0167] For example, in a medical setting, if there are deficiencies in the diagnosis of certain rare cases based on the model's feedback on case diagnosis, the model's feature extraction and inference network parameters can be updated. At the same time, the alignment of relevant disease nodes in the medical knowledge graph can be optimized, and the feature supplementation of long-tail case data can be improved, thereby enhancing the diagnostic effect.

[0168] Through the above embodiments, a continuous optimization cycle of the model and knowledge graph is formed, which can continuously improve the model performance and knowledge graph quality based on feedback from actual applications, and enhance the model's generalization ability and adaptability to cross-domain and long-tail scenarios.

[0169] As can be seen from the above technical solutions, this invention can construct a target domain knowledge graph and a source domain knowledge graph, and perform cross-domain knowledge alignment between the target domain knowledge graph and the source domain knowledge graph to obtain a cross-domain knowledge graph, unifying knowledge from different domains into a shared semantic space, promoting the association and reuse of knowledge from different domains; it uses density peak clustering algorithm to identify long-tail scene data in multimodal initial features, and can accurately identify long-tail data through data distribution analysis; it uses generative adversarial networks and cross-domain knowledge graphs to enhance long-tail scene data, and adds supplementary features to multimodal initial features to obtain multimodal cross-domain enhanced features, supplementing the feature information of long-tail data and expanding the coverage of training data, thereby improving the processing capability for low-frequency, cross-domain and complex long-tail scenes.

[0170] like Figure 3 The diagram shown is a schematic representation of the structure of a computer device that implements a preferred embodiment of the cross-domain long-tail data augmentation method of the present invention.

[0171] The computer device 1 may include a memory 12, a processor 13, and a bus (the arrow in the figure represents the bus), and may also include a computer program stored in the memory 12 and executable on the processor 13, such as a cross-domain long-tail data augmentation program.

[0172] Those skilled in the art will understand that the schematic diagram is merely an example of computer device 1 and does not constitute a limitation on computer device 1. Computer device 1 can be either a bus topology or a star topology. Computer device 1 may also include more or fewer other hardware or software than shown in the diagram, or different component arrangements. For example, computer device 1 may also include input / output devices, network access devices, etc.

[0173] It should be noted that the computer device 1 described is merely an example. Other existing or future electronic products that are adaptable to this invention should also be included within the scope of protection of this invention and are incorporated herein by reference.

[0174] The memory 12 includes at least one type of readable storage medium, such as flash memory, portable hard drive, multimedia card, card-type memory (e.g., SD or DX memory), magnetic storage, magnetic disk, optical disk, etc. In some embodiments, the memory 12 can be an internal storage unit of the computer device 1, such as a portable hard drive of the computer device 1. In other embodiments, the memory 12 can be an external storage device of the computer device 1, such as a plug-in portable hard drive, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the computer device 1. Furthermore, the memory 12 can include both internal and external storage units of the computer device 1. The memory 12 can be used not only to store application software and various types of data installed on the computer device 1, such as the code of cross-domain long-tail data augmentation programs, but also to temporarily store data that has been output or will be output.

[0175] In some embodiments, processor 13 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits packaged with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. Processor 13 is the control unit of the computer device 1, connecting various components of the computer device 1 via various interfaces and lines. It performs various functions of the computer device 1 and processes data by running or executing programs or modules stored in the memory 12 (e.g., executing cross-domain long-tail data augmentation programs) and calling data stored in the memory 12.

[0176] The processor 13 executes the operating system of the computer device 1 and various installed applications. The processor 13 executes these applications to implement the steps described in the various cross-domain long-tail data augmentation method embodiments above, for example... Figure 1 The steps are shown.

[0177] For example, the computer program may be divided into one or more modules / units, which are stored in the memory 12 and executed by the processor 13 to complete the present invention. The one or more modules / units may be a series of computer-readable instruction segments capable of performing specific functions, which describe the execution process of the computer program in the computer device 1. For example, the computer program may be divided into a construction unit 110, an alignment unit 111, an extraction unit 112, an identification unit 113, an enhancement unit 114, and an addition unit 115.

[0178] The integrated unit implemented as a software functional module described above can be stored in a computer-readable storage medium. This software functional module, stored in a storage medium, includes several instructions to cause a computer device (which may be a personal computer, computer equipment, or network device, etc.) or processor to execute portions of the cross-domain long-tail data augmentation method described in the various embodiments of this invention.

[0179] If the modules / units integrated in the computer device 1 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware devices. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above.

[0180] The computer program includes computer program code, which may be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium may include any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory, etc.

[0181] Furthermore, the computer-readable storage medium may primarily include a stored program area and a stored data area, wherein the stored program area may store the operating system, an application program required for at least one function, etc.; and the stored data area may store data created based on the use of blockchain nodes, etc.

[0182] The blockchain referred to in this invention is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.

[0183] The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, in... Figure 3 The bus is represented by only one straight line, but this does not mean that there is only one bus or one type of bus. The bus is configured to enable communication between the memory 12 and at least one processor 13, etc.

[0184] Although not shown, the computer device 1 may also include a power supply (such as a battery) to power various components. Preferably, the power supply can be logically connected to the at least one processor 13 through a power management device, thereby enabling functions such as charging management, discharging management, and power consumption management. The power supply may also include one or more DC or AC power supplies, recharging devices, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The computer device 1 may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be described in detail here.

[0185] Furthermore, the computer device 1 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a Wi-Fi interface, a Bluetooth interface, etc.), which is typically used to establish a communication connection between the computer device 1 and other computer devices.

[0186] Optionally, the computer device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), and optionally, a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen, etc. The display may also be appropriately referred to as a screen or display unit, used to display information processed in the computer device 1 and to display a visual user interface.

[0187] It should be understood that the embodiments described are for illustrative purposes only and are not limited to this structure in the scope of the patent application.

[0188] It will be understood by those skilled in the art that Figure 3 The structure shown does not constitute a limitation on the computer device 1, and may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0189] Combination Figure 1 The memory 12 in the computer device 1 stores multiple instructions to implement a cross-domain long-tail data augmentation method, and the processor 13 can execute the multiple instructions to achieve:

[0190] In response to the instruction to enhance long-tail data within the target domain, a target domain knowledge graph of the target domain is constructed, and a source domain knowledge graph of the source domain is constructed.

[0191] Cross-domain knowledge alignment is performed between the target domain knowledge graph and the source domain knowledge graph to obtain a cross-domain knowledge graph;

[0192] The input data is obtained according to the enhancement instructions, and multimodal domain features are extracted from the input data to obtain multimodal initial features;

[0193] Density peak clustering algorithm is used to identify long-tailed scene data in the initial features of the multimodal dataset;

[0194] The long-tail scene data is enhanced using generative adversarial networks and the cross-domain knowledge graph to obtain supplementary features;

[0195] The supplementary features are added to the initial multimodal features to obtain multimodal cross-domain enhanced features.

[0196] Specifically, the processor 13's implementation method for the above instructions can be found in [reference needed]. Figure 1 The descriptions of the relevant steps in the corresponding embodiments are not repeated here.

[0197] It should be noted that all data involved in this case was legally obtained. Software tools or components not belonging to this company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.

[0198] In the several embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.

[0199] This invention can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This invention can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0200] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0201] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.

[0202] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0203] Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within the invention. No appended diagram markings in the claims should be construed as limiting the scope of the claims.

[0204] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices described in this invention can also be implemented by a single unit or device through software or hardware. Terms such as "first," "second," etc., are used to indicate names and do not indicate any specific order.

[0205] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A cross-domain long-tail data augmentation method, characterized in that, The cross-domain long-tail data augmentation method includes: In response to the instruction to enhance long-tail data within the target domain, a target domain knowledge graph of the target domain is constructed, and a source domain knowledge graph of the source domain is constructed. Cross-domain knowledge alignment is performed between the target domain knowledge graph and the source domain knowledge graph to obtain a cross-domain knowledge graph; The input data is acquired according to the enhancement instructions, and multimodal domain feature extraction is performed on the input data to obtain multimodal initial features, including: extracting features from video frames in the input data using a domain-adaptive ResNeXt network to obtain domain visual features; encoding the language text in the input data using a pre-trained multilingual BERT model to obtain first encoded features; acquiring domain labels for the target domain; selecting a target word embedding matrix from a pre-constructed set of word embedding matrices based on the domain labels; fusing the first encoded features with the target word embedding matrix to obtain domain language features; encoding motion sensor data in the input data using a Transformer-based time-series encoder to obtain second encoded features; acquiring domain knowledge of the target domain; optimizing the second encoded features using the domain knowledge to obtain domain action features; and integrating the domain visual features, the domain language features, and the domain action features to obtain the multimodal initial features. Density peak clustering algorithm is used to identify long-tailed scene data in the initial features of the multimodal dataset; The long-tail scene data is enhanced using generative adversarial networks and the cross-domain knowledge graph to obtain supplementary features, including: enhancing the long-tail scene data using generative adversarial networks to obtain similar samples; querying the long-tail scene data in the cross-domain knowledge graph to obtain triple data; and concatenating the triple data with the similar samples to obtain the supplementary features. The supplementary features are added to the initial multimodal features to obtain multimodal cross-domain enhanced features.

2. The cross-domain long-tail data augmentation method as described in claim 1, characterized in that, The construction of the target domain knowledge graph for the target domain includes: Collect structured and unstructured knowledge within the target domain; Identify entities and relationships between entities in the structured knowledge and the unstructured knowledge; The entities are identified as nodes, and the nodes are connected according to the relationships between the entities to form edges, thereby obtaining the target domain knowledge graph.

3. The cross-domain long-tail data augmentation method as described in claim 1, characterized in that, The step of performing cross-domain knowledge alignment between the target domain knowledge graph and the source domain knowledge graph to obtain a cross-domain knowledge graph includes: Identify nodes and edges in the target domain knowledge graph and the source domain knowledge graph that are semantically similar; Construct a cross-domain mapping matrix from the source domain knowledge graph to the target domain knowledge graph based on the identified nodes and edges; The target domain knowledge graph is aligned with the source domain knowledge graph based on the cross-domain mapping matrix to obtain the cross-domain knowledge graph.

4. The cross-domain long-tail data augmentation method as described in claim 1, characterized in that, The method of using density peak clustering algorithm to identify long-tailed scene data in the multimodal initial features includes: The density peak clustering algorithm is used to cluster the multimodal initial features to obtain multiple clusters; Calculate the sample size for each cluster, and calculate the global sample size for the initial multimodal features; The distribution percentage of each cluster is calculated based on the sample size of each cluster and the global sample size. Calculate the average local density of each cluster and the global average density of the initial multimodal features; Obtain the pre-configured percentage threshold and density threshold; When a cluster is detected whose distribution proportion is less than the proportion threshold and whose average local density is less than the density threshold, the features within the detected cluster are identified as the long-tail scene data.

5. The cross-domain long-tail data augmentation method as described in claim 1, characterized in that, After obtaining the multimodal cross-domain enhanced features, the method further includes: In response to the processing instruction for the target task, an initial model corresponding to the target task is obtained; The initial model is trained using the multimodal cross-domain enhancement features to obtain the target model; The target task is performed using the target model.

6. A cross-domain long-tail data augmentation device, characterized in that, The cross-domain long-tail data augmentation device includes: The construction unit is configured to construct a target domain knowledge graph of the target domain and a source domain knowledge graph of the source domain in response to an enhancement instruction for long-tail data in the target domain. An alignment unit is used to perform cross-domain knowledge alignment between the target domain knowledge graph and the source domain knowledge graph to obtain a cross-domain knowledge graph. An extraction unit is configured to acquire input data according to the enhancement instructions and perform multimodal domain feature extraction on the input data to obtain multimodal initial features, including: extracting features from video frames in the input data using a domain-adaptive ResNeXt network to obtain domain visual features; encoding language text in the input data using a pre-trained multilingual BERT model to obtain first encoded features; acquiring domain labels for the target domain; selecting a target word embedding matrix from a pre-constructed set of word embedding matrices based on the domain labels; fusing the first encoded features with the target word embedding matrix to obtain domain language features; encoding motion sensor data in the input data using a Transformer-based time-series encoder to obtain second encoded features; acquiring domain knowledge of the target domain; optimizing the second encoded features using the domain knowledge to obtain domain action features; and integrating the domain visual features, the domain language features, and the domain action features to obtain the multimodal initial features. The identification unit is used to identify long-tail scene data in the multimodal initial features using a density peak clustering algorithm; An enhancement unit is used to enhance the long-tail scene data using a generative adversarial network and the cross-domain knowledge graph to obtain supplementary features. This includes: enhancing the long-tail scene data using a generative adversarial network to obtain similar samples; querying the long-tail scene data in the cross-domain knowledge graph to obtain triple data; and concatenating the triple data with the similar samples to obtain the supplementary features. An addition unit is used to add the supplementary features to the multimodal initial features to obtain multimodal cross-domain enhanced features.

7. A computer device, characterized in that, The computer device includes: Memory, storing at least one instruction; and The processor executes instructions stored in the memory to implement the cross-domain long-tail data augmentation method as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction, which is executed by a processor in a computer device to implement the cross-domain long-tail data augmentation method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Knowledge graph representation learning method through combination of entity hierarchy category

    CN107423820A