Cross-domain long-tail data enhancement method and device, equipment and medium

By constructing and aligning knowledge graphs, and combining density peak clustering and generative adversarial networks, the problem of augmenting long-tail data across domains is solved, and the model's decision-making ability in new domains is improved.

CN120974183AActive Publication Date: 2025-11-18PING AN TECH (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511059531.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-11-18
Estimated Expiration
2045-07-29

AI Technical Summary

Technical Problem

Existing models struggle to effectively utilize long-tail data when applied across different domains, leading to misjudgments or inability to make decisions in new domains.

Method used

We construct knowledge graphs for target and source domains, generate cross-domain knowledge graphs through cross-domain knowledge alignment, and combine density peak clustering and generative adversarial networks to identify and enhance the features of long-tail scenario data.

Benefits of technology

It improves the model's processing capabilities in cross-domain scenarios, especially its adaptability and decision-making accuracy in low-frequency and complex long-tail scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120974183A_ABST
    Figure CN120974183A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, provides a cross-domain long-tail data enhancement method and device, equipment and a medium, can be applied to financial and medical health old-age service scenes, can construct a target domain knowledge graph and a source domain knowledge graph, and carries out cross-domain knowledge alignment on the target domain knowledge graph and the source domain knowledge graph, so that the user experience is improved. Different domain knowledge is unified to a shared semantic space, and association and reuse of the different domain knowledge are promoted; long-tail scene data in the multi-modal initial features are identified by using a density peak clustering algorithm, and the long-tail data can be accurately identified through data distribution analysis; according to the method, the generative adversarial network and the cross-domain knowledge graph are utilized to enhance the long-tail scene data, and the supplementary features are added to the multi-modal initial features to obtain the multi-modal cross-domain enhanced features, so that the feature information of the long-tail data is supplemented, and the coverage range of the training data is expanded; therefore, the processing capability of low-frequency, cross-domain and complex long-tail scenes is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and in particular to a cross-domain long-tail data enhancement method, device, equipment and medium. BACKGROUND

[0002] There are a large number of low-frequency and complex long-tail data in real-world multi-modal data, and traditional models are usually trained only for specific domain data.

[0003] Because the model is limited by the distribution of training data, it is difficult to reuse existing domain knowledge in new domains, so when encountering long-tail scenarios such as rare fault handling and special environment operation, misjudgment or inability to make decisions may occur.

[0004] For example, a robot used in a financial service hall may not be able to transfer the knowledge of carrying objects to the task of delivering surgical equipment when switching to a medical surgery scene due to the existence of long-tail data, which not only requires the model to be retrained, but also results in poor training results.

[0005] Therefore, how to enhance long-tail data by combining cross-domain knowledge to make it more suitable for the execution of VLA (Vision-Language-Action) tasks has become a problem to be solved. SUMMARY

[0006] In view of the above, it is necessary to provide a cross-domain long-tail data enhancement method, device, equipment and medium, which aims to solve the problem of being unable to enhance long-tail data across domains.

[0007] A cross-domain long-tail data enhancement method, the cross-domain long-tail data enhancement method comprising:

[0008] In response to an enhancement instruction for long-tail data in a target domain, a target domain knowledge graph of the target domain is constructed, and a source domain knowledge graph of a source domain is constructed;

[0009] The target domain knowledge graph and the source domain knowledge graph are aligned for cross-domain knowledge to obtain a cross-domain knowledge graph;

[0010] Input data is obtained according to the enhancement instruction, and multi-modal domain feature extraction is performed on the input data to obtain multi-modal initial features;

[0011] Density peak clustering algorithm is used to identify long-tail scenario data in the multi-modal initial features;

[0012] The long-tail scenario data is enhanced using a generative adversarial network and the cross-domain knowledge graph to obtain supplementary features;

[0013] add the supplementary feature to the multi-modal initial feature to obtain a multi-modal cross-domain enhanced feature.

[0014] A cross-domain long-tail data enhancement apparatus, comprising:

[0015] a construction unit configured to, in response to an enhancement instruction for long-tail data in a target domain, construct a target domain knowledge graph of the target domain and construct a source domain knowledge graph of a source domain;

[0016] an alignment unit configured to perform cross-domain knowledge alignment on the target domain knowledge graph and the source domain knowledge graph to obtain a cross-domain knowledge graph;

[0017] an extraction unit configured to acquire input data according to the enhancement instruction and perform multi-modal domain feature extraction on the input data to obtain multi-modal initial features;

[0018] a recognition unit configured to recognize long-tail scene data in the multi-modal initial features by using a density peak clustering algorithm;

[0019] an enhancement unit configured to enhance the long-tail scene data by using a generative adversarial network and the cross-domain knowledge graph to obtain supplementary features;

[0020] an adding unit configured to add the supplementary features to the multi-modal initial features to obtain multi-modal cross-domain enhanced features.

[0021] A computer device, comprising:

[0022] a memory configured to store at least one instruction; and

[0023] a processor configured to execute the instruction stored in the memory to implement the cross-domain long-tail data enhancement method.

[0024] A computer-readable storage medium, having stored therein at least one instruction, the at least one instruction being executed by a processor in a computer device to implement the cross-domain long-tail data enhancement method.

[0025] As can be seen from the above technical solutions, this invention can construct a target domain knowledge graph and a source domain knowledge graph, and perform cross-domain knowledge alignment between the target domain knowledge graph and the source domain knowledge graph to obtain a cross-domain knowledge graph, unifying knowledge from different domains into a shared semantic space, promoting the association and reuse of knowledge from different domains; it uses density peak clustering algorithm to identify long-tail scene data in multimodal initial features, and can accurately identify long-tail data through data distribution analysis; it uses generative adversarial networks and cross-domain knowledge graphs to enhance long-tail scene data, and adds supplementary features to multimodal initial features to obtain multimodal cross-domain enhanced features, supplementing the feature information of long-tail data and expanding the coverage of training data, thereby improving the processing capability for low-frequency, cross-domain and complex long-tail scenes. Attached Figure Description

[0026] Figure 1 This is a flowchart of a preferred embodiment of the cross-domain long-tail data augmentation method of the present invention.

[0027] Figure 2 This is a functional block diagram of a preferred embodiment of the cross-domain long-tail data enhancement device of the present invention.

[0028] Figure 3 This is a schematic diagram of the structure of a computer device that implements a preferred embodiment of the cross-domain long-tail data augmentation method of the present invention. Detailed Implementation

[0029] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0030] like Figure 1 The diagram shown is a flowchart of a preferred embodiment of the cross-domain long-tail data augmentation method of the present invention. The order of the steps in this flowchart can be changed, and some steps can be omitted, depending on different requirements.

[0031] The cross-domain long-tail data augmentation method is applied to one or more computer devices. The computer device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0032] The computer device can be any electronic product that can interact with a user, such as a personal computer, a tablet computer, a smartphone, a personal digital assistant (PDA), a game console, an interactive Internet Protocol Television (IPTV), a smart wearable device, and the like.

[0033] The computer device can also include a network device and / or a user device. The network device includes, but is not limited to, a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of hosts or network servers based on cloud computing.

[0034] The server can be a standalone server, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, and the like.

[0035] Artificial intelligence (AI) is the theory, method, technology and application system for using digital computers or computer-controlled machines to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0036] Artificial intelligence basic technologies generally include technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, and the like. Artificial intelligence software technologies mainly include computer vision technology, robot technology, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning, and the like.

[0037] The network in which the computer device is located includes, but is not limited to, the Internet, a wide area network, a metropolitan area network, a local area network, a virtual private network (VPN), and the like.

[0038] S10, in response to the enhancement instruction of the long-tail data in the target field, constructing a target field knowledge graph of the target field, and constructing a source field knowledge graph of the source field.

[0039] In this embodiment, the target field can be a field that needs to perform a task, and the source field can be a field that needs to perform knowledge transfer.

[0040] For example, when a robot in a financial service hall needs to be switched to a medical operating room for service, the target field is the medical health field, and the source field is the financial field.

[0041] In this embodiment, the long-tail data refers to a distribution phenomenon in which a small number of mainstream categories occupy most of the data volume, and a large number of non-mainstream categories only occupy a small amount of data, which is common in the fields of Internet, business, and machine learning.

[0042] For example, in the financial field, the long-tail data can be rare financial fraud methods, low-frequency abnormal transaction behaviors, risk events under special macroeconomic environments, etc.; in the medical health field, the long-tail data can be rare disease cases, adverse reactions of special individuals, combination conditions of complex conditions, etc.

[0043] In this embodiment, the enhancement instruction can be triggered by staff in the relevant field according to actual needs.

[0044] In this embodiment, the target field knowledge graph of the target field includes:

[0045] Collecting structured knowledge and unstructured knowledge in the target field;

[0046] Identifying entities and relationships between entities in the structured knowledge and the unstructured knowledge;

[0047] Determining the entities as nodes and connecting the nodes according to the relationships between the entities to form edges to obtain the target field knowledge graph.

[0048] For example, for the industrial assembly field, structured knowledge such as industrial assembly manuals and household item usage instructions, and unstructured knowledge such as technical documents and user reviews can be collected. Entities can include “screwdriver”, “dining table”, “tightening operation”, “placement rules”, etc., and edges can include “tool-operation”, “item-location”, etc.

[0049] Similarly, the source field knowledge graph can be constructed in a similar way to the construction of the target field knowledge graph, which is not described here.

[0050] S11, aligning the target field knowledge graph and the source field knowledge graph across fields to obtain a cross-field knowledge graph.

[0051] In this embodiment, the cross-field knowledge graph obtained by aligning the target field knowledge graph and the source field knowledge graph across fields includes:

[0052] Identifying nodes and edges with similar semantics in the target field knowledge graph and the source field knowledge graph;

[0053] constructing a cross-domain mapping matrix from the source domain knowledge graph to the target domain knowledge graph according to the identified nodes and edges;

[0054] aligning the target domain knowledge graph and the source domain knowledge graph according to the cross-domain mapping matrix to obtain the cross-domain knowledge graph.

[0055] Specifically, the target domain knowledge graph and the source domain knowledge graph can be subjected to graph convolution operation, and the semantic similarity of a cross-domain node pair can be calculated.

[0056] The cross-domain node pair can include a pair of nodes and edges with similar semantics.

[0057] The semantic similarity of the cross-domain node pair can be calculated by using a cosine similarity or a Softmax function with a temperature coefficient.

[0058] The embedding mapping from the source domain to the target domain can be learned by optimal transport or linear transformation matrix, so as to form the cross-domain mapping matrix, and the target domain knowledge graph and the source domain knowledge graph can be aligned according to the cross-domain mapping matrix, thereby ensuring the consistency of the semantic space.

[0059] S12, according to the enhancement instruction, input data is obtained, and multi-modal domain feature extraction is performed on the input data to obtain multi-modal initial features.

[0060] In this embodiment, the multi-modal domain feature extraction on the input data to obtain multi-modal initial features includes:

[0061] A ResNeXt (Residual Networks with Next) network with domain adaptation is used to extract features from video frames in the input data to obtain domain visual features.

[0062] A pre-trained multilingual BERT (Bidirectional Encoder Representation from Transformers) model is used to encode language text in the input data to obtain first encoding features; domain labels of the target domain are obtained; a target word embedding matrix is selected from a pre-constructed word embedding matrix set according to the domain labels; and the first encoding features and the target word embedding matrix are fused to obtain domain language features.

[0063] encoding the action sensor data in the input data using a Transformer-based time series encoder to obtain second encoded features; obtaining domain knowledge of the target domain; optimizing the second encoded features using the domain knowledge to obtain domain action features;

[0064] integrating the domain visual features, the domain language features, and the domain action features to obtain the multi-modal initial features.

[0065] The domain visual features are extracted by the domain-adaptive ResNeXt network, which can reduce the feature difference between domains while retaining general visual features.

[0066] The domain language features are extracted by combining the pre-trained multi-language BERT model and the domain-specific word embedding matrix, which can enhance the expression ability of language features for different domain semantics.

[0067] The domain action features are extracted by combining the Transformer-based time series encoder and the domain knowledge, which can enhance the action features by combining the domain knowledge. For example, in the medical field, surgical actions can be associated with human anatomy knowledge; in the financial service robot field, mechanical kinematics knowledge can be combined to optimize the action feature representation.

[0068] Through the above embodiments, cross-domain features such as vision, language, and action in different domains can be effectively extracted, the feature difference between domains can be reduced, and the expression ability of features for different domain semantics can be enhanced, providing high-quality initial features for subsequent processing.

[0069] S13, identifying long-tail scene data in the multi-modal initial features using a density peak clustering algorithm.

[0070] In this embodiment, the identification of the long-tail scene data in the multi-modal initial features using the density peak clustering algorithm includes:

[0071] Clustering the multi-modal initial features using the density peak clustering algorithm to obtain a plurality of clusters;

[0072] Calculating the sample size of each cluster and calculating the global sample size of the multi-modal initial features;

[0073] Calculating the distribution proportion of each cluster according to the sample size of each cluster and the global sample size;

[0074] Calculating the average local density of each cluster and the global average density of the multi-modal initial features;

[0075] Obtaining a proportion threshold and a density threshold configured in advance;

[0076] When it is detected that the distribution proportion of the cluster corresponds to is less than the proportion threshold value, and the corresponding average local density is less than the density threshold value, the features in the detected cluster are determined as the long-tail scene data.

[0077] The proportion threshold value and the density threshold value can be optimal values selected according to a large number of experiments.

[0078] Through the above embodiment, the long-tail data can be accurately identified.

[0079] S14, using a generative adversarial network and the cross-domain knowledge graph to enhance the long-tail scene data to obtain supplementary features.

[0080] In the embodiment, the using a generative adversarial network and the cross-domain knowledge graph to enhance the long-tail scene data to obtain supplementary features includes:

[0081] Using a generative adversarial network to enhance the long-tail scene data to obtain similar samples;

[0082] Using the long-tail scene data to query in the cross-domain knowledge graph to obtain triple data;

[0083] Splicing the triple data and the similar samples to obtain the supplementary features.

[0084] The generator of the generative adversarial network can take the long-tail scene data and random noise as input.

[0085] The generator of the generative adversarial network can also incorporate a scene prior vector extracted from the cross-domain knowledge graph, such as a subgraph embedding of "equipment failure-repair operation" extracted from an industrial knowledge graph.

[0086] The discriminator of the generative adversarial network can adopt a multi-branch discriminant structure to respectively discriminate the text consistency, visual authenticity and cross-modal correlation of the generated samples, so as to constrain the multi-modal consistency of the generated data.

[0087] The relevant triplets can be extracted from the cross-domain knowledge graph by subgraph walk algorithm, etc. For example, in the financial abnormal transaction mode scene, the triplets can be (abnormal transaction, associated risk, abnormal behavior feature).

[0088] After converting the triple data into graph embedding vectors, splicing to the similar samples, and fusing through a multi-layer perception machine, the supplementary features can be obtained, thereby supplementing the knowledge background of the long-tail data, such as the abnormal transaction knowledge in the financial abnormal transaction mode scene.

[0089] Through the above embodiment, effective data enhancement can be performed after accurate identification of long-tail data, the feature information of the long-tail data is supplemented, the training data coverage range is expanded, and the processing capability for low-frequency and complex long-tail scenes is improved.

[0090] S15, adding the supplemented features to the multi-modal initial features to obtain multi-modal cross-domain enhanced features.

[0091] In this embodiment, after the multi-modal cross-domain enhanced features are obtained, the method further comprises:

[0092] In response to a processing instruction of a target task, an initial model corresponding to the target task is acquired;

[0093] The initial model is trained by using the multi-modal cross-domain enhanced features to obtain a target model;

[0094] The target task is executed by using the target model.

[0095] For example, when the target task is auxiliary diagnosis of a disease in the medical field, a classification model can be trained by using the multi-modal cross-domain enhanced features, so that the model can refer to the diagnosis experience of different departments and different cases for reasoning, thereby assisting doctors to make more accurate diagnosis decisions.

[0096] For another example, when the target task is an object carrying task of a financial hall service robot, a decision model can be trained by using the multi-modal cross-domain enhanced features, so that the model can refer to the carrying knowledge of an industrial field carrying robot for reasoning, thereby assisting the financial hall service robot to be more effectively controlled.

[0097] Through the above embodiment, the model can be trained and processed in combination with the deep fusion of multi-modal feature enhancement and cross-domain knowledge graph, so that the model can solve new problems by referring to knowledge in different fields, and the reasoning ability and decision accuracy of the model in the cross-domain scene are improved.

[0098] In this embodiment, during execution of the target task by using the target model, feedback of an execution result of model decision can also be collected, including action execution accuracy, language reply rationality, cross-domain knowledge transfer effect evaluation, etc., decision performance in a long-tail scene is recorded, and deficiencies of the model in a low-frequency scene are analyzed.

[0099] Further, the result feedback can be used for bidirectional optimization, on one hand, model parameters are updated by using a back propagation algorithm, and on the other hand, alignment relationships and node features of the cross-domain knowledge graph are optimized, so that a long-tail data enhancement strategy and a feature supplement method are adjusted.

[0100] For example, in financial scenarios, feedback results from the model on the assessment of financial transaction risks can be collected. If a high misjudgment rate is found for a certain type of new risk, the model parameters can be updated on the one hand, and the relationship between relevant risk nodes in the financial knowledge graph can be optimized on the other hand, and the enhancement method of long-tail risk data can be adjusted, thereby improving the accuracy of subsequent risk assessments.

[0101] For example, in a medical setting, if there are deficiencies in the diagnosis of certain rare cases based on the model's feedback on case diagnosis, the model's feature extraction and inference network parameters can be updated. At the same time, the alignment of relevant disease nodes in the medical knowledge graph can be optimized, and the feature supplementation of long-tail case data can be improved, thereby enhancing the diagnostic effect.

[0102] Through the above embodiments, a continuous optimization cycle of the model and knowledge graph is formed, which can continuously improve the model performance and knowledge graph quality based on feedback from actual applications, and enhance the model's generalization ability and adaptability to cross-domain and long-tail scenarios.

[0103] As can be seen from the above technical solutions, this invention can construct a target domain knowledge graph and a source domain knowledge graph, and perform cross-domain knowledge alignment between the target domain knowledge graph and the source domain knowledge graph to obtain a cross-domain knowledge graph, unifying knowledge from different domains into a shared semantic space, promoting the association and reuse of knowledge from different domains; it uses density peak clustering algorithm to identify long-tail scene data in multimodal initial features, and can accurately identify long-tail data through data distribution analysis; it uses generative adversarial networks and cross-domain knowledge graphs to enhance long-tail scene data, and adds supplementary features to multimodal initial features to obtain multimodal cross-domain enhanced features, supplementing the feature information of long-tail data and expanding the coverage of training data, thereby improving the processing capability for low-frequency, cross-domain and complex long-tail scenes.

[0104] like Figure 2 The diagram shown is a functional block diagram of a preferred embodiment of the cross-domain long-tail data augmentation device of the present invention. The cross-domain long-tail data augmentation device 11 includes a construction unit 110, an alignment unit 111, an extraction unit 112, an identification unit 113, an augmentation unit 114, and an addition unit 115. The module / unit referred to in this invention refers to a series of computer program segments that can be executed by a processor and perform a fixed function, and are stored in memory. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.

[0105] The construction unit 110 is configured to construct a target domain knowledge graph of the target domain and a source domain knowledge graph of the source domain in response to an enhancement instruction for long-tail data in the target domain.

[0106] In this embodiment, the target domain can be a domain in which a task needs to be performed, and the source domain can be a domain in which knowledge needs to be transferred.

[0107] For example, when a robot in a financial service hall needs to be switched to a medical operating room for service, the target domain is the medical health domain, and the source domain is the financial domain.

[0108] In this embodiment, the long-tail data refers to a distribution phenomenon in which a small number of mainstream categories occupy most of the data volume, and a large number of non-mainstream categories only occupy a small amount of data in a data set, which is common in the fields of the Internet, business, and machine learning.

[0109] For example, in the financial field, the long-tail data can be rare financial fraud methods, low-frequency abnormal transaction behaviors, risk events under special macroeconomic environments, etc.; in the medical health field, the long-tail data can be rare disease cases, adverse reactions of special individuals, combination conditions of complex conditions, etc.

[0110] In this embodiment, the enhancement instruction can be triggered by a staff member in the relevant domain according to actual needs.

[0111] In this embodiment, the constructing unit 110 constructs the target domain knowledge graph of the target domain includes:

[0112] Collecting structured knowledge and unstructured knowledge in the target domain;

[0113] Identifying entities and relationships between entities in the structured knowledge and the unstructured knowledge;

[0114] Determining the entities as nodes, and connecting the nodes according to the relationships between the entities to form edges to obtain the target domain knowledge graph.

[0115] For example, for the industrial assembly domain, structured knowledge such as industrial assembly manuals and household item usage instructions, and unstructured knowledge such as technical documents and user reviews can be collected. Entities can include “screwdriver”, “dining table”, “tightening operation”, “placement rules”, etc., and edges can include “tool-operation”, “item-position”, etc.

[0116] Similarly, the source domain knowledge graph can be constructed in a similar manner to the construction of the target domain knowledge graph, which is not described here.

[0117] The aligning unit 111 is configured to perform cross-domain knowledge alignment on the target domain knowledge graph and the source domain knowledge graph to obtain a cross-domain knowledge graph.

[0118] In the embodiment, the alignment unit 111 performs cross-domain knowledge alignment on the target domain knowledge graph and the source domain knowledge graph, and obtains a cross-domain knowledge graph including:

[0119] Identify nodes and edges with similar semantics in the target domain knowledge graph and the source domain knowledge graph;

[0120] According to the identified nodes and edges, a cross-domain mapping matrix from the source domain knowledge graph to the target domain knowledge graph is constructed;

[0121] According to the cross-domain mapping matrix, the target domain knowledge graph and the source domain knowledge graph are aligned, and the cross-domain knowledge graph is obtained.

[0122] Specifically, the target domain knowledge graph and the source domain knowledge graph can be subjected to graph convolution operation, and the semantic similarity of cross-domain node pairs can be calculated.

[0123] Among them, the cross-domain node pair can include a pair of nodes and edges with similar semantics.

[0124] Among them, the semantic similarity of the cross-domain node pair can be calculated by using a cosine similarity or a Softmax function with a temperature coefficient.

[0125] Among them, the optimal transport or linear transformation matrix can be used to learn the embedding mapping from the source domain to the target domain, so as to form the cross-domain mapping matrix, and the target domain knowledge graph and the source domain knowledge graph are aligned according to the cross-domain mapping matrix, so as to ensure the consistency of the semantic space.

[0126] The extraction unit 112 is configured to obtain input data according to the enhancement instruction, and perform multi-modal domain feature extraction on the input data to obtain multi-modal initial features.

[0127] In the embodiment, the extraction unit 112 performs multi-modal domain feature extraction on the input data to obtain multi-modal initial features, including:

[0128] A ResNeXt (Residual Networks with Next, extended residual network) network with domain adaptation is used to extract features from video frames in the input data to obtain domain visual features;

[0129] The language text in the input data is encoded by using a pre-trained multilingual BERT (Bidirectional Encoder Representation from Transformers) model to obtain first encoding features; a domain label of the target domain is obtained; a target word embedding matrix is selected from a pre-constructed word embedding matrix set according to the domain label; the first encoding features and the target word embedding matrix are fused to obtain domain language features;

[0130] The action sensor data in the input data is encoded by using a Transformer-based time series encoder to obtain second encoding features; domain knowledge of the target domain is obtained; the second encoding features are optimized by using the domain knowledge to obtain domain action features;

[0131] The domain visual features, the domain language features and the domain action features are integrated to obtain the multi-modal initial features.

[0132] The domain visual features are extracted by the domain adaptive ResNeXt network, which can reduce the feature difference between domains while retaining general visual features.

[0133] The domain language features are extracted by combining the pre-trained multilingual BERT model and the domain-specific word embedding matrix, which can enhance the expression ability of language features for different domain semantics.

[0134] The domain action features are extracted by combining the Transformer-based time series encoder and the domain knowledge, which can enhance the action features by combining the domain knowledge. For example, in the medical field, surgical actions can be associated with human anatomy knowledge; in the financial service robot field, mechanical kinematics knowledge can be combined to optimize the action feature representation.

[0135] Through the above embodiments, cross-domain features such as vision, language and action in different domains can be effectively extracted, the feature difference between domains can be reduced, and the expression ability of the features for different domain semantics can be enhanced, thereby providing high-quality initial features for subsequent processing.

[0136] The recognition unit 113 is configured to recognize long-tail scene data in the multi-modal initial features by using a density peak clustering algorithm.

[0137] In this embodiment, the recognition unit 113 recognizes long-tail scene data in the multi-modal initial features by using a density peak clustering algorithm, which includes:

[0138] The multi-modal initial features are clustered by using the density peak clustering algorithm to obtain a plurality of clusters.

[0139] calculate a sample quantity of each cluster, and calculate a global sample quantity of the multi-modal initial features;

[0140] calculate a distribution proportion of each cluster according to the sample quantity of each cluster and the global sample quantity;

[0141] calculate an average local density of each cluster and a global average density of the multi-modal initial features;

[0142] obtain a proportion threshold and a density threshold pre-configured;

[0143] when it is detected that the distribution proportion corresponding to a cluster is less than the proportion threshold, and the average local density corresponding to the cluster is less than the density threshold, determine the features in the detected cluster as the long-tail scene data.

[0144] The proportion threshold and the density threshold can be optimal values selected according to a large number of experiments.

[0145] Through the above embodiment, the long-tail data can be accurately identified.

[0146] The enhancement unit 114 is configured to enhance the long-tail scene data by using a generative adversarial network and the cross-domain knowledge graph to obtain supplementary features.

[0147] In this embodiment, the enhancement unit 114 enhances the long-tail scene data by using a generative adversarial network and the cross-domain knowledge graph to obtain supplementary features, including:

[0148] enhancing the long-tail scene data by using a generative adversarial network to obtain similar samples;

[0149] querying the long-tail scene data in the cross-domain knowledge graph to obtain triple data;

[0150] splicing the triple data and the similar samples to obtain the supplementary features.

[0151] The generator of the generative adversarial network can take the long-tail scene data and random noise as input.

[0152] The generator of the generative adversarial network can also incorporate a scene prior vector extracted from the cross-domain knowledge graph, such as a subgraph embedding of “equipment failure-repair operation” extracted from an industrial knowledge graph.

[0153] The discriminator of the generative adversarial network can adopt a multi-branch discrimination structure to respectively discriminate the text consistency, visual authenticity and cross-modal correlation of the generated samples, so as to constrain the multi-modal consistency of the generated data.

[0154] Among them, relevant triples can be extracted from the cross-domain knowledge graph by subgraph walk algorithm and the like. For example, in the financial abnormal transaction mode scenario, the triples can be (abnormal transaction, associated risk, abnormal behavior characteristics).

[0155] Among them, after the triple data is converted into a graph embedding vector and spliced to the similar sample, the supplementary feature can be obtained by fusing through a multi-layer perception machine, thereby supplementing the knowledge background of the long-tail data, such as abnormal transaction knowledge in the financial abnormal transaction mode scenario.

[0156] Through the above embodiments, effective data enhancement can be performed after accurately identifying the long-tail data, the feature information of the long-tail data is supplemented, and the training data coverage range is expanded, thereby improving the processing capability for low-frequency and complex long-tail scenarios.

[0157] The adding unit 115 is configured to add the supplementary feature to the multi-modal initial feature to obtain a multi-modal cross-domain enhanced feature.

[0158] In this embodiment, after the multi-modal cross-domain enhanced feature is obtained, an initial model corresponding to a target task is acquired in response to a processing instruction of the target task.

[0159] The initial model is trained by using the multi-modal cross-domain enhanced feature to obtain a target model.

[0160] The target task is executed by using the target model.

[0161] For example, when the target task is disease auxiliary diagnosis in the medical field, a classification model can be trained by using the multi-modal cross-domain enhanced feature, so that the model can refer to the diagnosis experience of different departments and different cases for reasoning, thereby assisting doctors to make more accurate diagnosis decisions.

[0162] For another example, when the target task is the object carrying task of a financial hall service robot, a decision model can be trained by using the multi-modal cross-domain enhanced feature, so that the model can refer to the carrying knowledge of the carrying robot in the industrial field for reasoning, thereby assisting the more effective control of the financial hall service robot.

[0163] Through the above embodiments, the model can be trained and processed in a targeted manner by combining multi-modal feature enhancement and deep fusion of cross-domain knowledge graphs, so that the model can learn from different domain knowledge to solve new problems, thereby improving the reasoning ability and decision accuracy of the model in cross-domain scenarios.

[0164] In the embodiment, in the process of performing the target task by using the target model, the execution result feedback of the model decision can also be collected, including action execution accuracy, language reply rationality, cross-domain knowledge transfer effect evaluation, etc., the decision performance in the long-tail scene is recorded, and the deficiency of the model in the low-frequency scene is analyzed.

[0165] Further, the result feedback can be used for bidirectional optimization, on the one hand, the model parameters are updated by using the back propagation algorithm, and on the other hand, the alignment relationship and node feature of the cross-domain knowledge graph are optimized, so that the long-tail data enhancement strategy and the feature supplement method are adjusted.

[0166] For example, in the financial scene, the feedback result of the model to the financial transaction risk assessment can be collected, if it is found that the misjudgment rate of a certain type of new risk is high, on the one hand, the model parameters are updated, and on the other hand, the relationship of the related risk nodes in the financial knowledge graph is optimized, and the long-tail risk data enhancement method is adjusted, so as to improve the subsequent risk assessment accuracy.

[0167] For another example, in the medical scene, according to the feedback of the model to the case diagnosis, if there is a deficiency in the diagnosis of some rare cases, the feature extraction and reasoning network parameters of the model can be updated, and the alignment of the related disease nodes in the medical knowledge graph is optimized, and the feature supplement of the long-tail case data is improved, so as to improve the diagnosis effect.

[0168] Through the above embodiment, a continuous optimization cycle of the model and the knowledge graph is formed, the model performance and the knowledge graph quality can be continuously improved according to the actual application feedback, and the generalization ability and the adaptability to the cross-domain and long-tail scene of the model are enhanced.

[0169] As can be seen from the above technical solutions, the present application can construct a target domain knowledge graph and a source domain knowledge graph, and perform cross-domain knowledge alignment on the target domain knowledge graph and the source domain knowledge graph to obtain a cross-domain knowledge graph, unify different domain knowledge to a shared semantic space, and promote the association and reuse of different domain knowledge; the density peak clustering algorithm is used to identify long-tail scene data in the multi-modal initial feature, which can accurately identify long-tail data through data distribution analysis; the long-tail scene data is enhanced by using the generative adversarial network and the cross-domain knowledge graph, and the supplementary features are added to the multi-modal initial feature to obtain the multi-modal cross-domain enhanced feature, the feature information of the long-tail data is supplemented, and the training data coverage is expanded, so as to improve the processing capability of low-frequency, cross-domain and complex long-tail scenes.

[0170] As Figure 3 shown is a structure schematic diagram of a computer device of a preferred embodiment of a cross-domain long-tail data enhancement method of the present application.

[0171] The computer device 1 can include a memory 12, a processor 13 and a bus (the arrow in the figure is the bus), and can further include a computer program, such as a cross-domain long tail data enhancement program, stored in the memory 12 and executable on the processor 13.

[0172] Those skilled in the art can understand that the schematic diagram is only an example of the computer device 1 and does not constitute a limitation on the computer device 1, which can be a bus type structure or a star type structure, and can further include more or less other hardware or software or different component arrangement, such as an input / output device, a network access device, etc.

[0173] It should be noted that the computer device 1 is only an example, and other existing or future electronic products, such as those adaptable to the present application, should also be included in the protection scope of the present application and included by reference.

[0174] The memory 12 includes at least one type of readable storage medium, such as a flash memory, a mobile hard disk, a multimedia card, a card type memory (such as an SD or DX memory, etc.), a magnetic memory, a magnetic disk, an optical disk, etc. The memory 12 can be an internal storage unit of the computer device 1 in some embodiments, such as a mobile hard disk of the computer device 1. The memory 12 can also be an external storage device of the computer device 1 in other embodiments, such as a plug-in mobile hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 12 can include both an internal storage unit and an external storage device of the computer device 1. The memory 12 can be used not only to store application software and various data installed in the computer device 1, such as the code of the cross-domain long tail data enhancement program, but also to temporarily store data that has been output or will be output.

[0175] The processor 13 can be composed of integrated circuits in some embodiments, for example, can be composed of a single packaged integrated circuit, or can be composed of multiple packaged integrated circuits with the same function or different functions, including one or more central processing units (CPU), microprocessors, digital processing chips, graphics processors, combinations of various control chips, etc. The processor 13 is the control core (Control Unit) of the computer device 1, which connects various components of the entire computer device 1 through various interfaces and lines, and performs various functions and processes data of the computer device 1 by running or executing programs or modules stored in the memory 12 (such as executing cross-domain long tail data enhancement programs, etc.), and calling data stored in the memory 12.

[0176] The processor 13 executes the operating system of the computer device 1 and various installed application programs. The processor 13 executes the application programs to implement the steps in each of the above cross-domain long tail data enhancement method embodiments, for example Figure 1 The steps shown.

[0177] For example, the computer program can be divided into one or more modules / units, which are stored in the memory 12 and executed by the processor 13 to complete the present application. The one or more modules / units can be a series of computer readable instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program in the computer device 1. For example, the computer program can be divided into a construction unit 110, an alignment unit 111, an extraction unit 112, an identification unit 113, an enhancement unit 114, and an adding unit 115.

[0178] The integrated units implemented in the form of software function modules described above can be stored in a computer readable storage medium. The software function modules described above are stored in a storage medium, including a number of instructions for causing a computer device (which can be a personal computer, a computer device, or a network device, etc.) or a processor to execute part of the cross-domain long tail data enhancement method described in each embodiment of the present application.

[0179] The modules / units integrated by the computer device 1, if implemented in the form of software function units and sold or used as independent products, can be stored in a computer readable storage medium. Based on this understanding, all or part of the processes in the above-mentioned embodiments can also be instructed by a computer program to complete related hardware devices, and the computer program can be stored in a computer readable storage medium. When the processor executes the computer program, the steps of each method embodiment described above can be implemented.

[0180] The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory, etc.

[0181] Further, the computer readable storage medium can mainly include a storage program area and a storage data area, wherein the storage program area can store an operating system, an application required by at least one function, etc.; and the storage data area can store data created according to the use of the blockchain node, etc.

[0182] The blockchain referred to in the present application is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. The blockchain is essentially a decentralized database, which is a series of data blocks associated by using cryptographic methods, each data block contains the information of a batch of network transactions, and is used to verify the validity (anti-fake) of the information and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer, etc.

[0183] The bus can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one straight line is used in the Figure 3 However, it does not mean that there is only one bus or only one type of bus. The bus is arranged to realize the connection and communication between the memory 12, the at least one processor 13, etc.

[0184] Although not shown, the computer device 1 can also include a power supply (such as a battery) for powering each component. Preferably, the power supply can be logically connected to the at least one processor 13 through a power management device, so as to realize functions such as charge management, discharge management, and power consumption management through the power management device. The power supply can also include one or more direct current or alternating current power supplies, recharging devices, power failure detection circuits, power converters or inverters, power status indicators, etc. The computer device 1 can also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be described here.

[0185] Further, the computer device 1 can further include a network interface, which can optionally include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), and is generally used to establish a communication connection between the computer device 1 and other computer devices.

[0186] Optionally, the computer device 1 can further include a user interface, which can be a display (Display), an input unit (such as a keyboard (Keyboard)), and optionally, the user interface can also be a standard wired interface, a wireless interface. Optionally, in some embodiments, the display can be an LED display, a liquid crystal display, a touch liquid crystal display, an OLED (Organic Light-Emitting Diode) touch, etc. The display can also be appropriately referred to as a display screen or a display unit, and is used to display information processed in the computer device 1 and to display a visualized user interface.

[0187] It should be understood that the embodiments are only for illustration and are not limited in the scope of the patent application by the structure.

[0188] Those skilled in the art can understand that, Figure 3 The structure shown does not constitute a limitation on the computer device 1, and can include fewer or more components than shown, or combine certain components, or different component arrangements.

[0189] In combination Figure 1 The memory 12 in the computer device 1 stores a plurality of instructions to implement a cross-domain long-tail data enhancement method, and the processor 13 can execute the plurality of instructions to implement:

[0190] In response to an enhancement instruction for long-tail data in a target domain, a target domain knowledge graph of the target domain is constructed, and a source domain knowledge graph of a source domain is constructed;

[0191] Cross-domain knowledge alignment is performed on the target domain knowledge graph and the source domain knowledge graph to obtain a cross-domain knowledge graph;

[0192] Input data is obtained according to the enhancement instruction, and multi-modal domain feature extraction is performed on the input data to obtain multi-modal initial features;

[0193] A density peak clustering algorithm is used to identify long-tail scene data in the multi-modal initial features;

[0194] A generative adversarial network and the cross-domain knowledge graph are used to enhance the long-tail scene data to obtain supplementary features;

[0195] The supplementary features are added to the multi-modal initial features to obtain multi-modal cross-domain enhanced features.

[0196] Specifically, the specific implementation method of the processor 13 to the above instructions can refer to Figure 1 The description of related steps in the corresponding embodiments will not be repeated here.

[0197] It should be noted that the data involved in the present case are all legally obtained. The non-company software tools or components appearing in the embodiments of the present application are only examples for introduction and do not represent actual use.

[0198] In several embodiments provided by the present application, it should be understood that the disclosed system, device and method can be implemented by other manners. For example, the device embodiments described above are only schematic, for example, the division of the modules is only a logical function division, and actual implementation can have another division manner.

[0199] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment, in which tasks are performed by remote processing devices connected by a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.

[0200] The modules described as separate components can or can not be physically separated, and the components shown as modules can or can not be physical units, i.e. they can be located in one place or distributed to multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the present embodiment.

[0201] In addition, each functional module in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically independently, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of hardware plus software functional modules.

[0202] It will be obvious to a person skilled in the art that the application is not limited to the details of the above-described exemplary embodiments but can be implemented in other embodiments without departing from the scope of the application.

[0203] The embodiments should therefore be considered in all respects as illustrative and not restrictive, the scope of the application being indicated by the appended claims rather than by the description given above, so that all changes coming within the meaning and equivalency range of the claims are intended to be embraced therein. Any reference signs in the claims should not be construed as limiting the scope of the claims.

[0204] Furthermore, the word "comprising" does not exclude other elements or steps, and the singular does not exclude the plural and vice versa. Use of the expression "one" or "the" in relation to an element or step of the application does not exclude the presence of more than one element or step. The data, steps and / or functions can be carried out in any other order than the one described above without departing from the scope of the application.

[0205] Finally, it should be noted that the above-described embodiments are merely intended to illustrate the technical solutions of the present application, rather than limit the scope of the present application. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or equivalent replaced without departing from the spirit and scope of the present application.

Claims

1. A cross-domain long-tail data augmentation method, characterized in that, The cross-domain long-tail data enhancement method comprises: In response to an enhancement instruction for long-tail data in a target domain, a target domain knowledge graph of the target domain is constructed, and a source domain knowledge graph of a source domain is constructed; Cross-domain knowledge alignment is performed on the target domain knowledge graph and the source domain knowledge graph to obtain a cross-domain knowledge graph; Input data is obtained according to the enhancement instruction, and multi-modal domain feature extraction is performed on the input data to obtain multi-modal initial features; A density peak clustering algorithm is used to identify long-tail scene data in the multi-modal initial features; A generative adversarial network and the cross-domain knowledge graph are used to enhance the long-tail scene data to obtain supplementary features; The supplementary features are added to the multi-modal initial features to obtain multi-modal cross-domain enhanced features.

2. The cross-domain long tail data augmentation method of claim 1, wherein, The construction of the target domain knowledge graph of the target domain comprises: Structured knowledge and unstructured knowledge in the target domain are collected; Entities and inter-entity relationships in the structured knowledge and the unstructured knowledge are identified; The entities are determined as nodes, and the nodes are connected according to the inter-entity relationships to form edges to obtain the target domain knowledge graph.

3. The cross-domain long tail data augmentation method of claim 1, wherein, The cross-domain knowledge alignment of the target domain knowledge graph and the source domain knowledge graph to obtain a cross-domain knowledge graph comprises: Nodes and edges with similar semantics in the target domain knowledge graph and the source domain knowledge graph are identified; A cross-domain mapping matrix from the source domain knowledge graph to the target domain knowledge graph is constructed according to the identified nodes and edges; The target domain knowledge graph and the source domain knowledge graph are aligned according to the cross-domain mapping matrix to obtain the cross-domain knowledge graph.

4. The cross-domain long tail data augmentation method of claim 1, wherein, The multi-modal domain feature extraction on the input data to obtain multi-modal initial features comprises: A domain self-adaptive ResNeXt network is used to extract features from video frames in the input data to obtain domain visual features; A pre-trained multilingual BERT model is used to encode language text in the input data to obtain first encoding features; domain labels of the target domain are obtained; a target word embedding matrix is selected from a pre-constructed word embedding matrix set according to the domain labels; the first encoding features and the target word embedding matrix are fused to obtain domain language features; A time series encoder based on Transformer is used to encode action sensor data in the input data to obtain second encoding features; domain knowledge of the target domain is obtained; the second encoding features are optimized using the domain knowledge to obtain domain action features; The domain visual features, the domain language features, and the domain action features are integrated to obtain the multi-modal initial features.

5. The cross-domain long tail data augmentation method of claim 1, wherein, The use of the density peak clustering algorithm to identify long-tail scene data in the multi-modal initial features comprises: The density peak clustering algorithm is used to cluster the multi-modal initial features to obtain multiple clusters; The sample size of each cluster is calculated, and the global sample size of the multi-modal initial features is calculated; According to the sample quantity of each cluster and the global sample quantity, a distribution proportion of each cluster is calculated; An average local density of each cluster and a global average density of the multi-modal initial feature are calculated; A pre-configured proportion threshold and a density threshold are obtained; When it is detected that the distribution proportion corresponding to a cluster is less than the proportion threshold and the corresponding average local density is less than the density threshold, the features in the detected cluster are determined as the long-tail scene data.

6. The cross-domain long tail data augmentation method of claim 1, wherein, The long-tail scene data is enhanced by using a generative adversarial network and the cross-domain knowledge graph to obtain supplementary features, including: The long-tail scene data is enhanced by using a generative adversarial network to obtain similar samples; The long-tail scene data is queried in the cross-domain knowledge graph to obtain triple data; The triple data and the similar samples are spliced to obtain the supplementary features.

7. The cross-domain long tail data augmentation method of claim 1, wherein, After the multi-modal cross-domain enhanced features are obtained, the method further includes: In response to a processing instruction of a target task, an initial model corresponding to the target task is obtained; The initial model is trained by using the multi-modal cross-domain enhanced features to obtain a target model; The target task is executed by using the target model.

8. A cross-domain long tail data augmentation apparatus, comprising: The cross-domain long-tail data enhancement device includes: A construction unit is configured to, in response to an enhancement instruction of long-tail data in a target domain, construct a target domain knowledge graph of the target domain and a source domain knowledge graph of a source domain; An alignment unit is configured to perform cross-domain knowledge alignment on the target domain knowledge graph and the source domain knowledge graph to obtain a cross-domain knowledge graph; An extraction unit is configured to obtain input data according to the enhancement instruction and perform multi-modal domain feature extraction on the input data to obtain multi-modal initial features; An identification unit is configured to identify long-tail scene data in the multi-modal initial features by using a density peak value clustering algorithm; An enhancement unit is configured to enhance the long-tail scene data by using a generative adversarial network and the cross-domain knowledge graph to obtain supplementary features; An adding unit is configured to add the supplementary features to the multi-modal initial features to obtain multi-modal cross-domain enhanced features.

9. A computer device, comprising: The computer device includes: a memory storing at least one instruction; and a processor executing the instruction stored in the memory to implement the cross-domain long-tail data enhancement method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction, which is executed by a processor in a computer device to implement the cross-domain long-tail data enhancement method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Knowledge graph representation learning method through combination of entity hierarchy category

    CN107423820A