Data detection method and device based on comparison embedding alignment, equipment and medium

By comparing the data detection methods with embedded alignment, using feature extraction modules and prototype networks for small sample classification, and combining them with rule engines to identify illegal content, we solve the problems of high cost and insufficient semantic alignment in cross-domain detection, and achieve high-precision and robust violation detection.

CN120763702APending Publication Date: 2025-10-10PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510918106.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Existing technologies in cross-domain detection have problems such as high detection cost, poor generalization of small samples, and insufficient semantic alignment. It is difficult to effectively align the characteristics of illegal content on different platforms, especially in the fields of financial technology and health, where cross-domain detection performance is reduced.

Method used

A data detection method based on contrast embedding alignment is adopted. By acquiring and preprocessing data, a feature alignment module is generated, a feature alignment module is obtained, small sample classification is performed through the prototype network, and violation identification is performed in combination with the rule engine to output the detection results.

Benefits of technology

It achieves high-precision and high-robust detection of illegal content, can eliminate inter-domain differences in cross-domain scenarios, quickly classify small sample data in new domains, and improve detection accuracy and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120763702A_ABST
    Figure CN120763702A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, can be applied to business system platforms of medical health, financial science and technology and the like, and discloses a data detection method based on comparison embedding alignment, which comprises the following steps: acquiring historical violation data, and preprocessing the historical violation data; inputting the preprocessed historical violation data into a feature extraction module for training, and generating a feature alignment module; acquiring to-be-detected data, and preprocessing the to-be-detected data; inputting the preprocessed to-be-detected data into a feature alignment module for feature extraction, and outputting a to-be-detected embedded vector; inputting the to-be-detected embedded vector into a prototype network for small sample classification, and outputting a classification result; and performing violation identification on the classification result according to a rule engine, and outputting a detection result. According to the method, inter-domain differences are eliminated through feature alignment optimization, prototype network classification and rule engine secondary recognition are combined, and high-precision and high-robustness illegal content detection is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a data detection method, device, equipment and storage medium based on comparison, embedding and alignment. Background Art

[0002] With the explosive growth of internet content, detecting illegal content (such as violence, hate speech, and false information) has become a core challenge in network governance. Traditional detection methods mainly rely on supervised learning, which requires a large amount of labeled data to train models for specific domains. However, in practical applications, illegal content exhibits significant cross-domain diversity and dynamic evolution. For example, there are significant differences in content distribution across different platforms, and new forms of violations are constantly emerging. These characteristics pose many problems for traditional methods: on the one hand, existing domain adaptation methods require retraining models for each new domain, making it difficult to meet real-time detection needs; on the other hand, conventional small-sample learning methods experience a sharp decline in performance in cross-domain scenarios due to domain shift. In addition, existing cross-domain alignment techniques mostly focus on matching global feature distributions, but ignore the fine-grained alignment of violation semantics.

[0003] To address these issues, existing research has attempted to combine contrastive learning with domain adaptation techniques. For example, some studies have proposed a cross-domain few-shot learning framework with semantic consistency constraints, but this framework fails to explicitly model the cross-domain invariance of violation categories. Other studies have implemented cross-domain detection of graph data through anomaly-aware contrastive alignment, but this approach has difficulty generalizing to non-graph-structured content. Furthermore, some studies have employed cross-spatial alignment strategies, but these strategies rely on unsupervised data from the target domain and are not suitable for zero-shot domain transfer scenarios. Overall, existing methods are unable to simultaneously achieve metric learning of cross-domain violation semantics (e.g., common features of "hate speech" across different platforms) and domain-independent discriminative representation learning (e.g., eliminating platform-specific style interference).

[0004] In FinTech, illegal content may involve fraudulent information, false transaction propaganda, and illegal financial activities. This content is disseminated through various channels such as social media and online advertising, and the data styles and content formats vary significantly across different platforms. For example, new financial fraud methods are constantly updated and cross-platform and cross-channel, making it difficult to update detection models in real time to respond to emerging fraud methods. Furthermore, when conducting cross-domain detection, user behavior data and transaction information differ significantly across different financial platforms (such as banks, payment platforms, and investment platforms), leading to serious domain shift issues and affecting detection performance. Existing cross-domain alignment technologies cannot effectively align the semantic features of fraudulent behavior across different platforms, making accurate detection difficult.

[0005] In the health sector, illegal content may include false medical advertisements, pseudoscientific health advice, and misleading medical information. This information is also disseminated through various channels, including social media, health forums, and online advertising. The content styles and audiences vary across platforms. For example, health rumors on social media may spread in the form of images and short videos, while health forums are primarily text-based, making cross-platform and cross-modal false medical information detection difficult. Existing cross-domain alignment techniques cannot effectively align the semantic features of false medical information across different platforms, resulting in reduced detection performance.

[0006] Finally, the difficulty of aligning and fusing multimodal or multi-source data is also a major issue currently facing the detection of illegal content. In FinTech, multi-source data such as transaction data, user behavior data, and social media data from different platforms need to be integrated and analyzed, but these data differ in format, dimension, and semantic features, making them difficult to align and fuse effectively. In the health sector, false medical information may be presented in multiple modalities such as text, images, and videos, making it difficult to extract and align semantic features from data in different modalities. In addition, in small or zero-sample scenarios, how to use limited labeled data to achieve accurate detection while avoiding misjudgments due to differences in platform styles is also an urgent problem to be solved. Summary of the Invention

[0007] The main purpose of the present invention is to provide a data detection method, device, equipment and storage medium based on comparative embedding alignment, aiming to solve the problems of high cross-domain detection cost, poor generalization of small samples, and insufficient semantic alignment in the existing technology.

[0008] To achieve the above objectives, the present invention provides a data detection method based on contrast embedding alignment, comprising: Acquire historical violation data and pre-process the historical violation data; The pre-processed historical violation data is input into the feature extraction module for training to generate the feature alignment module; Acquiring data to be detected and preprocessing the data to be detected; The pre-processed data to be detected is input into the feature alignment module for feature extraction, and the embedded vector to be detected is output; Input the embedding vector to be detected into the prototype network for small sample classification, and output the classification result; The rule engine performs violation identification on the classification results and outputs the detection results.

[0009] Furthermore, to achieve the above-mentioned purpose, the present invention provides a data detection device based on contrast embedding alignment, comprising: A historical data module, used to obtain historical violation data and pre-process the historical violation data; The module training module is used to input the pre-processed historical violation data into the feature extraction module for training and generate a feature alignment module; A detection data module is used to obtain the data to be detected and pre-process the data to be detected; The embedding vector module is used to input the preprocessed data to be detected into the feature alignment module for feature extraction and output the embedding vector to be detected; A classification result module is used to input the embedding vector to be detected into the prototype network for small sample classification and output the classification result; The detection result module is used to identify violations of the classification results according to the rule engine and output the detection results.

[0010] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer device, which includes a memory, a processor, and a data detection program based on comparison embedded alignment stored in the memory and runnable on the processor, and when the data detection program based on comparison embedded alignment is executed by the processor, the steps of the data detection method based on comparison embedded alignment as described above are implemented.

[0011] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which a data detection program based on comparison embedded alignment is stored. When the data detection program based on comparison embedded alignment is executed by a processor, the steps of the data detection method based on comparison embedded alignment as described above are implemented.

[0012] Beneficial effects: The present invention relates to the field of data processing technology and can be applied to business system platforms such as communications, medical health and financial technology. It discloses a data detection method based on comparative embedding alignment, including: obtaining historical violation data and preprocessing the historical violation data; inputting the preprocessed historical violation data into the feature extraction module for training to generate a feature alignment module; obtaining data to be detected and preprocessing the data to be detected; inputting the preprocessed data to be detected into the feature alignment module for feature extraction and outputting the embedding vector to be detected; inputting the embedding vector to be detected into the prototype network for small sample classification and outputting the classification result; performing violation identification on the classification result according to the rule engine and outputting the detection result. The present invention optimizes the embedding alignment by inputting the data to be detected into the feature alignment module to eliminate the differences between domains. Then, the small sample data in the new domain is quickly classified by the prototype network, and secondary violation identification is performed in combination with the rule engine to achieve high-precision and high-robustness detection of illegal content. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which: Figure 1An application environment schematic diagram of the data detection method based on contrast embedding alignment in an embodiment of the present application; Figure 2 A flow schematic diagram of the data detection method based on contrast embedding alignment in an embodiment of the present application; Figure 3 A system framework schematic diagram of the data detection method based on contrast embedding alignment in an embodiment of the present application; Figure 4 A feature extraction module and prototype network structure schematic diagram of the data detection method based on contrast embedding alignment in an embodiment of the present application; Figure 5 A function module schematic diagram of the data detection device based on contrast embedding alignment in a preferred embodiment of the present application; Figure 6 A structure schematic diagram of a computer device in an embodiment of the present application; Figure 7 Another structure schematic diagram of a computer device in an embodiment of the present application. DETAILED DESCRIPTION

[0014] It should be understood that the specific embodiments described herein are merely illustrative of the present application and do not limit the present application.

[0015] The data detection method based on contrast embedding alignment provided by the embodiments of the present application can be applied in an application environment such as Figure 1 , wherein a user end communicates with a service end through a network. The service end can obtain historical violation data through the user end, preprocess the historical violation data, input the preprocessed historical violation data into a feature extraction module for training to generate a feature alignment module, obtain to-be-detected data, preprocess the to-be-detected data, input the preprocessed to-be-detected data into the feature alignment module for feature extraction to output to-be-detected embedding vectors, input the to-be-detected embedding vectors into a prototype network for small sample classification to output a classification result, and perform violation identification on the classification result according to a rule engine to output a detection result. The present application optimizes embedding alignment by inputting to-be-detected data into a feature alignment module to eliminate domain differences. Then, the prototype network is used to quickly classify small sample data in a new domain, and the rule engine is used for secondary violation identification, thereby realizing high-precision and high-robustness violation content detection. The user end can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers and portable wearable devices. The service end can be realized by an independent server or a server cluster composed of multiple servers. The present application will be described in detail through specific embodiments.

[0016] Please refer to Figure 2 , Figure 2This is a flow chart of an embodiment of a data detection method based on comparison and embedding alignment provided by the present invention. It should be noted that although a logical order is shown in the flow chart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0017] like Figure 2 As shown, the data detection method based on comparative embedding alignment proposed by the present invention includes the following steps: S100: Acquire historical violation data and pre-process the historical violation data; S200, inputting the pre-processed historical violation data into the feature extraction module for training to generate a feature alignment module; S300, obtaining data to be detected and preprocessing the data to be detected; S400, inputting the pre-processed data to be detected into a feature alignment module for feature extraction, and outputting an embedding vector to be detected; S500, inputting the embedding vector to be detected into the prototype network for small sample classification, and outputting the classification result; S600: Identify violations on the classification results according to a rule engine and output detection results.

[0018] In this embodiment, if Figure 3 As shown, the system first normalizes historical violation data and the data to be tested to ensure consistency and processability. The preprocessed data is then fed into the feature extraction module, which generates a feature alignment module using cross-domain feature alignment techniques and extracts embedding vectors for the data to be tested. These embedding vectors are then fed into the prototypical network for small-sample classification. The prototypical network rapidly adapts to new domains using a small number of labeled samples, significantly improving the model's robustness and adaptability in dynamic, open environments.

[0019] To further improve the accuracy and reliability of detection results, the system incorporates a rules engine for secondary filtering of classification results. The rules engine further checks the classification results based on pre-set rules (such as keyword filtering and specific pattern matching), marking or filtering content that does not meet the rules. This dual verification mechanism ensures the high quality and credibility of the final detection results.

[0020] For content flagged as suspicious by the system, the system undergoes manual review to ensure the accuracy of the detection results. The results of these manual reviews are collected and used for iterative model optimization. Through this feedback data, the model can continuously adjust and optimize its parameters and rules to better adapt to the ever-changing nature of illegal content. This closed-loop optimization mechanism enables the model to continuously improve and respond to new challenges.

[0021] The system in this embodiment supports joint representation learning of image and text data, making it applicable to a variety of illegal content detection scenarios. Through the collaborative training of a shared encoder and domain discriminator, end-to-end cross-domain feature alignment is achieved in the content review field. Combining a rule engine with manual review, the system not only improves the accuracy of detection results but also enhances the model's adaptability in complex environments.

[0022] For example, in the healthcare sector, violations may include medical fraud, misuse of medical insurance, and non-standard diagnosis and treatment procedures. By acquiring and preprocessing historical violation data, we can remove noise, fill missing values, and standardize data formats, providing a high-quality data foundation for subsequent analysis. The preprocessed data is input into the feature extraction module for training, generating a feature alignment module that can extract key features from massive amounts of medical data, such as patient frequency of medical visits, abnormal fluctuations in medical expenses, and medication usage. The preprocessed data is then input into the feature alignment module for feature extraction, which outputs an embedding vector for testing. This embedding vector is then input into the prototype network for small-sample classification, enabling rapid identification of potential violation patterns. Finally, the rule engine identifies violations based on established medical compliance rules and outputs detection results, helping medical institutions promptly detect and correct violations, ensuring the rational use of medical insurance funds and improving the quality and standardization of medical services.

[0023] In the FinTech sector, violations may involve financial fraud, money laundering, illegal transactions, and more. The acquisition and preprocessing of historical violation data is fundamental. Data cleaning and organization ensure its accuracy and usability. The feature extraction module extracts key features from complex financial transaction data, such as transaction amount, frequency, time, and the identities of both parties. For the financial transaction data to be tested, after preprocessing and feature extraction, an embedding vector is generated and input into a prototype network for small-sample classification. The prototype network quickly classifies new transaction data based on known violation patterns to determine potential violation risks. The rule engine, combined with financial regulatory laws and regulations and internal compliance rules, further identifies violations based on the classification results and outputs the final detection results. This process effectively enhances the risk prevention and control capabilities of FinTech platforms, ensures the security and compliance of financial transactions, and maintains the stability of the financial market.

[0024] In one embodiment, step S100 specifically includes: S101. Acquire historical violation data, and add violation category labels and domain labels to the historical violation data; S102, classifying the historical violation data to obtain image violation data and text violation data; S103, performing normalization processing and channel standardization processing on the image violation data; S104: performing subword segmentation and dynamic filling processing on the text violation data.

[0025] In this embodiment, the data preprocessing module is mainly responsible for converting the original historical violation data into a format suitable for subsequent processing. Historical violation data usually comes from different data sources, such as social media platforms, e-commerce platforms, financial transaction records, etc. These data may contain images, text or other forms of content. Each violation sample will be assigned a violation category label (such as "counterfeit goods", "violent content", "hate speech", etc.) to indicate its violation type. These labels are the basis of supervised learning and are used to train the model to identify different types of violation content. In addition, each sample will be assigned a domain label to indicate its source domain (such as "North American social platform", "Southeast Asian e-commerce platform", etc.). The domain label plays an important role in subsequent domain adversarial training, helping the model learn cross-domain feature alignment.

[0026] Historical violation data is categorized into image violation data and text violation data based on its content type. Image violation data includes violation samples with image content, such as product images and images from social media; text violation data includes violation samples with text content, such as product descriptions and text content from social media. The purpose of classification is to feed different types of data into corresponding processing flows for subsequent feature extraction and alignment.

[0027] For image data, the data preprocessing module adjusts its pixel values ​​to a uniform range, such as [0, 1]. This is done by dividing each pixel value by 255 (assuming the original pixel value range is [0, 255]). Furthermore, each channel of the image (such as the RGB channels) is normalized to reduce statistical differences between different images. This is done by subtracting the mean and dividing the standard deviation from the pixel values ​​of each channel. After normalization and channel standardization, the image data is typically resized to a uniform size, such as 224×224 pixels, for input to a visual encoder (such as ViT).

[0028] For text data, the data preprocessing module uses a subword segmentation algorithm (such as BERT's WordPiece or SentencePiece) to split it into smaller units (subwords) so that the model can better handle vocabulary variations and combinations. For example, the word "unbelievable" may be segmented into "un", "be", "liev", and "able". Then, based on the length of the longest text sequence in the batch, other shorter text sequences are dynamically padded (usually with special fillers such as `[PAD]`) to make their length uniform so that they can be input into the text encoder (such as RoBERTa-large). After subword segmentation and dynamic padding, each text sequence is converted into a fixed-length subword index sequence.

[0029] From the above detailed explanation, it can be seen that the data preprocessing module plays an important role in data standardization and format unification, providing high-quality input data for subsequent feature extraction and model training.

[0030] For example, in the healthcare sector, violations may involve medical image falsification, medical record tampering, and health insurance fraud. By acquiring historical violation data and annotating it with violation categories and domain labels, the type of violation (e.g., image falsification, medical record tampering) and the data source (e.g., hospital information systems, health insurance reimbursement systems) can be clearly identified. Categorizing historical violation data into image violations (e.g., forged X-rays and CT images) and text violations (e.g., falsified medical records and false diagnostic reports) allows for targeted processing. For image violation data, normalization can unify images captured by different devices to the same scale and range, eliminating the impact of device differences. Channel normalization can adjust the image's color channels to meet model input requirements. This processed image data enables more accurate model recognition, effectively detecting forged or tampered images. For text violation data, subword segmentation can break down text in medical records or diagnostic reports into smaller semantic units, while dynamic padding ensures consistent text length for easier model processing. The text data after these processes can be better analyzed by natural language processing models to identify violations such as medical record tampering and false diagnosis.

[0031] In the FinTech sector, violations may involve financial fraud, money laundering, and fraudulent transactions. Labeling historical violation data with violation category and domain labels can clearly identify the type of violation (e.g., money laundering, fraud) and the data source (e.g., bank transaction records, payment platform data). Categorizing data into image violations (e.g., forged bills and seals) and text violations (e.g., false transaction records and altered contract texts) allows for targeted processing. For image violations, normalization and channel standardization ensure consistency in size and color across images of forged bills or seals, facilitating model recognition of forgery features. This processed image data can help financial institutions quickly identify counterfeit bills or seals and prevent financial fraud. For text violations, subword segmentation and dynamic filling can break down transaction records or contract texts into smaller semantic units and standardize text length. This processed text data can be better analyzed by natural language processing models to identify false transaction records or altered contract terms, effectively preventing financial fraud and money laundering.

[0032] In one embodiment, step S200 inputs the pre-processed historical violation data into a feature extraction module for training to generate a feature alignment module, including: S2011, the shared encoder includes a visual encoder and a text encoder; S2012, inputting the pre-processed historical violation data into a shared encoder, and outputting an embedding vector; S2013: Input the embedding vector into the domain discriminator to perform contrastive learning and domain adversarial training to generate a feature alignment module.

[0033] In this embodiment, Figure 4 As shown in Figure 1, the feature extraction module primarily consists of two components: a shared encoder and a domain discriminator. The shared encoder is a unified feature extractor for multimodal data (image and text), consisting of a visual encoder (such as VisionTransformer, ViT) and a text encoder (such as RoBERTa-large). Image data, after being resized (e.g., to 224×224 pixels) and channel-normalized, is fed into the visual encoder (e.g., ViT-B / 16). Image embeddings are generated through block segmentation, positional encoding, and a Transformer layer. Text data, after undergoing subword segmentation and dynamic padding, is fed into the text encoder. Subword segmentation, input embedding, and a Transformer layer generate text embeddings. The output of the shared encoder is an embedding vector for image and text data in a unified semantic space, enabling data from different modalities to be compared and processed in the same space.

[0034] The domain discriminator is a two-layer multi-layer perceptron (MLP) that distinguishes the domain information of input features. Through domain adversarial training, the feature extraction module utilizes a gradient reversal layer (GRL) to keep the input unchanged during forward propagation and multiply the gradient by a negative constant during backward propagation. This mechanism forces the encoder to generate domain-invariant features by minimizing the adversarial training loss function, making it impossible for the domain discriminator to distinguish the domain information of the input features, thus achieving feature alignment.

[0035] The feature extraction module generates the feature alignment module through contrastive learning and domain adversarial training. In contrastive learning, cross-domain sample pairs are constructed, including positive sample pairs (samples from different domains but belonging to the same category) and negative sample pairs (samples from different domains but belonging to different categories). A contrastive loss function (such as the InfoNCE loss) is used to calculate the similarity between embedding vectors, minimizing the embedding distance between positive sample pairs and maximizing the embedding distance between negative sample pairs. This results in samples of the same type from different domains being closer in the embedding space, while samples of different categories are further apart.

[0036] The goal of the feature alignment module is to achieve cross-domain semantic consistency, that is, similar samples in different domains have similar representations in the embedding space, eliminate domain-specific interference, and generate domain-independent discriminative representations to improve the model's generalization ability in cross-domain tasks. In addition, the shared encoder supports cross-modal alignment, can process image and text data simultaneously, and supports violation detection of multimodal content. The feature extraction module also supports end-to-end training, eliminating the need for manual design of feature extraction steps, simplifying the model training process and significantly improving the robustness of the model in cross-domain scenarios, enabling it to maintain high classification performance on different platforms or in different language environments.

[0037] From the above detailed explanation, it can be seen that the feature extraction module plays an important role in multimodal data processing and cross-domain feature alignment, providing high-quality feature representation for subsequent violation detection tasks and significantly improving the generalization ability and robustness of the model.

[0038] For example, in the healthcare sector, data sources are extensive and complex, such as medical records and imaging data from different hospitals and departments. These data may have significant differences in format, annotation standards, and distribution. By inputting pre-processed historical violation data into a shared encoder, data from different sources can be mapped into a unified feature space to generate embedding vectors. These embedding vectors can capture the core features of the data, such as key symptoms in medical records and pathological features in images. For example, in medical insurance fraud detection, the format and annotation of medical record data from different hospitals may differ. Through the feature alignment module, these data can be unified into a feature space, thereby more accurately identifying violations such as abnormal medical behavior and false medical records, thereby improving the efficiency and security of medical insurance fund use.

[0039] In the field of FinTech, data is also diverse and heterogeneous. For example, transaction records and user behavior logs from different financial institutions vary in format and distribution. By inputting preprocessed historical violation data into a shared encoder, key features such as transaction amount, transaction frequency, and user behavior patterns can be extracted and embedded into vectors. For example, in financial fraud detection, transaction data formats and annotations may differ between different banks. The feature alignment module can unify this data into a single feature space, enabling more accurate identification of violations such as abnormal transaction behavior and fraudulent user identities, thereby ensuring the security and compliance of financial transactions.

[0040] In one embodiment, step S2012 includes: The shared encoder is composed of a visual encoder and a text encoder; Inputting the image violation data into a visual encoder, wherein the visual encoder processes the image violation data to generate an image embedding vector; The text violation data is input into a text encoder, and the text encoder processes the text violation data to generate a text embedding vector.

[0041] In this embodiment, the shared encoder's main function is to map the input image and text data into a unified semantic space. It consists of two parts: a visual encoder and a text encoder, which are used to process image and text data respectively.

[0042] The visual encoder receives image data as input, which may include product images, social media content, and more. Before input, the image data undergoes preprocessing, such as resizing to a uniform size (e.g., 224×224 pixels) and channel normalization (e.g., subtracting the mean and dividing by the standard deviation). The visual encoder processes the image data using a deep learning model, such as the Vision Transformer (ViT). The Vision Transformer is a vision model based on the Transformer architecture that effectively captures both local and global features in images. Its workflow includes: patch embedding, which divides the input image into fixed-size patches (e.g., 16×16 pixels) and flattens them into a one-dimensional vector; positional encoding, which adds a positional encoding to each patch to preserve the spatial information of the image; and a Transformer layer, which encodes the image patches using a multi-head self-attention mechanism and a feed-forward network. The final output is a fixed-dimensional image embedding vector (e.g., 768 dimensions). For example, when using the ViT-B / 16 model, the input image is divided into 16×16 blocks and encoded through multiple layers of Transformer to generate a 768-dimensional image embedding vector.

[0043] The text encoder receives text violation data as input, which may include product descriptions, text content from social media, and so on. Before input, the text data goes through preprocessing modules such as subword tokenization and dynamic padding. Subword tokenization divides the text into subword units, such as using BERT's WordPiece tokenization method; dynamic padding pads the text sequence to a uniform length for input into the model. The text encoder uses a deep learning model (such as RoBERTa-large) to process the text data. RoBERTa-large is a pre-trained language model based on the Transformer architecture that can effectively capture semantic information in the text. Its workflow includes: input embedding, which maps subword units to embedding vectors and adds positional encoding; Transformer layer, which encodes the text sequence through a multi-head self-attention mechanism and a feedforward network; and finally outputs a fixed-dimensional text embedding vector (such as 768 dimensions). For example, when using the RoBERTa-large model, the input text is segmented into subword units and encoded through multiple layers of Transformer to generate a 768-dimensional text embedding vector.

[0044] The shared encoder maps the output embedding vectors of the visual encoder and text encoder into a semantic space of the same dimensionality, allowing image and text data to be compared and processed in the same space. Through contrastive learning and domain adversarial training, the shared encoder learns the semantic consistency of image and text data, achieving cross-modal alignment and ensuring that data from different modalities have similar representations in the embedding space. Furthermore, the embedding vectors generated by the shared encoder can be used for subsequent feature alignment and classification tasks, such as optimizing the embedding alignment of cross-domain samples using a contrastive loss function.

[0045] For example, in the healthcare field, image violation data may include forged medical images (such as X-rays, CT images, and MRI images), while text violation data may involve tampered medical records and false diagnosis reports. By feeding the image violation data into a visual encoder, key features in the image, such as the shape, size, and texture of the lesion area, can be extracted to generate image embedding vectors. These embedding vectors can capture subtle differences in the image, helping the model identify forged or tampered images. Simultaneously, feeding the text violation data into a text encoder can extract key information from medical records, such as symptom descriptions, diagnosis results, and treatment processes, and generate text embedding vectors. These embedding vectors can capture the semantic information in the text, helping the model identify medical record tampering or false diagnoses.

[0046] In the fintech sector, image violation data may include forged bills, seals, and ID documents, while text violation data may involve fraudulent transaction records or altered contract terms. By feeding image violation data into a visual encoder, features such as the text on the bill and the shape and texture of the seal can be extracted to generate image embeddings. These embeddings help the model identify forged bills or seals, thereby preventing financial fraud. Simultaneously, feeding text violation data into a text encoder extracts transaction records, including the amount, transaction time, and information about the two parties involved, to generate text embeddings. These embeddings help the model identify fraudulent transaction records or altered contract terms.

[0047] In one embodiment, the step S200 further includes: S2021. Construct a cross-domain sample pair according to the embedding vector, where the cross-domain sample pair includes a positive sample pair and a negative sample pair; S2022. Calculate the similarity between the embedded vectors using a contrast loss function; S2023. By minimizing the contrast loss function, dynamically adjust the embedding distance between positive sample pairs and negative sample pairs to generate a feature alignment module.

[0048] In this embodiment, in the contrastive learning component, the input embedding vectors are generated by a feature extraction module (such as VisionTransformer or RoBERTa-large). These embedding vectors represent samples in a unified semantic space. The core of contrastive learning is to learn the similarities between samples by constructing positive and negative sample pairs. Positive sample pairs are pairs of samples from different domains but belonging to the same category, such as "counterfeit goods" samples in the source domain and "counterfeit goods" samples in the target domain. Negative sample pairs are pairs of samples from different domains and belonging to different categories, such as "normal goods" samples in the source domain and "counterfeit goods" samples in the target domain.

[0049] For each positive sample pair, samples from different domains but belonging to the same category are selected to force these samples to be close in the embedding space. For example, the distance between the "counterfeit goods" sample in the source domain and the "counterfeit goods" sample in the target domain will be shortened in the embedding space. For each negative sample pair, samples from different domains and belonging to different categories are randomly selected to ensure that these samples are far apart in the embedding space. In this way, the model can learn the semantic consistency of cross-domain samples, so that similar samples from different domains have similar representations in the embedding space.

[0050] Contrastive learning uses a contrastive loss function to measure the similarity between samples. Common contrastive loss functions include the InfoNCE loss function. The goal of contrastive learning is to minimize the embedding distance between pairs of positive samples while maximizing the embedding distance between pairs of negative samples. This can be achieved by minimizing the contrastive loss function. By minimizing the contrastive loss function, positive pairs of samples from different domains but belonging to the same category are brought closer together in the embedding space. By minimizing the contrastive loss function, negative pairs of samples from different domains and belonging to different categories are pushed further apart in the embedding space.

[0051] Through contrastive learning, the feature alignment module generates an embedding space for cross-domain feature alignment. The goal of the feature alignment module is to ensure that similar samples from different domains have similar representations in the embedding space, thereby achieving cross-domain feature alignment. The construction of cross-domain positive sample pairs forces samples from different domains but belonging to the same category to be closer in the embedding space. For example, "counterfeit goods" samples in the source domain and "counterfeit goods" samples in the target domain are closer in the embedding space. In this way, the model can learn the semantic consistency of cross-domain samples, improving the accuracy of cross-domain classification.

[0052] Dynamic negative sampling and the construction of cross-domain positive sample pairs enable the model to be exposed to a wider variety of samples during training, enhancing its generalization ability. By minimizing the embedding distance between positive sample pairs and maximizing the embedding distance between negative sample pairs, the model can more accurately distinguish between samples of different categories, thereby improving classification accuracy.

[0053] For example, in the healthcare sector, data comes from a wide range of sources, such as medical records and imaging data from different hospitals and departments. This data may have significant differences in format, annotation standards, and distribution. By constructing cross-domain sample pairs (including positive and negative pairs), data from different domains with the same violation can be used as positive pairs, while data from different domains with different behaviors can be used as negative pairs. For example, a positive pair could be a case of the same type of image manipulation from different hospitals, while a negative pair could be a combination of real and fabricated images. By calculating the similarity between embedding vectors using a contrastive loss function, the model can learn which features are associated with the violation and which features are noise due to differences in data sources. By minimizing the contrastive loss function and dynamically adjusting the embedding distance between positive and negative pairs, the model can bring samples with the same violation closer together and push samples with different violations further apart. The resulting feature alignment module can better handle data across hospitals and departments, improving the model's generalization capabilities across different medical scenarios.

[0054] In the fintech sector, data is similarly diverse and heterogeneous. For example, transaction records and user behavior logs from different financial institutions differ in format and distribution. By constructing cross-domain sample pairs, data from different financial institutions with the same violation can be used as positive pairs, while data with different violations can be used as negative pairs. For example, a positive pair could be a fraudulent transaction of the same type from different banks, while a negative pair could be a combination of normal and fraudulent transactions. By calculating the similarity between embedding vectors using a contrastive loss function, the model can learn which features are associated with violations and which features are noise due to differences in data sources. By minimizing the contrastive loss function and dynamically adjusting the embedding distance between positive and negative pairs, the model can bring samples with the same violation closer together and push samples with different violations further apart. The resulting feature alignment module can better handle data across institutions and businesses, improving the model's generalization capabilities in different scenarios.

[0055] In one embodiment, step S200 includes: S2031, inputting the embedding vector into the first fully connected layer of the domain classifier for mapping processing to generate an intermediate dimension; S2032. Map the intermediate dimension to the second fully connected layer of the domain classifier, and output the domain probability of the embedded vector in each domain; S2033, inputting the embedding vector into a gradient reversal layer for adversarial training to generate domain-invariant features; S2034. Train the model according to the domain probabilities and domain-invariant features to generate a feature alignment module.

[0056] In this embodiment, the domain adversarial training component uses a two-layer multilayer perceptron (MLP) with a first hidden layer dimension of 512. The embedding vector is first input into the first fully connected layer of the domain classifier for mapping, generating feature representations of intermediate dimensions. These intermediate features are then input into the second fully connected layer of the domain classifier to calculate the probability distribution for each domain.

[0057] To make the feature representation insensitive to domain information, the embedding vector is also fed into a gradient reversal layer (GRL). GRL is the core mechanism of domain adversarial training. It keeps the input unchanged during forward propagation but multiplies the gradient by a negative constant during backward propagation. This design aims to deceive the domain classifier into being unable to distinguish the domain information of the input features. In this way, the encoder is forced to generate domain-invariant features, that is, features that have a similar distribution across different domains.

[0058] During training, the model aims to make the output of the domain classifier as close to random guessing as possible—that is, to make it impossible to distinguish domain information about the input features. Domain adversarial training achieves this by maximizing the uncertainty of the domain classifier. Specifically, the model aims to make the output of the domain classifier as close to random guessing as possible, thereby forcing the encoder to generate domain-invariant features. This design ensures that the gradients of the domain classifier, when backpropagated to the encoder, have the opposite effect, further promoting domain invariance of the features.

[0059] Through domain adversarial training, the model can generate domain-invariant features, thereby eliminating feature differences between different domains and improving the model's generalization ability in cross-domain tasks. In cross-domain scenarios, the model's tolerance to domain shift is significantly improved, and it can maintain high classification performance across different platforms or language environments. In addition, the domain adversarial training component is combined with the feature alignment module to achieve end-to-end training, eliminating the need for manual adjustment of feature alignment steps and simplifying the model training process.

[0060] For example, in the healthcare sector, data comes from a wide range of sources, such as medical records and imaging data from different hospitals. This data can have significant differences in format, annotation standards, and distribution. For example, in health insurance fraud detection, medical record data formats and annotations may vary from hospital to hospital. Using the feature alignment module, the model can more accurately identify anomalous medical behavior, such as false medical records or forged images, while ignoring data format differences between hospitals. This helps improve the efficiency and security of medical insurance funds.

[0061] In the FinTech sector, data is also diverse and heterogeneous. For example, transaction records and user behavior logs from different financial institutions vary in format and distribution. For example, in financial fraud detection, transaction data formats and annotations may differ between banks. Using the feature alignment module, the model can more accurately identify anomalous transaction behavior, such as fraudulent transaction records or forged user identities, while ignoring data format differences between banks. This helps improve the security and compliance of financial transactions.

[0062] In one embodiment, step S500 includes: S501, input the embedding vector to be detected for each category into the prototype network; S502, calculating the mean of the embedding vector to be detected of each category by the prototype network, and using the mean as the prototype vector of the category; S503, calculating a distance metric based on the to-be-detected embedded vector of each category and the prototype vector of each category; S504: Classify the data to be detected according to the distance metric to obtain a classification result.

[0063] In this embodiment, Figure 4 As shown in the figure, in the small sample classification module, the data to be tested is first input into the prototype network. This data has been processed by the data preprocessing module and the feature alignment module (SNDA) to generate embedding vectors. These embedding vectors represent the data to be tested in a unified semantic space and are used for subsequent classification operations.

[0064] The core idea of ​​the prototypical network is to construct a prototype vector for each category using a support set. The support set is a collection of a small number of samples with well-labeled categories, which is used to define the feature center of each category. For each category, the prototypical network calculates the mean of the embedding vectors of all samples in the support set to obtain the category prototype vector. In traditional prototypical networks, the Euclidean distance is typically used to measure the similarity between the sample to be detected and the category prototype vector. However, this paper introduces a learnable Mahalanobis distance to replace the Euclidean distance to improve classification accuracy. Specifically, the covariance matrix captures the distribution characteristics of samples within a category, thereby adaptively adjusting the importance of different feature dimensions. The inverse of the covariance matrix plays a weighting role in the Mahalanobis distance formula: if the variance of a feature dimension is large (i.e., the sample distribution of that dimension is more dispersed), then the weight of that dimension in the distance calculation will be smaller; conversely, if the variance of a feature dimension is small (i.e., the sample distribution of that dimension is more concentrated), then the weight of that dimension in the distance calculation will be larger. This weighting mechanism enables the model to more accurately measure the similarity between the sample to be detected and the category prototype vector.

[0065] After calculating the Mahalanobis distance between the sample under test and each class prototype vector, the class with the closest distance is selected as the classification result. To convert distances into classification confidence, the present invention employs softmax normalization. Specifically, the distance metric for each class is converted into a similarity score, which is then normalized using the softmax function to obtain the classification confidence for each class. Finally, the class with the highest confidence score is selected as the classification result for the sample under test.

[0066] Furthermore, the present invention introduces a dynamic update mechanism. After calculating the mean distances between data features across all categories, the category with the closest distance is selected, and the samples corresponding to the current features are added to that category. Subsequently, the mean feature vector for that category is recalculated. Through repeated iterations of the model, the mean feature vector for each category is gradually converged. This dynamic update mechanism enables the model to continuously adjust the category prototype vector based on newly added samples, thereby better adapting to new data and improving classification accuracy and robustness.

[0067] By introducing a learnable covariance matrix and Mahalanobis distance, the model can more accurately measure the similarity between the sample to be tested and the class prototype vector, thereby improving classification accuracy. The dynamic update mechanism for the class prototype vector enables the model to continuously adjust to new data, enhancing its adaptability to new samples. Only a small number of labeled samples are required to achieve high detection accuracy in new domains, significantly reducing labeling costs. In cross-domain scenarios, the model's tolerance to domain shift is significantly improved, maintaining high classification performance across different platforms and languages.

[0068] For example, in the healthcare sector, data classification can help identify unusual medical practices, such as medical insurance fraud, medical record tampering, or image fabrication. For example, in healthcare fraud detection, the model can input the embedding vectors of medical record or image data into a prototype network to generate prototype vectors for normal and fraudulent behavior. For new data to be tested, the model calculates the distance between its embedding vector and the two prototype vectors. If the distance to the prototype vector for fraudulent behavior is closer, it is classified as fraudulent. This approach can effectively identify unusual medical behavior or forged images, improving the efficiency and security of medical insurance funds.

[0069] In the field of FinTech, data classification can help identify unusual financial transactions, such as fraudulent transactions and money laundering. For example, in financial fraud detection, the model can input the embedding vectors of transaction records into a prototype network to generate prototype vectors for normal and fraudulent transactions. For new transaction data to be detected, the model calculates the distance between its embedding vector and the two prototype vectors. If the distance to the prototype vector of a fraudulent transaction is closer, it is classified as fraudulent. This approach can effectively identify unusual transaction behavior and improve the security and compliance of financial transactions.

[0070] In one embodiment, a data detection device based on contrast embedding alignment is provided, and the data detection device based on contrast embedding alignment corresponds to the data detection method based on contrast embedding alignment in the above embodiment. Figure 5 , Figure 5 This is a functional module diagram of a preferred embodiment of a data detection device based on comparative embedding alignment according to the present invention. It includes a historical data module 10, a module training module 20, a detection data module 30, an embedding vector module 40, a classification result module 50, and a detection result module 60. Each functional module is described in detail as follows: A historical data module 10 is used to obtain historical violation data and pre-process the historical violation data; The module training module 20 is used to input the pre-processed historical violation data into the feature extraction module for training to generate a feature alignment module; The detection data module 30 is used to obtain the data to be detected and pre-process the data to be detected; Embedding vector module 40, used to input the pre-processed data to be detected into the feature alignment module for feature extraction and output the embedding vector to be detected; A classification result module 50 is used to input the embedding vector to be detected into the prototype network for small sample classification and output the classification result; The detection result module 60 is used to identify violations of the classification results according to the rule engine and output the detection results.

[0071] In one embodiment, the historical data module 10 specifically includes: A historical data unit, configured to obtain historical violation data and add violation category labels and domain labels to the historical violation data; A data classification unit, configured to classify the historical violation data to obtain image violation data and text violation data; An image data processing unit, configured to perform normalization processing and channel standardization processing on the image violation data; The text data processing unit is used to perform subword segmentation processing and dynamic filling processing on the text violation data.

[0072] In one embodiment, the module training module 20 includes: The feature extraction module includes a shared encoder and a domain discriminator; an embedding vector unit, configured to input the preprocessed historical violation data into a shared encoder and output an embedding vector; The module training unit is used to input the embedding vector into the domain discriminator to perform contrastive learning and domain adversarial training to generate a feature alignment module.

[0073] In one embodiment, the embedding vector unit includes: The shared encoder includes a visual encoder and a text encoder; Inputting the image violation data into a visual encoder, wherein the visual encoder processes the image violation data to generate an image embedding vector; The text violation data is input into a text encoder, and the text encoder processes the text violation data to generate a text embedding vector.

[0074] In one embodiment, the module training module 20 further includes: A sample pair unit, configured to construct a cross-domain sample pair according to the embedding vector, wherein the cross-domain sample pair includes a positive sample pair and a negative sample pair; Similarity unit, used to calculate the similarity between embedding vectors using contrast loss function; The module training unit is used to dynamically adjust the embedding distance between positive and negative sample pairs by minimizing the contrast loss function to generate a feature alignment module.

[0075] In one embodiment, the module training module 20 includes: An intermediate dimension unit, configured to input the embedding vector into a first fully connected layer of a domain classifier for mapping processing to generate an intermediate dimension; A domain probability unit, configured to map the intermediate dimension to the second fully connected layer of the domain classifier and output the domain probability of the embedded vector in each domain; A feature unit, configured to input the embedding vector into a gradient reversal layer for adversarial training to generate domain-invariant features; The module training unit is used to train the model according to the domain probability and the domain invariant features to generate a feature alignment module.

[0076] In one embodiment, the classification result module 50 includes: The prototype network unit is used to input the embedding vector of each category to be detected into the prototype network; A prototype vector unit, configured to calculate the mean of the embedding vectors to be detected for each category using the prototype network, and use the mean as the prototype vector of the category; A distance metric unit, configured to calculate a distance metric based on the embedding vector to be detected of each category and the prototype vector of each category; The classification result unit is used to classify the data to be detected according to the distance metric to obtain a classification result.

[0077] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external user terminal via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a data detection method based on comparative embedded alignment.

[0078] In one embodiment, a computer device is provided. The computer device may be a user terminal, and its internal structure diagram may be as follows: Figure 7As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps of a data detection method based on comparison embedding alignment on the user side. In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed: Acquire historical violation data and pre-process the historical violation data; The pre-processed historical violation data is input into the feature extraction module for training to generate the feature alignment module; Acquiring data to be detected and preprocessing the data to be detected; The pre-processed data to be detected is input into the feature alignment module for feature extraction, and the embedded vector to be detected is output; Input the embedding vector to be detected into the prototype network for small sample classification, and output the classification result; The rule engine performs violation identification on the classification results and outputs the detection results.

[0079] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: Acquire historical violation data and pre-process the historical violation data; The pre-processed historical violation data is input into the feature extraction module for training to generate the feature alignment module; Acquiring data to be detected and preprocessing the data to be detected; The pre-processed data to be detected is input into the feature alignment module for feature extraction, and the embedded vector to be detected is output; Input the embedding vector to be detected into the prototype network for small sample classification, and output the classification result; The rule engine performs violation identification on the classification results and outputs the detection results.

[0080] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the user side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0081] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the above-described method embodiments. Any reference to memory, storage, database, or other media used in the various embodiments provided herein may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAMbus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM).

[0082] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0083] It should be noted that if any software tools or components other than those of the Company appear in the embodiments of this application, they are merely for illustration and do not represent actual use. The above embodiments are intended only to illustrate the technical solutions of the present invention, not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some of the technical features therein with equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A data detection method based on contrast embedding alignment, characterized in that: The following steps are involved: Acquire historical violation data and pre-process the historical violation data; The pre-processed historical violation data is input into the feature extraction module for training to generate the feature alignment module; Acquiring data to be detected and preprocessing the data to be detected; The pre-processed data to be detected is input into the feature alignment module for feature extraction, and the embedded vector to be detected is output; Input the embedding vector to be detected into the prototype network for small sample classification, and output the classification result; The rule engine performs violation identification on the classification results and outputs the detection results.

2. The data detection method based on contrast embedding alignment according to claim 1, characterized in that: The obtaining of historical violation data and pre-processing of the historical violation data specifically includes: Acquire historical violation data, and add violation category labels and domain labels to the historical violation data; Classify the historical violation data to obtain image violation data and text violation data; Performing normalization processing and channel standardization processing on the image violation data; Subword segmentation and dynamic filling processing are performed on the text violation data.

3. The data detection method based on contrast embedding alignment according to claim 2, characterized in that: The pre-processed historical violation data is input into the feature extraction module for training to generate a feature alignment module, including: The feature extraction module includes a shared encoder and a domain discriminator; Inputting the preprocessed historical violation data into a shared encoder and outputting an embedding vector; The embedding vector is input into the domain discriminator for contrastive learning and domain adversarial training to generate a feature alignment module.

4. The data detection method based on contrast embedding alignment according to claim 2, characterized in that: The step of inputting the pre-processed historical violation data into a shared encoder and outputting an embedding vector comprises: The shared encoder includes a visual encoder and a text encoder; Inputting the image violation data into a visual encoder, wherein the visual encoder processes the image violation data to generate an image embedding vector; The text violation data is input into a text encoder, and the text encoder processes the text violation data to generate a text embedding vector.

5. The data detection method based on contrast embedding alignment according to claim 2, characterized in that: The step of inputting the pre-processed historical violation data into a feature extraction module for training to generate a feature alignment module further includes: Constructing a cross-domain sample pair according to the embedding vector, wherein the cross-domain sample pair includes a positive sample pair and a negative sample pair; Calculate the similarity between embedding vectors through contrast loss function; By minimizing the contrastive loss function, the embedding distance between positive and negative sample pairs is dynamically adjusted to generate a feature alignment module.

6. The data detection method based on contrast embedding alignment according to claim 2, characterized in that: The pre-processed historical violation data is input into the feature extraction module for training. Generate feature alignment module, including: Inputting the embedding vector into the first fully connected layer of the domain classifier for mapping processing to generate an intermediate dimension; Mapping the intermediate dimension to the second fully connected layer of the domain classifier, outputting the domain probability of the embedding vector in each domain; Inputting the embedding vector into the gradient reversal layer for adversarial training to generate domain-invariant features; The model is trained according to the domain probabilities and domain-invariant features to generate a feature alignment module.

7. The data detection method based on contrast embedding alignment according to claim 1, characterized in that: The step of inputting the embedding vector to be detected into the prototype network for small sample classification and outputting the classification result includes: Input the embedding vector of each category to be detected into the prototype network; The prototype network calculates the mean of the embedding vector to be detected for each category, and uses the mean as the prototype vector of the category; Calculating a distance metric based on the embedding vector to be detected of each category and the prototype vector of each category; The data to be detected is classified according to the distance metric to obtain a classification result.

8. A data detection device based on contrast embedding alignment, characterized in that: The data detection device based on comparison embedding alignment includes: A historical data module, used to obtain historical violation data and pre-process the historical violation data; The module training module is used to input the pre-processed historical violation data into the feature extraction module for training and generate a feature alignment module; A detection data module is used to obtain the data to be detected and pre-process the data to be detected; The embedding vector module is used to input the preprocessed data to be detected into the feature alignment module for feature extraction and output the embedding vector to be detected; A classification result module is used to input the embedding vector to be detected into the prototype network for small sample classification and output the classification result; The detection result module is used to identify violations of the classification results according to the rule engine and output the detection results.

9. A computer device, characterized in that: The computer device includes a memory, a processor, and a data detection program based on comparison and embedded alignment stored in the memory and capable of running on the processor. When the data detection program based on comparison and embedded alignment is executed by the processor, the steps of the data detection method based on comparison and embedded alignment as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that The storage medium stores a data detection program based on contrast, embedding, and alignment. When the data detection program based on contrast, embedding, and alignment is executed by a processor, the steps of the data detection method based on contrast, embedding, and alignment as described in any one of claims 1 to 7 are implemented.