Cross-language information data acquisition and structured processing method based on adaptive learning

By introducing adaptive learning methods in cross-language information data processing, including noise prior learning and covariance matrix prediction network, the problems of high noise, difficulty in alignment, and poor adaptability of low-resource languages ​​are solved, and more efficient and accurate data processing is achieved.

CN120011479AActive Publication Date: 2025-05-16XIAN BONNIE XINGZHI INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510503387.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-16
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

Currently, there are problems such as high noise, difficulty in alignment, and poor language adaptability in cross-language information data processing, resulting in low data processing efficiency and accuracy.

Method used

Adaptive learning-based approach is adopted to improve the efficiency and accuracy of cross-language data processing by introducing noise prior learning, covariance matrix prediction network, pseudo-gradient descent update strategy and a three-stage training framework.

Benefits of technology

It improves the efficiency and accuracy of cross-language data processing, especially in low-resource language and small sample scenarios, significantly improving the effects of data cleaning and feature extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011479A_ABST
    Figure CN120011479A_ABST
Patent Text Reader

Abstract

The invention provides a cross-language information data collection and structured processing method based on adaptive learning, and relates to the technical field of artificial intelligence natural language processing, and the method comprises the steps: collecting information data through a cross-language information data source, and carrying out the preprocessing; the preprocessed cross-language text data is input into a self-adaptive learnable semantic model of a neural network framework combining dynamic feature extraction and meta-learning strategies for feature extraction and classification, the semantic model is trained and optimized based on a covariance matrix prediction network, the disturbance direction is dynamically adjusted, and the feature extraction and classification are performed. Enabling the semantic model to adaptively adjust a feature extraction mode according to different contexts in a multi-language environment; and extracting structured information of a cross-language text through a natural language processing technology, and performing cross-language entity alignment and normalization on the feature representation output by the semantic model to generate a unified data structure representation. According to the invention, the cross-language data processing efficiency and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of artificial intelligence natural language processing, and in particular to a method for cross-language information data collection and structured processing based on adaptive learning. Background Art

[0002] In today's information age, multilingual information data is growing exponentially around the world, and enterprises, research institutions, and government departments are increasingly in urgent need of effectively acquiring, processing, and utilizing multilingual information. These multilingual data come from a wide range of sources, covering news websites, social media platforms, academic journals, blogs, and other fields. However, due to the large differences in grammatical structure, expression, and encoding format of each language, the collection, processing, and structuring of cross-language information data face many challenges. Traditional methods usually use rule-based or translation-based methods to align and normalize data, rely on a large amount of annotated data, and often require separate development models for small languages ​​such as Swahili, resulting in high development costs and low efficiency, and poor performance when faced with low-resource languages, semantic variability, and data noise.

[0003] At the same time, cross-language data is usually accompanied by noise such as inconsistent formats, damaged characters, scanning errors, and encoding problems, which seriously affect the data quality and the accuracy of information extraction. In the case of limited privacy budget, how to reasonably allocate the budget to achieve efficient training is also an urgent problem to be solved.

[0004] Therefore, how to improve the efficiency of cross-language data collection, preprocessing, and structured processing through adaptive learning mechanisms has become an important research direction in the current fields of artificial intelligence and natural language processing. Summary of the invention

[0005] In view of the above defects, the present invention aims to propose a cross-language information data collection and structured processing method based on adaptive learning, aiming to solve the problems of high noise, difficult alignment, poor adaptability to low-resource languages, etc. in the current cross-language information data processing. By introducing innovative ideas such as noise prior learning, covariance matrix prediction network, pseudo gradient descent update strategy, and three-stage training framework, the efficiency and accuracy of cross-language data processing can be improved.

[0006] The purpose of the present invention can be achieved through the following technical solutions: The present invention provides a cross-language information data collection and structured processing method based on adaptive learning, comprising: Collect and pre-process information data through cross-language information data sources; The preprocessed cross-language text data is input into an adaptive learnable semantic model of a neural network framework combining dynamic feature extraction and meta-learning strategy for feature extraction and classification, wherein the semantic model is trained and optimized based on a covariance matrix prediction network constructed based on a multi-layer perceptron, and the covariance matrix prediction network takes sample depth features as input and outputs a predicted covariance matrix, and dynamically adjusts the disturbance direction, so that the semantic model can adaptively adjust the feature extraction method according to different contexts in a multi-language environment; The structured information of cross-language texts is extracted through natural language processing technology, and the feature representation output by the semantic model is cross-language entity aligned and normalized to generate a unified data structure representation.

[0007] In some preferred embodiments, the pre-processing comprises: Use synthetic data to generate adversarial networks to generate noisy text data, train feature extractors through contrastive representation learning, and obtain data prior knowledge to filter noise; Gaussian perturbation is applied to the feature vector at the feature level to perform implicit semantic data enhancement and generate enhanced features. The formula is:

[0008] in, is the original eigenvector; is the enhanced feature vector; is a scaling factor; The mean is 0 and the covariance matrix is Gaussian noise.

[0009] In some preferred embodiments, the training and optimization of the semantic model further comprises: Covariance matrix prediction network optimization based on meta-learning strategy and a three-stage training framework are used to improve the classification ability and cross-language processing ability of the model, including: adopting a meta-learning strategy to train the covariance matrix prediction network and classifier through three-stage alternating optimization of pseudo gradient descent update, meta update and real update; Among them, the three-stage training framework includes: stage one: pre-training feature extractor on synthetic data, stage two: training linear classifier with small privacy budget, stage three: end-to-end joint training to reasonably allocate privacy budget and ensure the best optimization effect of the model at different training stages.

[0010] In some preferred embodiments, the meta-learning of the semantic model is optimized by minimizing the implicit semantic data augmentation loss based on the prediction covariance matrix; The parameters of the covariance matrix prediction network are optimized using the cross entropy loss of metadata so that its predicted covariance matrix can accurately model the feature distribution after data enhancement.

[0011] In some preferred embodiments, the covariance matrix prediction network takes sample depth features as input and outputs a predicted covariance matrix, including:

[0012] in, are input features, It is a covariance matrix prediction network based on a multi-layer perceptron. is the predicted covariance matrix.

[0013] In some preferred embodiments, in the three-stage training framework, Pre-training the feature extractor on the synthetic data includes: training the feature extractor using a contrastive loss function to optimize parameters of the feature extractor; Training a linear classifier using a small privacy budget includes: training a linear classifier on features extracted from private data to reduce the impact of gradient noise; End-to-end joint training involves using the remaining budget to perform end-to-end training to adapt the feature extractor and classifier to private data.

[0014] In some preferred embodiments, the implicit semantic data enhancement loss function based on the predicted covariance matrix is:

[0015] in, It is Characteristics of the samples; is the predicted feature; is the number of samples in the batch; is the covariance matrix predicted by the covariance matrix prediction network; is the cross entropy loss:

[0016] in, is the total number of categories; Is a class Labels; is the predicted category probability.

[0017] In some preferred embodiments, the optimization process of the meta-learning strategy includes: Pseudo gradient descent update:

[0018] in, For the Classification parameters during round optimization; To optimize the number of iterations; is the learning rate; Enhance the classification parameters for the current implicit semantic data loss function The gradient of is the classification parameter after pseudo-update; Meta Update:

[0019] in, For the The covariance matrix during round optimization predicts network parameters; Predict the learning rate of the network for the covariance matrix; Predicting network parameters for the covariance matrix using the cross entropy loss function The gradient of Predict network parameters for the updated covariance matrix of the performer; Real update:

[0020] in, The loss function for implicit semantic data enhancement is The gradient of is the final updated classification parameter.

[0021] In some preferred embodiments, the privacy budget allocation rule of the three-stage training framework is: under low privacy budget conditions, stage two allocates a higher budget ratio; as the total privacy budget increases, the budget ratio of stage two is reduced and the budget ratio of stage three is increased.

[0022] In some preferred embodiments, the cross-language entity alignment is achieved through cross-language embedding mapping, mapping entities in different languages ​​to a unified semantic space and normalizing them based on similarity.

[0023] Compared with the prior art, the present invention has the following beneficial effects: The present invention proposes a covariance matrix prediction network for dynamically predicting semantic direction, replacing the traditional statistical estimation method. By learning the semantic relationship of texts in different languages, it can adaptively adjust the feature extraction method and improve the semantic consistency and information alignment ability of cross-language texts, especially in low-resource languages ​​and few-sample scenarios.

[0024] The present invention independently proposes a noise prior learning method, which uses synthetic data for comparative representation learning, so that it has strong data prior knowledge and can identify and filter common noise in cross-language data. Compared with traditional random initialization methods, the present invention can significantly improve the effects of data cleaning and feature extraction under low privacy budget conditions, and improve the reliability and availability of cross-language information data.

[0025] To address the problems of high computational overhead and slow convergence in traditional cross-language model training, the present invention proposes a pseudo-gradient descent update, which splits the classification optimization process into three stages: pseudo-update, meta-update, and real update. This strategy can effectively reduce unnecessary calculations and improve training efficiency, so that the model can achieve better classification results with a smaller privacy budget.

[0026] In the privacy protection scenario, data collection and processing are usually strictly restricted. Therefore, the present invention proposes a three-stage training framework including synthetic data pre-training, small privacy budget linear classifier training and end-to-end training to reasonably allocate the privacy budget and ensure the best optimization effect of the model in different training stages.

[0027] It should be understood that the contents described in the summary of the invention are not intended to limit the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, among which: Figure 1 It is a schematic diagram of the steps of the cross-language information data collection and structured processing method based on adaptive learning provided in an embodiment of the present application; Figure 2 It is a data flow chart of the cross-language information data collection and structured processing method based on adaptive learning provided in an embodiment of the present application; Figure 3 This is a performance comparison curve of the algorithm proposed in this application and the traditional Transformer algorithm; Figure 4 It is a scatter plot of the feature distribution of cross-language data of the algorithm proposed in this application. DETAILED DESCRIPTION

[0029] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.

[0030] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0031] Figure 1 is a schematic diagram of the steps of a cross-language information data collection and structured processing method based on adaptive learning provided by an embodiment of the present disclosure, Figure 2 This is a data flow chart of a cross-language information data collection and structured processing method based on adaptive learning provided by an embodiment of the present disclosure. Figure 1 and Figure 2 As shown, a method 100 for collecting and structuring cross-language information data based on adaptive learning includes: S110: Collecting information data through a cross-language information data source and performing pre-processing; The step S110 specifically includes: S111: Cross-language data collection First, you need to identify cross-language information data sources, including news websites, social media platforms, academic journals, blogs, etc. in multiple languages. Then use crawler technology or application programming interfaces to crawl information data from multiple language environments, supporting multi-language web crawling and application programming interface data acquisition.

[0032] Among them, the hardware devices include servers, data storage devices, processors, network interfaces and other components, which are used to support data capture, processing, storage and push.

[0033] S112: Cross-language data preprocessing The captured raw data is usually disorganized and needs to be cleaned, denoised and formatted, including removing hypertext markup language tags and garbled code processing.

[0034] In order to improve data quality and provide better representation for subsequent modeling, preferably, the present invention also combines noise prior learning and implicit semantic data enhancement to improve the robustness and generalization ability of the data. Specifically: S1121: Noise Prior Learning Use synthetic data to generate adversarial networks to generate noisy text data, train feature extractors through contrastive representation learning, and obtain data prior knowledge to filter noise. More specifically: Use synthetic data to obtain data priors that are beneficial to the task, and this process has no privacy cost, and the use of synthetic data does not involve privacy issues. Use style generative adversarial networks (GANs) to generate text data that is close to the distribution of real data; use shaders to generate complex text backgrounds and interference samples to simulate text noise in the real world (such as optical character recognition scanning errors, text blur, character damage, etc.). Perform contrastive representation learning on these data, train feature extractors, and obtain data priors. Use contrastive learning methods such as simple contrastive learning representations or momentum contrastive learning, and generate positive and negative sample pairs through data enhancement (random character transformation, character spacing transformation, optical character recognition distortion, etc.) to enable the model to learn data priors and improve robustness to noise. When the privacy budget is low, using these priors can effectively improve performance, especially compared to random initialization. The loss function of contrastive learning is contrastive loss , the formula is:

[0035] in: is the number of samples, representing the total number of samples in a batch. Sample pairs , each sample pair contains a positive sample pair or a negative sample pair. For example, if the batch size is 128, then N=128, which means that the batch has 128 sample pairs. The loss function calculates the average loss of the entire batch, so there is a Normalization factor; It is a sample The label of (1 represents a positive sample, 0 represents a negative sample). Value range , express and is a positive sample pair (similar samples should be brought closer); express and is a negative sample pair (dissimilar sample, should be pushed away). It is a sample and The distance between them (such as Euclidean distance, cosine distance, etc.) is used to calculate the similarity measure between two samples. is a hyperparameter (minimum distance threshold) that controls the minimum distance between negative samples (dissimilar samples). If the distance between negative sample pairs is Less than , a loss will be incurred, forcing the model to push away negative sample pairs; if , then the loss of the negative sample pair is 0, indicating that they are sufficiently separated and no further optimization is required.

[0036] The model will eventually form reasonable clusters in the embedding space, so that data of the same type are clustered together and data of different types have clear boundaries. Through the above processing of step S1121, data prior knowledge that is beneficial to the task can be obtained, and common noise in cross-language data can be identified and filtered.

[0037] S1122: Implicit Semantic Data Enhancement This step S1122 is used to apply Gaussian perturbation to the feature vector at the feature level to perform implicit semantic data enhancement and generate enhanced features.

[0038] Traditional data augmentation is usually performed at the input layer (such as character replacement, insertion, deletion, etc.). The perturbation direction of data augmentation is fixed and cannot adapt to semantic changes in different languages ​​or scenarios. Implicit semantic data augmentation, on the other hand, performs augmentation at the feature level, making the semantic information richer and improving the model's adaptability to cross-language data. Data augmentation is performed at the feature level, expanding features by sampling and transforming directions from a Gaussian distribution. By expanding the feature hierarchy through data augmentation, model robustness is improved, especially in fine-grained scenarios.

[0039] Sampling from Gaussian distribution transforms the direction and expands the features. In this process, the spatial features of the data are disturbed by transformation to enhance the generalization ability of the model. For the feature vector, the enhanced features are generated by Gaussian distribution. The formula is:

[0040] in, is the original eigenvector; is the enhanced feature vector; is a scaling factor; The mean is 0 and the covariance matrix is Gaussian noise.

[0041] This enhancement method improves the model's tolerance to input data noise by sampling from a Gaussian distribution to perform feature perturbations.

[0042] Implicit semantic data enhancement plays multiple roles in this invention, including cross-language feature consistency builder, noise robustness enhancer, and low-resource language generalization booster. It directly improves the performance of the model in cross-language classification, entity alignment, and noise scenarios, and seamlessly connects with the covariance matrix prediction network and meta-learning strategy to form a complete technical closed loop. It provides efficient solutions for scenarios such as multilingual data analysis for global enterprises.

[0043] S120: inputting the preprocessed cross-language text data into an adaptive learnable semantic model of a neural network framework combining dynamic feature extraction and meta-learning strategy for feature extraction and classification, wherein the semantic model is trained and optimized based on a covariance matrix prediction network constructed by a multi-layer perceptron, the covariance matrix prediction network takes sample depth features as input and outputs a predicted covariance matrix, and dynamically adjusts the disturbance direction, so that the semantic model can adaptively adjust the feature extraction method according to different contexts in a multi-language environment; The step S120 specifically includes: S121: Covariance matrix prediction network construction For the captured data, first determine the language type of the text through a language recognition model (such as a language recognition tool based on natural language processing) and store it in categories. Design an adaptive learnable semantic model so that in a multilingual environment, the processing strategy can be dynamically adjusted according to the characteristics of different languages, such as grammar, vocabulary, and syntactic structure. The model is based on covariance matrix prediction network training, so that it can adjust the feature extraction method according to different contexts. Pre-train the feature extractor on data in different languages ​​to learn common features across languages. Enhance the model's ability to learn low-resource languages ​​and improve its ability to process low-frequency languages.

[0044] Among them, the adaptive learnable semantic model is a neural network framework that combines dynamic feature extraction and meta-learning strategies. The basic network of the dynamic feature extractor in the model can use a multilingual Transformer encoder (such as XLM-RoBERTa). The purpose and function of this semantic model are: cross-language feature extraction: extract semantic consistency features of multilingual texts to eliminate interference caused by language grammar and vocabulary differences; semantic alignment and classification: map texts in different languages ​​to a unified feature space to achieve cross-language classification (such as sentiment classification, topic classification); dynamic adaptation to low-resource languages: through the covariance matrix prediction network and meta-learning strategy, adaptively adjust the feature extraction method to improve the modeling ability of low-resource languages ​​(such as minority languages). Noise robustness enhancement: combined with implicit semantic data enhancement, improve the model's tolerance to data noise (such as OCR errors, character damage).

[0045] This semantic model can achieve cross-language information classification: automatically identify text categories in multilingual mixed data (such as news classification, spam detection); structured information extraction: support named entity recognition (NER), relationship extraction and other tasks, and generate cross-language knowledge graphs. Low-resource language processing: achieve high-precision semantic analysis in languages ​​that lack labeled data (such as Swahili). Privacy-sensitive scenario applications: securely process multilingual texts under differential privacy constraints (such as medical and financial data).

[0046] The covariance matrix prediction network dynamically predicts the semantic direction according to the deep features of the samples to replace the class conditional covariance matrix in implicit semantic data enhancement, improve the semantic consistency and information alignment capabilities of cross-language texts, and thus improve the quality of semantic direction.

[0047] A covariance matrix prediction network is constructed based on a multi-layer perceptron. The semantic direction of the sample is predicted using deep features as input. The semantic direction of the sample is predicted, replacing the class-conditional covariance matrix estimated statistically in implicit semantic data enhancement, and improving the quality of the semantic direction. The structure of the covariance matrix prediction network includes: input layer: the dimension is 768, which is consistent with the output dimension of the feature extractor; hidden layer: two fully connected layers, the dimension is 512, and the activation function is ReLU; output layer: the dimension is 768×(768+1) / 2, and the lower triangular parameters of the output covariance matrix are output.

[0048] Assume that the input depth feature is , the predicted covariance matrix is , the output of the covariance matrix prediction network is the covariance matrix prediction value:

[0049] in, is the input feature, the covariance matrix predicts the network is a multilayer perceptron model, the output is the predicted covariance matrix The covariance matrix of Gaussian noise is dynamically adjusted through the covariance matrix prediction network , so that the perturbation direction is aligned with the semantic direction.

[0050] The meta-learning of the model is optimized by minimizing the implicit semantic data enhancement loss based on the predicted covariance matrix. The covariance matrix prediction network is optimized using the cross entropy loss of the metadata to avoid the degradation problem when optimizing two networks at the same time and predict the direction of semantic perturbation, thereby improving the quality of feature enhancement. The meta-update process is accelerated by freezing part of the classification network, improving training efficiency without sacrificing model performance. It is used to optimize the feature consistency of predictions during data augmentation and measure the uncertainty of feature distribution through the covariance matrix:

[0051] in, It is Characteristics of the samples; is the predicted feature; is the number of samples in the batch.

[0052] The covariance matrix of the prediction network is predicted through the covariance matrix, which measures the uncertainty of the model prediction and intuitively reflects the distribution of feature changes after data enhancement.

[0053] Cross Entropy Loss Used to optimize the parameters of the covariance matrix prediction network so that its predicted covariance matrix can accurately model the feature distribution after data enhancement:

[0054] in, is the total number of categories; Is a class Labels; is the predicted category probability.

[0055] The implicit semantic data is used to enhance the loss Used to train the covariance matrix prediction network, it can improve the robustness of data augmentation, model the uncertainty of feature changes, make the features of augmented samples more stable, improve the generalization ability of the model under different data transformations, and reduce the risk of overfitting. Cross entropy loss It is used to optimize the parameters of the covariance matrix prediction network. The purpose is to train the covariance matrix prediction network so that the features after data enhancement are more reliable.

[0056] S122: Covariance Matrix Prediction Network Optimization Based on Meta-Learning Strategy The optimization process consists of three steps, including pseudo-updates (temporary updates) of the classification, meta-updates of the covariance matrix prediction network and real updates of the classification, which are performed alternately in each optimization iteration, and the training set and meta-set are ensured to be different when sampling batches.

[0057] The pseudo-update of classification is based on pseudo-gradient descent update, and the formula is:

[0058] in, For the Classification parameters during round optimization; To optimize the number of iterations; is the learning rate; Enhance the classification parameters for the current implicit semantic data loss function The gradient of is the classification parameter after pseudo-update; The meta-update is based on the state after the classification pseudo-update, and the formula is:

[0059] in, For the The covariance matrix during round optimization predicts network parameters; Predict the learning rate of the network for the covariance matrix; Predicting network parameters for the covariance matrix using the cross entropy loss function The gradient of Predict network parameters for the updated covariance matrix of the performer; Real update:

[0060] in, The loss function for implicit semantic data enhancement is The gradient of is the final updated classification parameter.

[0061] The step-by-step update allows the model to learn the enhanced feature distribution while taking into account the uncertainty of the data, thereby improving the generalization ability. This meta-optimization strategy can improve the effectiveness of data enhancement and make the classifier more robust.

[0062] In addition, the system will adjust and optimize the model online based on feedback data from actual applications (such as user needs, changes in hot information, etc.), so that it can dynamically adapt to new language features and information trends. The captured data is updated regularly to ensure the timeliness of information, and the accuracy and efficiency of data collection and processing are continuously improved through the feedback mechanism of the model.

[0063] S123: Three-stage training of covariance matrix prediction network Using a three-stage training framework, stage one pre-trains the feature extractor on synthetic data, and private training is divided into two stages. Stage two uses a small privacy budget to train a linear classifier on private data extracted features. Since the linear layer has fewer parameters, gradient noise can be reduced. Stage three uses the remaining budget for end-to-end training to adapt the feature extractor and classifier to private data. It is crucial to allocate the privacy budget reasonably. The privacy budget allocation rule of the three-stage training framework is: under low privacy budget conditions, stage two allocates a higher budget ratio; as the total privacy budget increases, the budget ratio of stage two is reduced and the budget ratio of stage three is increased. Optionally, in some embodiments, the privacy budget allocation rule satisfies: the privacy budget of stage two accounts for 60% of the total budget, and the noise standard deviation σ=0.1; the privacy budget of stage three accounts for 40%, and the noise standard deviation σ=0.05. In an optional embodiment, during the training process, optimizer: AdamW (learning rate = 1e-4, weight decay = 0.01); batch size: 128; number of training rounds: stage one (50 rounds), stage two (30 rounds), stage three (100 rounds).

[0064] S1231: Synthetic Data Pre-training Pre-train the feature extractor on the synthetic data. Similar to the contrastive representation learning in noise prior learning, use the contrastive loss function in step S1121 To train the feature extractor and optimize the parameters of the feature extractor , We will gradually learn feature representations that can effectively distinguish different categories of data in the feature space:

[0065] in, are the parameters of the feature extractor at the current iteration step (i.e. the parameters before optimization); are the updated parameters of the feature extractor (i.e., optimized parameters). is the learning rate of the feature extractor, which controls the step size of parameter update. For the contrastive loss function, a contrastive learning based loss is defined to bring the representations of similar samples closer and separate samples of different categories. is the contrast loss function About Feature Extractor Parameters gradient.

[0066] Repeat step S1231 stage 1 training process until the feature extractor learns a stable and robust feature representation.

[0067] S1232: Small Privacy Budget Linear Classifier Training Using a small privacy budget to train a linear classifier on private data extraction features, since the linear layer has fewer parameters, it can reduce the impact of differential privacy mechanisms (such as adding gradient noise) and improve the learning ability of the classifier. Assume that the linear classifier is ,in is the weight vector, is the bias. Use cross-pick loss Training is used to measure the prediction results of the classifier and the actual label The gap between:

[0068] is the index of the training sample, indicating the sample number currently being processed. The losses of all samples in the training dataset are accumulated and calculated. is the category index, indicating the categories (if it is a binary classification, then Takes 0 or 1. If it is multi-classification, From 0 to . is the true label, used to represent the sample Belongs to category . The logarithm is taken to calculate the cross-cutting loss, making the optimization more stable (avoiding numerical underflow problems). : Prediction sample Belongs to category The probability of is calculated by the softmax function:

[0069] Where j is the index value; Represents a linear classifier in the category The score on is calculated as follows:

[0070] Yes Category The corresponding weight vector; Yes Category The corresponding bias term.

[0071] Optimize the linear classifier parameters and:

[0072] in, is the learning rate of the linear classifier; : Weight parameter of the current iteration step (before optimization); is the updated weight parameter. is the bias parameter of the current iteration (before optimization); is the updated bias parameter; is the cross-extraction loss function About weight gradient. represents the cross-picking loss function The gradient with respect to bias b.

[0073] Due to the small number of parameters, the impact of gradient noise on model training can be reduced. Differential privacy mechanisms such as gradient clipping and noise injection can be combined to protect data privacy.

[0074] S1233: End-to-end joint training We use the remaining budget to train the feature extractor and classifier end-to-end to adapt them to the private data. We combine the feature extractor and classifier and set the joint parameter to , using the overall loss (Combined with implicit semantic data enhancement loss and classification loss, etc.) for optimization:

[0075] are joint parameters, all model parameters optimized during end-to-end training, including: . are the parameters of the feature extractor, which is used to extract high-dimensional features from the input data. are the parameters of the classifier, which are used to make the final classification based on the features. : Old parameters, which represent the model parameters before optimization, that is, the parameters in the previous training iteration. : New parameters, representing the optimized model parameters, that is, the parameters updated by gradient descent in the current iteration. : The learning rate of joint training controls the learning rate of the optimization step and affects the amplitude of parameter updates. : Overall loss function, the total loss function used in end-to-end training, contains multiple sub-loss items. : Loss gradient, calculate the loss function For model parameters The gradient is:

[0076] This gradient is used to guide parameter updates so that the model is optimized in the direction of minimizing loss.

[0077] Throughout the process, the allocation of the privacy budget affects the training intensity of different stages. For example, when the privacy budget is small, more can be allocated to stage 2. As the privacy budget increases, the proportion of stage 2 is reduced.

[0078] The processing flow of step S120 specifically includes: Language identification and dynamic adaptation: Determine the text language type (such as Chinese, English, etc.) through the language identification module and generate a language type embedding vector; dynamically adjust the parameters of the feature extractor according to the language type (such as inserting a language adaptation layer).

[0079] Feature extraction and covariance prediction: Use multilingual Transformer (such as XLM-R) to extract basic features. Covariance matrix prediction network (CovNet) predicts the perturbation direction (covariance matrix of Gaussian noise) based on features.

[0080] Implicit semantic data augmentation: Perturbations are applied to features according to the predicted covariance matrix to generate enhanced features.

[0081] Classification and Optimization: The classifier outputs category probabilities based on the enhanced features. The classifier and CovNet parameters are optimized in stages through a meta-learning strategy.

[0082] S130: extracting structured information of cross-language texts through natural language processing technology, and performing cross-language entity alignment and normalization on the feature representation output by the semantic model to generate a unified data structure representation.

[0083] In step S130, structured processing and information extraction are performed. Cross - language entity alignment is achieved through cross - language embedding mapping, which maps entities in different languages to a unified semantic space and normalizes them based on similarity (such as using cosine similarity). Specifically, structured information such as entities like time, location, event, and person is extracted from information texts in different languages through natural language processing techniques (such as named entity recognition, relation extraction, event extraction, etc.). Through a cross - language mapping mechanism (such as an alignment model or cross - language embedding), for example, using the pre - trained LaBSE model to map entities to a unified semantic space to achieve cross - language embedding mapping, the same or similar entities in different languages are normalized to ensure the effective docking of information in multilingual texts. The normalization rule is that if an entity has multiple names in multiple languages (such as "Paris" and "巴黎"), the English name is used as the standard entity ID. The structured information data is fused to generate a unified data structure representation (such as a knowledge graph, relational database, etc.) and stored for subsequent analysis and query.

[0084] According to the above - mentioned embodiments of the present invention, beneficial data prior knowledge for the task is obtained through noise prior learning to identify and filter common noises in cross - language data; feature extraction and classification are performed through an adaptive learnable semantic model of a neural network framework that combines dynamic feature extraction and meta - learning strategies, and the semantic direction is dynamically predicted through a covariance matrix prediction network to improve the semantic consistency and information alignment ability of cross - language texts.

[0085] Result comparison and effect verification: Figure 3 This is a curve graph for algorithm performance comparison. To distinguish the proposed algorithm from the traditional Transformer algorithm, the curve of the proposed algorithm is drawn as a solid line with dots, and the traditional Transformer algorithm is drawn as a dashed line with squares. Among them Figure 3 This is for algorithm loss convergence comparison. As Figure 3 shown, for the proposed algorithm, at the 5th round, the loss drops to about 0.2, and at the 10th round, it approaches 0.05, with overall fast convergence. For the Transformer algorithm, the loss is still 0.35 at the 5th round and only drops to 0.1 at the 10th round, with slow convergence. The results show that the proposed algorithm can quickly reduce the loss within fewer training rounds and improve the optimization efficiency.

[0086] Table 1 is the comparison of cross - language classification accuracy (XGLUE benchmark test), and Table 2 is the noise robustness test (including 20% OCR noise data). As shown in Table 1 and Table 2, the algorithm proposed in the present invention shows obvious advantages compared with the traditional Transformer in both cross - language classification accuracy tests and noise robustness tests.

[0087] Table 1: Cross-language classification accuracy comparison (XGLUE benchmark)

[0088] Table 2: Noise robustness test (including 20% ​​OCR noise data)

[0089] Furthermore, Figure 4 A scatter plot depicting the feature distribution of cross-language data, simulating the distribution of English (blue), Chinese (red), and French (green) in the feature space.

[0090] The mean of the feature points of English data (blue) is (2.05, 2.10), the standard deviation is about (0.45, 0.48), and the data is mainly distributed in the range of (1.5, 2.5). The feature distribution is compact, indicating that the English features extracted by the algorithm are stable. The mean of the feature points of Chinese (red) data is (-2.00, 2.05), the standard deviation is about (0.50, 0.47), and the data is mainly distributed in the range of (-2.5, -1.5), indicating that the model has a good normalization ability for Chinese features. The distribution of Chinese data is more scattered, and there are some points that are off-center, which may indicate that the diversity of Chinese text features is higher. The mean of the feature points of French data (green) is (0.05, -1.95), the standard deviation is about (0.48, 0.49), and the data is mainly distributed in the range of (-0.5, 0.5), indicating that the feature extraction of French is also relatively stable, but slightly extended in the negative direction. These feature vectors show that the proposed algorithm can distinguish data of different languages ​​well, making the feature distribution of the same language more compact. Data of different languages ​​still maintain a certain distribution relationship, which means that the proposed algorithm can better capture the feature similarity across languages. The average intra-group distance is the Euclidean distance from the language data point to its central mean. The average intra-group distance of English data is 0.52, the average intra-group distance of Chinese data is 0.55, and the average intra-group distance of French data is 0.51. The small intra-group distance shows that the clustering effect of feature points of the same language is good, and the proposed algorithm can distinguish data of different languages ​​well. The inter-group distance is the Euclidean distance between the center points of different language categories. The inter-group distance of English-Chinese is 4.08, the inter-group distance of English-French is 4.05, and the inter-group distance of Chinese-French is 4.12. The large inter-group distance shows that the data points of different languages ​​are obviously separated, the cross-language feature discrimination is high, and information confusion is reduced.

[0091] The above analysis shows that the proposed algorithm has obvious advantages in tasks such as cross-language information extraction, machine translation, and knowledge graph construction.

[0092] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0093] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A cross-language information data collection and structured processing method based on adaptive learning, characterized in that: include: Collect and pre-process information data through cross-language information data sources; The preprocessed cross-language text data is input into an adaptive learnable semantic model of a neural network framework combining dynamic feature extraction and meta-learning strategy for feature extraction and classification, wherein the semantic model is trained and optimized based on a covariance matrix prediction network constructed based on a multi-layer perceptron, and the covariance matrix prediction network takes sample depth features as input and outputs a predicted covariance matrix, and dynamically adjusts the disturbance direction, so that the semantic model can adaptively adjust the feature extraction method according to different contexts in a multi-language environment; The structured information of cross-language texts is extracted through natural language processing technology, and the feature representation output by the semantic model is cross-language entity aligned and normalized to generate a unified data structure representation.

2. The cross-language information data collection and structured processing method based on adaptive learning according to claim 1, characterized in that: in, Preprocessing includes: Use synthetic data to generate adversarial networks to generate noisy text data, train feature extractors through contrastive representation learning, and obtain data prior knowledge to filter noise; Gaussian perturbation is applied to the feature vector at the feature level to perform implicit semantic data enhancement and generate enhanced features. The formula is: , in, is the original eigenvector; is the enhanced feature vector; is a scaling factor; The mean is 0 and the covariance matrix is Gaussian noise.

3. The cross-language information data collection and structured processing method based on adaptive learning according to claim 2, characterized in that: in, The training and optimization of the semantic model also includes: Covariance matrix prediction network optimization based on meta-learning strategy and a three-stage training framework are used to improve the classification ability and cross-language processing ability of the model, including: adopting a meta-learning strategy to train the covariance matrix prediction network and classifier through three-stage alternating optimization of pseudo gradient descent update, meta update and real update; Among them, the three-stage training framework includes: stage one: pre-training feature extractor on synthetic data, stage two: training linear classifier with small privacy budget, stage three: end-to-end joint training to reasonably allocate privacy budget and ensure the best optimization effect of the model at different training stages.

4. The cross-language information data collection and structured processing method based on adaptive learning according to claim 3, characterized in that: in, The meta-learning of the semantic model is optimized by minimizing the implicit semantic data augmentation loss based on the prediction covariance matrix; The parameters of the covariance matrix prediction network are optimized using the cross entropy loss of metadata so that its predicted covariance matrix can accurately model the feature distribution after data enhancement.

5. The cross-language information data collection and structured processing method based on adaptive learning according to claim 4, characterized in that: The covariance matrix prediction network takes sample depth features as input and outputs a predicted covariance matrix, including: , in, are input features, It is a covariance matrix prediction network based on a multi-layer perceptron. is the predicted covariance matrix.

6. The cross-language information data collection and structured processing method based on adaptive learning according to claim 5, characterized in that: in, In the three-stage training framework, Pre-training the feature extractor on the synthetic data includes: training the feature extractor using a contrastive loss function to optimize parameters of the feature extractor; Training a linear classifier using a small privacy budget includes: training a linear classifier on features extracted from private data to reduce the impact of gradient noise; End-to-end joint training involves using the remaining budget to perform end-to-end training to adapt the feature extractor and classifier to private data.

7. The cross-language information data collection and structured processing method based on adaptive learning according to claim 4, characterized in that: The implicit semantic data enhancement loss function based on the predicted covariance matrix is: , in, It is Characteristics of the samples; is the predicted feature; is the number of samples in the batch; is the covariance matrix predicted by the covariance matrix prediction network; is the cross entropy loss: , in, is the total number of categories; Is a class Labels; is the predicted category probability.

8. The cross-language information data collection and structured processing method based on adaptive learning according to claim 7, characterized in that: in, The optimization process of the meta-learning strategy includes: Pseudo gradient descent update: , in, For the Classification parameters during round optimization; To optimize the number of iterations; is the learning rate; Enhance the classification parameters for the current implicit semantic data loss function The gradient of is the classification parameter after pseudo-update; Meta Update: , in, For the The covariance matrix during round optimization predicts network parameters; Predict the learning rate of the network for the covariance matrix; Predicting network parameters for the covariance matrix using the cross entropy loss function The gradient of Predict network parameters for the updated covariance matrix of the performer; Real update: , in, The loss function for implicit semantic data enhancement is The gradient of is the final updated classification parameter.

9. The cross-language information data collection and structured processing method based on adaptive learning according to claim 6, characterized in that: in, The privacy budget allocation rule of the three-stage training framework is: under low privacy budget conditions, stage two allocates a higher budget ratio; as the total privacy budget increases, the budget ratio of stage two is reduced and the budget ratio of stage three is increased.

10. The cross-language information data collection and structured processing method based on adaptive learning according to claim 9, characterized in that: in, The cross-language entity alignment is achieved through cross-language embedding mapping, which maps entities in different languages ​​to a unified semantic space and normalizes them based on similarity.

Citation Information

Patent Citations

  • Cross-language abstract abstract method based on robust self-learning strategy in low-resource scene

    CN117271761A

  • Mongolian-Chinese cross-language sentiment analysis method based on knowledge sharing and adaptive learning

    CN117332081A

  • Denoising semantic communication system optimization method based on semantic compression

    CN119210652A

  • Artificial Intelligence-Based Cross-Language Speech Transcription Method and Apparatus, Device and Readable Medium

    US20180336900A1

Cited By

  • Prompt injection attack detection method and system based on low-resource language

    CN121435238A

  • Method and system for low-resource language-based prompt injection attack detection

    CN121435238B