Cross-language Information Data Collection and Structured Processing Method Based on Adaptive Learning

Through the adaptive learning mechanism, noise prior learning and covariance matrix prediction network are used to solve the problem of noise and alignment difficulties in cross-language information data processing, improve the adaptability and processing efficiency of low-resource languages, and realize efficient and accurate data acquisition and structure under limited privacy budgets.

CN120011479BActive Publication Date: 2025-07-04XIAN BONNIE XINGZHI INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510503387.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-04
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

The collection, processing and structure of cross-language information data faces the problems of high noise, difficulty in alignment, and poor adaptability of low-resource languages, especially in the case of limited privacy budgets.

Method used

Adaptive learning mechanism is adopted to collect and structure the cross-language information data through noise prior learning, covariance matrix prediction network, pseudo-gradient descent update strategy and three-stage training framework, and combine the neural network framework of dynamic feature extraction and meta-learning strategy.

Benefits of technology

It improves the efficiency and accuracy of cross-language data processing, especially in low-resource language and noise scenarios, enhances the semantic consistency and information alignment capabilities of the model, and reduces the limitations of computing overhead and privacy budget.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011479B_ABST
    Figure CN120011479B_ABST
Patent Text Reader

Abstract

This application provides a cross - language information data collection and structured processing method based on adaptive learning, which relates to the technical field of artificial intelligence natural language processing. The method includes: collecting information data through a cross - language information data source and performing pre - processing; inputting the pre - processed cross - language text data into an adaptive learnable semantic model of a neural network framework that combines dynamic feature extraction and meta - learning strategies for feature extraction and classification. Among them, the semantic model is trained and optimized based on a covariance matrix prediction network, dynamically adjusting the perturbation direction so that the semantic model can adaptively adjust the feature extraction method according to different contexts in a multilingual environment; extracting structured information from the cross - language text through natural language processing techniques, and performing cross - language entity alignment and normalization on the feature representation output by the semantic model to generate a unified data structure representation. The present invention improves the efficiency and accuracy of cross - language data processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of artificial intelligence natural language processing, and particularly to a cross-language information data collection and structured processing method based on adaptive learning. Background Art

[0002] In today's information age, multi-language information data globally has grown exponentially. Enterprises, research institutions, government departments, etc. have an increasingly urgent need to effectively acquire, process, and utilize multi-language information. These multi-language data sources are extensive, covering multiple fields such as news websites, social media platforms, academic journals, blogs, etc. However, due to significant differences in grammar structures, expression methods, and encoding formats among languages, the collection, processing, and structuring of cross-language information data face many challenges. Traditional methods usually use rule-based or translation-based means for data alignment and normalization, relying on a large amount of labeled data. For small languages such as Swahili, models often need to be developed separately, resulting in high development costs and low efficiency, and poor performance in the face of low-resource languages, semantic variability, and data noise.

[0003] At the same time, cross-language data is usually accompanied by noise such as inconsistent formats, character corruption, scanning errors, encoding problems, etc., seriously affecting data quality and the accuracy of information extraction. In the case of limited privacy budgets, how to reasonably allocate budgets to achieve efficient training is also an urgent problem to be solved.

[0004] Therefore, how to improve the efficiency of cross-language data collection, preprocessing, and structured processing through an adaptive learning mechanism has become an important research direction in the current fields of artificial intelligence and natural language processing. Summary of the Invention

[0005] Aiming at the above defects, the purpose of the present invention is to propose a cross-language information data collection and structured processing method based on adaptive learning, aiming to solve problems such as a large amount of noise, difficult alignment, and poor adaptability to low-resource languages in current cross-language information data processing. By introducing innovative ideas such as noise prior learning, covariance matrix prediction network, pseudo-gradient descent update strategy, and three-stage training framework, the efficiency and accuracy of cross-language data processing are improved.

[0006] The purpose of the present invention can be achieved through the following technical solutions:

[0007] The present invention provides a cross-language information data collection and structured processing method based on adaptive learning, including:

[0008] Collecting information data through cross-language information data sources and performing preprocessing;

[0009] Input the preprocessed cross - language text data into the adaptive learnable semantic model of the neural network framework that combines dynamic feature extraction and meta - learning strategies for feature extraction and classification. Among them, the semantic model is trained and optimized based on the covariance matrix prediction network constructed by a multi - layer perceptron. The covariance matrix prediction network takes the sample depth features as input and outputs the predicted covariance matrix, dynamically adjusting the perturbation direction so that the semantic model can adaptively adjust the feature extraction method according to different contexts in a multi - language environment;

[0010] Extract the structured information of the cross - language text through natural language processing techniques, and perform cross - language entity alignment and normalization on the feature representation output by the semantic model to generate a unified data structure representation.

[0011] In some preferred embodiments, the preprocessing includes:

[0012] Use a synthetic data generation adversarial network to generate noisy text data, train a feature extractor through contrastive representation learning, and obtain prior knowledge of the data to filter out noise;

[0013] Apply Gaussian perturbation to the feature vector at the feature level for implicit semantic data augmentation to generate enhanced features. The formula is:

[0014]

[0015] where, is the original feature vector; is the enhanced feature vector; is a scaling factor; is Gaussian noise with a mean of 0 and a covariance matrix of ;

[0016] In some preferred embodiments, the training and optimization of the semantic model further include:

[0017] Optimization of the covariance matrix prediction network based on the meta - learning strategy and the adoption of a three - stage training framework for enhancing the classification ability and cross - language processing ability of the model, including: adopting the meta - learning strategy and training the covariance matrix prediction network and the classifier through three - stage alternating optimization of pseudo - gradient descent update, meta - update, and real - update;

[0018] Among them, the three - stage training framework includes: Stage 1: Pre - train the feature extractor on synthetic data, Stage 2: Train a linear classifier with a small privacy budget, and Stage 3: End - to - end joint training to reasonably allocate the privacy budget and ensure the best optimization effect of the model in different training stages.

[0019] In some preferred embodiments, the meta - learning of the semantic model is optimized by minimizing the implicit semantic data augmentation loss based on the predicted covariance matrix;

[0020] The parameters of the covariance matrix prediction network are optimized using the cross - entropy loss of the metadata, so that the predicted covariance matrix can accurately model the feature distribution after data augmentation.

[0021] In some preferred embodiments, the covariance matrix prediction network takes the sample depth features as input and outputs the predicted covariance matrix, including:

[0022]

[0023] where, is the input feature, is the covariance matrix prediction network based on the multi - layer perceptron, is the predicted covariance matrix.

[0024] In some preferred embodiments, in the three - stage training framework,

[0025] Pre - training the feature extractor on synthetic data includes: training the feature extractor using the contrast loss function to optimize the parameters of the feature extractor;

[0026] Training the linear classifier with a small privacy budget includes: training the linear classifier on the features extracted from private data to reduce the impact of gradient noise;

[0027] End - to - end joint training includes: performing end - to - end training with the remaining budget to make the feature extractor and the classifier adapt to the private data.

[0028] In some preferred embodiments, the implicit semantic data augmentation loss function based on the predicted covariance matrix is:

[0029]

[0030] where, is the feature of the th sample; is the predicted feature; is the number of samples within the batch; is the covariance matrix predicted by the covariance matrix prediction network; is the cross - entropy loss:

[0031]

[0032] where, is the total number of classes; is the label of the class; is the predicted class probability.

[0033] In some preferred embodiments, the optimization process of the meta - learning strategy includes:

[0034] Pseudo - gradient descent update:

[0035]

[0036] wherein, is the classification parameter at the round of optimization; is the number of steps of the optimization iteration; is the learning rate; is the gradient of the current implicit semantic data augmentation loss function with respect to the classification parameter ; is the classification parameter after performing the pseudo - update;

[0037] Meta - update:

[0038]

[0039] wherein, is the covariance matrix prediction network parameter at the round of optimization; is the learning rate of the covariance matrix prediction network; is the gradient of the cross - entropy loss function with respect to the covariance matrix prediction network parameter ; is the covariance matrix prediction network parameter after performing the meta - update;

[0040] True update:

[0041]

[0042] wherein, is the gradient of the implicit semantic data augmentation loss function with respect to ; is the classification parameter after the final update.

[0043] In some preferred embodiments, the privacy budget allocation rule of the three - stage training framework is: under the condition of low privacy budget, a higher budget proportion is allocated in stage two; as the total privacy budget increases, the budget proportion of stage two is reduced and the budget proportion of stage three is increased.

[0044] In some preferred embodiments, the cross - language entity alignment is achieved through cross - language embedding mapping, mapping entities in different languages to a unified semantic space and normalizing based on similarity.

[0045] Compared with the prior art, the present invention has the following beneficial effects:

[0046] The present invention proposes a covariance matrix prediction network for dynamically predicting semantic directions, replacing traditional statistical estimation methods. By learning the semantic relationships of different language texts, it can adaptively adjust the feature extraction method, improve the semantic consistency and information alignment ability of cross-language texts, especially outstanding in low-resource languages and few-shot scenarios.

[0047] The present invention independently proposes a noise prior learning method, which uses synthetic data for contrastive representation learning to enable it to have strong data prior knowledge and can identify and filter common noises in cross-language data. Compared with traditional random initialization methods, the present invention can significantly improve the effects of data cleaning and feature extraction under low privacy budget conditions, and improve the reliability and availability of cross-language information data.

[0048] Aiming at the problems of large computational overhead and slow convergence speed in traditional cross-language model training, the present invention proposes pseudo-gradient descent update, which splits the classification optimization process into three stages: pseudo-update, meta-update, and real update. This strategy can effectively reduce unnecessary computations and improve training efficiency, enabling the model to achieve better classification effects with less privacy budget.

[0049] In privacy protection scenarios, the collection and processing of data are usually strictly restricted. Therefore, the present invention proposes a three-stage training framework including synthetic data pre-training, small privacy budget linear classifier training, and end-to-end training, which reasonably allocates privacy budget to ensure the best optimization effects of the model in different training stages.

[0050] It should be understood that the content described in the Summary of the Invention section is not intended to limit the key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] With reference to the accompanying drawings and the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. In the drawings, the same or similar reference numerals represent the same or similar elements, where:

[0052] Figure 1 is a schematic flow chart of the steps of the cross-language information data collection and structured processing method based on adaptive learning provided by an embodiment of the present application;

[0053] Figure 2 is a data flow chart of the cross-language information data collection and structured processing method based on adaptive learning provided by an embodiment of the present application;

[0054] Figure 3 It is a performance comparison curve graph of the algorithm proposed in this application and the traditional Transformer algorithm;

[0055] Figure 4 It is a scatter plot of the feature distribution of cross - language data of the algorithm proposed in this application. Detailed implementation manners

[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.

[0057] To make the above - mentioned objectives, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners.

[0058] Figure 1 It is a schematic flowchart of the steps of a cross - language information data acquisition and structured processing method based on adaptive learning provided by an embodiment of the present disclosure, Figure 2 It is a data flow diagram of a cross - language information data acquisition and structured processing method based on adaptive learning provided by an embodiment of the present disclosure. Refer to Figure 1 and Figure 2 As shown in, a cross - language information data acquisition and structured processing method 100 based on adaptive learning includes:

[0059] S110: Collect information data through cross - language information data sources and perform pre - processing;

[0060] This step S110 specifically includes:

[0061] S111: Cross - language data collection

[0062] First, it is necessary to determine cross - language information data sources, including news websites, social media platforms, academic journals, blogs, etc. in multiple languages. Then, web crawler technology or application programming interfaces are used to scrape information data from multiple language environments, supporting multi - language web scraping and application programming interface data acquisition.

[0063] Among them, the hardware device includes components such as a server, a data storage device, a processor, and a network interface, which are used to support data scraping, processing, storage, and pushing.

[0064] S112: Cross - language data pre - processing

[0065] The original data captured is usually chaotic and needs to be cleaned, denoised, and formatted, including removing hypertext markup language tags, handling garbled characters, etc.

[0066] To improve data quality and provide a better representation for subsequent modeling, preferably, the present invention also combines noise prior learning and implicit semantic data augmentation to improve the robustness and generalization ability of the data. Specifically:

[0067] S1121: Noise prior learning

[0068] Use a synthetic data generation adversarial network to generate text data with noise, and train a feature extractor through contrastive representation learning to obtain prior knowledge of the data to filter noise. More specifically: Use synthetic data to obtain prior knowledge of the data beneficial to the task, and this process has no privacy cost, and the use of synthetic data does not involve privacy issues. Use a style generation adversarial network (GAN) to generate text data similar to the real data distribution; use a shader to generate complex text backgrounds and interference samples to simulate text noise in the real world (such as optical character recognition scanning errors, text blurring, character damage, etc.). Perform contrastive representation learning on these data, train the feature extractor, and obtain the data prior. Use contrastive learning methods such as simple contrastive learning representation or momentum contrastive learning to generate positive and negative sample pairs through data augmentation (random character transformation, character spacing transformation, optical character recognition distortion, etc.), so that the model learns the data prior and improves the robustness to noise. In the case of low privacy budget, using these priors can effectively improve performance, especially compared to the case of random initialization. The loss function of contrastive learning is the contrastive loss , and its formula is:

[0069]

[0070] Where: is the number of samples, representing that there are a total of sample pairs in a batch , and each sample pair contains a positive sample pair or a negative sample pair. For example, if the batch size is 128, then N = 128, indicating that there are 128 sample pairs in this batch. The loss function calculates the average loss of the entire batch, so there is a normalization factor in front; is the label of the sample . The value range , indicates and are positive sample pairs (similar samples, which should be pulled closer); indicates and is a negative sample pair (dissimilar samples, which should be pushed apart). is a sample and is the distance between (e.g., Euclidean distance, cosine distance, etc.), used to calculate the similarity measure between two samples. is a hyperparameter (minimum distance threshold), controlling the minimum distance between negative samples (dissimilar samples). If the distance of the negative sample pair is less than , a loss will be generated, forcing the model to push the negative sample pair apart; if , the loss of this negative sample pair is 0, indicating that they are already separated enough and do not need further optimization.

[0071] The model will finally form reasonable clusters in the embedding space, making data of the same class gather together and data of different classes have clear boundaries. Through the above processing process of step S1121, prior knowledge of data beneficial to the task can be obtained, and common noises in cross-language data can be identified and filtered.

[0072] S1122: Implicit Semantic Data Augmentation

[0073] This step S1122 is used to perform implicit semantic data augmentation by applying Gaussian perturbation to the feature vector at the feature level to generate enhanced features.

[0074] Traditional data augmentation is usually performed at the input layer (such as character replacement, insertion, deletion, etc.), and the perturbation direction of data augmentation is fixed, unable to adapt to semantic changes in different languages or scenarios. While implicit semantic data augmentation is performed at the feature level, making the semantic information richer and improving the model's adaptability to cross-language data. Data augmentation at the feature level expands the features by sampling the transformation direction from the Gaussian distribution. By expanding the feature hierarchy through data augmentation, especially in fine-grained scenarios, the model's robustness is improved.

[0075] Sample the transformation direction from the Gaussian distribution to expand the features. During this process, the spatial features of the data are perturbed through the transformation to enhance the model's generalization ability. For the feature vector, it is enhanced through the Gaussian distribution to generate the enhanced feature, and the formula is:

[0076]

[0077] where is the original feature vector; is the enhanced feature vector; is a scaling factor; is Gaussian noise with a mean of 0 and a covariance matrix of .

[0078] This enhancement method perturbs features by sampling from a Gaussian distribution, improving the model's tolerance to input data noise.

[0079] Implicit semantic data augmentation plays multiple roles in the present invention, including being a cross - language feature consistency builder, a noise robustness enhancer, and a low - resource language generalization booster. It directly improves the model's performance in cross - language classification, entity alignment, and noise scenarios, seamlessly connects with the covariance matrix prediction network and meta - learning strategies to form a complete technical closed - loop, providing an efficient solution for scenarios such as multilingual data analysis for global enterprises.

[0080] S120: Input the pre - processed cross - language text data into the adaptive learnable semantic model of the neural network framework that combines dynamic feature extraction and meta - learning strategies for feature extraction and classification. Among them, the semantic model is trained and optimized based on the covariance matrix prediction network constructed by a multi - layer perceptron. The covariance matrix prediction network takes the sample depth features as input and outputs the predicted covariance matrix, dynamically adjusting the perturbation direction so that the semantic model can adaptively adjust the feature extraction method according to different contexts in a multilingual environment.

[0081] This step S120 specifically includes:

[0082] S121: Construction of the covariance matrix prediction network

[0083] For the captured data, first determine the language type of the text through a language recognition model (such as a language recognition tool based on natural language processing) and classify and store it. Design an adaptive learnable semantic model that can dynamically adjust the processing strategy according to the characteristics of different languages in a multilingual environment, such as grammar, vocabulary, syntactic structure, etc. The model is trained based on the covariance matrix prediction network so that it can adjust the feature extraction method according to different contexts. Pre - train the feature extractor on data in different languages to learn cross - language common features. Enhance the model's learning ability for low - resource languages and improve its ability to process low - frequency languages.

[0084] Among them, the adaptive learnable semantic model is a neural network framework that combines dynamic feature extraction and meta-learning strategies. The basic network of the dynamic feature extractor in the model can adopt a multilingual Transformer encoder (such as XLM-RoBERTa). The purpose and functions of this semantic model are as follows: Cross-lingual feature extraction: Extract semantic consistency features of multilingual texts and eliminate interference caused by language grammar and vocabulary differences; Semantic alignment and classification: Map texts in different languages to a unified feature space to achieve cross-lingual classification (such as sentiment classification, topic classification); Dynamically adapt to low-resource languages: Through the covariance matrix prediction network and meta-learning strategies, adaptively adjust the feature extraction method to improve the modeling ability for low-resource languages (such as minority languages); Enhance noise robustness: Combine implicit semantic data augmentation to improve the model's tolerance to data noise (such as OCR errors, character corruption).

[0085] This semantic model can achieve cross-lingual information classification: Automatically identify text categories in multilingual mixed data (such as news classification, spam detection); Structured information extraction: Support tasks such as named entity recognition (NER), relation extraction, etc., and generate cross-lingual knowledge graphs; Low-resource language processing: Achieve high-precision semantic analysis in languages lacking labeled data (such as Swahili); Application in privacy-sensitive scenarios: Securely process multilingual texts under differential privacy constraints (such as medical, financial data).

[0086] The covariance matrix prediction network dynamically predicts the semantic direction based on the deep features of the samples to replace the class-conditional covariance matrix in the implicit semantic data augmentation, improve the semantic consistency and information alignment ability of cross-lingual texts, and thus enhance the quality of the semantic direction.

[0087] The covariance matrix prediction network is constructed based on a multi-layer perceptron. Taking the deep features as input, it predicts the semantic direction of the samples, replacing the statistically estimated class-conditional covariance matrix in the implicit semantic data augmentation to improve the quality of the semantic direction. Among them, the structure of the covariance matrix prediction network includes: Input layer: with a dimension of 768, consistent with the output dimension of the feature extractor; Hidden layer: two fully connected layers with a dimension of 512, and the activation function is ReLU; Output layer: with a dimension of 768×(768 + 1) / 2, outputting the lower triangular parameters of the covariance matrix.

[0088] Assume the input deep feature is , and the predicted covariance matrix is . The output of the covariance matrix prediction network is the covariance matrix prediction value:

[0089]

[0090] Among them, is the input feature, and the covariance matrix prediction network is a multi - layer perceptron model, and the output is the predicted covariance matrix . The covariance matrix prediction network dynamically adjusts the covariance matrix of Gaussian noise to align the perturbation direction with the semantic direction.

[0091] The meta - learning of the model is optimized by minimizing the implicit semantic data augmentation loss based on the predicted covariance matrix, while the covariance matrix prediction network is optimized using the cross - entropy loss of the meta - data, avoiding the degradation problem when optimizing two networks simultaneously, predicting the semantic perturbation direction, and thus improving the quality of feature augmentation. The meta - update process is accelerated by freezing part of the classification network, improving the training efficiency without sacrificing the model performance. Among them, the implicit semantic data augmentation loss is used to optimize the predicted feature consistency during the data augmentation process and measure the uncertainty of the feature distribution through the covariance matrix:

[0092]

[0093] where, is the feature of the -th sample; is the predicted feature; is the number of samples within a batch.

[0094] is the covariance matrix passed through the covariance matrix prediction network, measuring the uncertainty of the model prediction and intuitively reflecting the distribution of feature changes after data augmentation.

[0095] The cross - entropy loss is used to optimize the parameters of the covariance matrix prediction network so that the predicted covariance matrix can accurately model the feature distribution after data augmentation:

[0096]

[0097] where, is the total number of classes; is the label of the class; is the predicted class probability.

[0098] Using the implicit semantic data augmentation loss to train the covariance matrix prediction network can improve the robustness of data augmentation, model the uncertainty of feature changes, make the features of augmented samples more stable, improve the generalization ability of the model under different data transformations, and reduce the risk of overfitting at the same time. The cross - entropy loss is used to optimize the parameters of the covariance matrix prediction network, aiming to train the covariance matrix prediction network so that the features after data augmentation are more reliable.

[0099] S122: Optimization of Covariance Matrix Prediction Network Based on Meta-Learning Strategy

[0100] The optimization process consists of three steps, including pseudo-update of classification (temporary update), meta-update of the covariance matrix prediction network, and true update of classification, which are alternated in each optimization iteration, and the training set and meta-dataset are ensured to be different during batch sampling.

[0101] The pseudo-update of classification is based on pseudo-gradient descent update, and the formula is:[[]]

[0102]

[0103] where is the classification parameter at the round of optimization; is the number of steps of the optimization iteration; is the learning rate; is the gradient of the current implicit semantic data augmentation loss function with respect to the classification parameter ; is the classification parameter after performing the pseudo-update;

[0104] The meta-update is based on the state after the classification pseudo-update, and the formula is:[[]]

[0105]

[0106] where is the covariance matrix prediction network parameter at the round of optimization; is the learning rate of the covariance matrix prediction network; is the gradient of the cross-entropy loss function with respect to the covariance matrix prediction network parameter ; is the covariance matrix prediction network parameter after performing the meta-update;

[0107] True update:

[0108]

[0109] where is the gradient of the implicit semantic data augmentation loss function with respect to ; is the final updated classification parameter.

[0110] The step-by-step update enables the model to learn both the enhanced feature distribution and consider the uncertainty of the data, thereby improving the generalization ability. This meta-optimization strategy can improve the effectiveness of data augmentation and make the classifier more robust.

[0111] In addition, the system will adjust and optimize the model online based on feedback data from actual applications (such as user needs, changes in hot information, etc.), so that it can dynamically adapt to new language features and information trends. The captured data is updated regularly to ensure the timeliness of information, and the accuracy and efficiency of data collection and processing are continuously improved through the feedback mechanism of the model.

[0112] S123: Three-stage training of covariance matrix prediction network

[0113] Using a three-stage training framework, stage one pre-trains the feature extractor on synthetic data, and private training is divided into two stages. Stage two uses a small privacy budget to train a linear classifier on private data extracted features. Since the linear layer has fewer parameters, gradient noise can be reduced. Stage three uses the remaining budget for end-to-end training to adapt the feature extractor and classifier to private data. It is crucial to allocate the privacy budget reasonably. The privacy budget allocation rule of the three-stage training framework is: under low privacy budget conditions, stage two allocates a higher budget ratio; as the total privacy budget increases, the budget ratio of stage two is reduced and the budget ratio of stage three is increased. Optionally, in some embodiments, the privacy budget allocation rule satisfies: the privacy budget of stage two accounts for 60% of the total budget, and the noise standard deviation σ=0.1; the privacy budget of stage three accounts for 40%, and the noise standard deviation σ=0.05. In an optional embodiment, during the training process, optimizer: AdamW (learning rate = 1e-4, weight decay = 0.01); batch size: 128; number of training rounds: stage one (50 rounds), stage two (30 rounds), stage three (100 rounds).

[0114] S1231: Synthetic Data Pre-training

[0115] Pre-train the feature extractor on the synthetic data. Similar to the contrastive representation learning in noise prior learning, use the contrastive loss function in step S1121 To train the feature extractor and optimize the parameters of the feature extractor , We will gradually learn feature representations that can effectively distinguish different categories of data in the feature space:

[0116]

[0117] in, are the parameters of the feature extractor at the current iteration step (i.e. the parameters before optimization); are the updated parameters of the feature extractor (i.e., optimized parameters). is the learning rate of the feature extractor, which controls the step size of parameter update. For the contrastive loss function, a contrastive learning based loss is defined to bring the representations of similar samples closer and separate samples of different categories. For the contrast loss function Regarding the parameters of the feature extractor Gradient of

[0118] Repeat the training process in Phase 1 of Step S1231 until the feature extractor learns a stable and robust feature representation.

[0119] S1232: Training of the linear classifier with a small privacy budget

[0120] Train a linear classifier on the features extracted from private data with a small privacy budget. Since the linear layer has few parameters, the impact of the differential privacy mechanism (such as adding gradient noise) can be reduced, improving the learning ability of the classifier. Assume the linear classifier is where is the weight vector, is the bias. Use the cross-entropy loss for training, which is used to measure the gap between the prediction result of the classifier and the true label :

[0121]

[0122] is the index of the training sample, representing the number of the currently processed sample. The losses of all samples in the training dataset will be accumulated and calculated. is the class index, representing the th class (if it is binary classification, then takes 0 or 1; if it is multi-class classification, then ranges from 0 to . is the true label, used to indicate that the sample belongs to the class . is to take the logarithm to calculate the cross-entropy loss, making the optimization more stable (avoiding the problem of numerical underflow). : The probability that the predicted sample belongs to the class , calculated by the softmax function:

[0123]

[0124] where j is the index value; represents the score of the linear classifier on the class , calculated as follows:

[0125]

[0126] is the weight vector corresponding to the class ; is the category corresponding bias term.

[0127] Optimize the linear classifier parameters and:

[0128]

[0129] where is the learning rate of the linear classifier; : the weight parameter at the current iteration step (before optimization); is the updated weight parameter. is the bias parameter at the current iteration step (before optimization); is the updated bias parameter; is the cross-entropy loss function with respect to the weight gradient. represents the cross-entropy loss function gradient with respect to the bias b.

[0130] Due to the small number of parameters, the impact of gradient noise on model training can be reduced. Differential privacy mechanisms such as gradient clipping and noise injection can be combined to protect data privacy.

[0131] S1233: End-to-end joint training

[0132] Use the remaining budget for end-to-end training to adapt the feature extractor and classifier to the private data. Combine the feature extractor and classifier, and set the joint parameter as and use the overall loss (combining implicit semantic data augmentation loss and classification loss, etc.) for optimization:

[0133]

[0134] is the joint parameter, all model parameters optimized during the end-to-end training process, including: . are the parameters of the feature extractor for extracting high-dimensional features from the input data. are the parameters of the classifier for performing the final classification based on the features. : old parameters, representing the model parameters before optimization, i.e., the parameters in the previous training iteration. : new parameters, representing the model parameters after optimization, i.e., the parameters updated by gradient descent in the current iteration. : the learning rate of joint training, the learning rate that controls the optimization step size and affects the magnitude of parameter update. : the overall loss function, the total loss function used in end-to-end training, containing multiple sub-loss terms. : Loss gradient, calculating the loss function for the model parameters gradient:

[0135]

[0136] This gradient is used to guide parameter updates, optimizing the model in the direction of minimizing the loss.

[0137] Throughout the process, the allocation of the privacy budget affects the training intensity at different stages. For example, with a small privacy budget, more can be allocated to stage two, and as the privacy budget increases, the proportion of stage two is reduced.

[0138] The processing flow of this step S120 specifically includes:

[0139] Language recognition and dynamic adaptation: Determine the text language type (such as Chinese, English, etc.) through the language recognition module, generating a language type embedding vector; Dynamically adjust the parameters of the feature extractor according to the language type (such as inserting a language adaptation layer).

[0140] Feature extraction and covariance prediction: Use a multilingual Transformer (such as XLM-R) to extract basic features. The covariance matrix prediction network (CovNet) predicts the perturbation direction (covariance matrix of Gaussian noise) based on the features.

[0141] Implicit semantic data augmentation: Apply perturbations to the features according to the predicted covariance matrix to generate augmented features.

[0142] Classification and optimization: The classifier outputs class probabilities based on the augmented features. Optimize the parameters of the classifier and CovNet in stages through the meta-learning strategy.

[0143] S130: Extract the structured information of cross-lingual texts through natural language processing techniques, and perform cross-lingual entity alignment and normalization on the feature representations output by the semantic model to generate a unified data structure representation.

[0144] This step S130 performs structured processing and information extraction. Cross-language entity alignment is achieved through cross-language embedding mapping, mapping entities in different languages ​​to a unified semantic space, and normalizing based on similarity (for example, using cosine similarity). Specifically, natural language processing techniques (such as named entity recognition, relationship extraction, event extraction, etc.) are used to extract structured information, such as entities such as time, place, event, and person, from information texts in different languages. Through a cross-language mapping mechanism (such as an alignment model or cross-language embedding), for example, a pre-trained LaBSE model is used to map entities to a unified semantic space to achieve cross-language embedding mapping, and the same or similar entities in different languages ​​are normalized to ensure that information in multilingual texts can be effectively connected. The normalization rules include that if the same entity has multiple names in multiple languages ​​(such as "Paris" and "巴黎"), the English name is used as the standard entity ID. The structured information data is integrated to generate a unified data structure representation (such as a knowledge graph, a relational database, etc.), and stored for subsequent analysis and query.

[0145] According to the above-mentioned embodiment of the present invention, data prior knowledge that is beneficial to the task is acquired through noise prior learning to identify and filter common noise in cross-language data; feature extraction and classification are performed through an adaptive learnable semantic model of a neural network framework that combines dynamic feature extraction and meta-learning strategies, and semantic direction is dynamically predicted through a covariance matrix prediction network to improve the semantic consistency and information alignment capabilities of cross-language texts.

[0146] Result comparison and effect verification:

[0147] Figure 3 This is a graph comparing the performance of the algorithms. To distinguish the proposed algorithm from the traditional Transformer algorithm, the proposed algorithm curve is drawn with a solid dotted line, and the traditional Transformer algorithm curve is drawn with a dashed square line. Figure 3 This is a comparison of algorithm loss convergence. Figure 3 As shown in the figure, the loss of the proposed algorithm dropped to about 0.2 in the 5th round, and approached 0.05 in the 10th round, and the overall convergence was fast. The loss of the Transformer algorithm was still 0.35 in the 5th round, and only dropped to 0.1 in the 10th round, and the convergence was slow. The results show that the proposed algorithm can quickly reduce the loss in a fewer number of training rounds and improve the optimization efficiency.

[0148] Table 1 is a comparison of cross-language classification accuracy (XGLUE benchmark test), and Table 2 is a noise robustness test (including 20% ​​OCR noise data). As shown in Tables 1 and 2, the algorithm proposed in this invention has obvious advantages over the traditional Transformer in both cross-language classification accuracy test and noise robustness test.

[0149] Table 1: Comparison of Cross - language Classification Accuracy (XGLUE Benchmark Test)

[0150]

[0151] Table 2: Noise Robustness Test (Including 20% OCR Noisy Data)

[0152]

[0153] Furthermore, Figure 4 A scatter plot of the feature distribution of cross - language data is described, simulating the distribution of English (blue), Chinese (red), and French (green) in the feature space.

[0154] Among them, the mean of the English data (blue) feature points is (2.05, 2.10), and the standard deviation is approximately (0.45, 0.48). The data is mainly distributed in the range of (1.5, 2.5). The feature distribution is compact, indicating that the features of English extracted by the algorithm are stable. The mean of the Chinese (red) data feature points is (-2.00, 2.05), and the standard deviation is approximately (0.50, 0.47). The data is mainly distributed in the range of (-2.5, -1.5), indicating that the model has good normalization ability for Chinese features. The distribution of Chinese data is more dispersed, with some points deviating from the center, which may indicate that the Chinese text features are more diverse. The mean of the French data (green) feature points is (0.05, -1.95), and the standard deviation is approximately (0.48, 0.49). The data is mainly distributed in the range of (-0.5, 0.5), indicating that the feature extraction of French is also relatively stable, but it extends slightly in the negative direction. These feature vectors show that the proposed algorithm can better distinguish data of different languages, making the feature distributions of the same language more compact. The data of different languages still maintain a certain distribution relationship, meaning that the proposed algorithm can better capture the feature similarities across languages. The average within - group distance is the Euclidean distance from the language data points to their central mean. The average within - group distance of English data is 0.52, that of Chinese data is 0.55, and that of French data is 0.51. The small within - group distance indicates that the clustering effect of feature points of the same language is good, and the proposed algorithm can well distinguish data of different languages. The between - group distance is the Euclidean distance between the central points of different language categories. Among them, the between - group distance between English and Chinese is 4.08, the between - group distance between English and French is 4.05, and the between - group distance between Chinese and French is 4.12. The large between - group distance indicates that the data points of different languages are clearly separated, and the cross - language feature discrimination degree is high, reducing information confusion.

[0155] The above analysis shows that the proposed algorithm has obvious advantages in tasks such as cross - language information extraction, machine translation, and knowledge graph construction.

[0156] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.

[0157] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A cross-language information data collection and structured processing method based on adaptive learning, characterized in that Including: Collecting information data from cross - language information data sources and performing pre - processing; Inputting the pre - processed cross - language text data into an adaptive learnable semantic model of a neural network framework that combines dynamic feature extraction and meta - learning strategies for feature extraction and classification. Among them, the semantic model is trained and optimized based on a covariance matrix prediction network constructed by a multi - layer perceptron. The covariance matrix prediction network takes sample depth features as input and outputs a predicted covariance matrix, dynamically adjusting the perturbation direction so that the semantic model can adaptively adjust the feature extraction method according to different contexts in a multi - language environment; Among them, the training and optimization of the semantic model also include: Optimization of the covariance matrix prediction network based on meta - learning strategies and the adoption of a three - stage training framework for improving the classification ability and cross - language processing ability of the model, including: adopting meta - learning strategies, and training the covariance matrix prediction network and the classifier through three - stage alternating optimization of pseudo - gradient descent update, meta - update, and real - update; Among them, the three - stage training framework includes: Stage 1: Pre - training the feature extractor on synthetic data, Stage 2: Training a linear classifier with a small privacy budget, and Stage 3: End - to - end joint training to allocate the privacy budget and ensure the optimization effect of the model in different training stages; Among them, the meta - learning of the semantic model is optimized by minimizing the implicit semantic data augmentation loss based on the predicted covariance matrix; among them, the implicit semantic data augmentation loss function based on the predicted covariance matrix is: , Among them, is the feature of the th sample; is the predicted feature; is the number of samples within a batch; is the covariance matrix predicted by the covariance matrix prediction network; is the cross-entropy loss: , Among them, is the total number of categories; is the label of the category; is the probability of the predicted category ; The parameters of the covariance matrix prediction network are optimized using the cross - entropy loss of meta - data so that the predicted covariance matrix can accurately model the feature distribution after data augmentation; Extracting the structured information of cross - language text through natural language processing techniques, and performing cross - language entity alignment and normalization on the feature representation output by the semantic model to generate a unified data structure representation.

2. The cross-language information data acquisition and structured processing method based on adaptive learning according to claim 1, characterized in that Among them, The pre - processing includes: Using a synthetic data generation adversarial network to generate noisy text data, training the feature extractor through contrastive representation learning to obtain prior knowledge of the data to filter noise; Applying Gaussian perturbation to the feature vector at the feature level for implicit semantic data augmentation to generate enhanced features, and the formula is: , Among them, is the original feature vector; is the enhanced feature vector; is a scaling factor; is Gaussian noise with a mean of 0 and a covariance matrix of ​ 3. The cross - language information data acquisition and structured processing method based on adaptive learning according to claim 1, characterized in that, The covariance matrix prediction network takes sample depth features as input and outputs a predicted covariance matrix, including: , Among them, is the input feature, is the covariance matrix prediction network based on a multi-layer perceptron, is the predicted covariance matrix.

4. The cross-language information data acquisition and structured processing method based on adaptive learning according to claim 3, characterized in that, Among them, In the three - stage training framework, Pre - training the feature extractor on synthetic data includes: training the feature extractor using a contrastive loss function to optimize the parameters of the feature extractor; Training the linear classifier with a small privacy budget includes: training the linear classifier on the features extracted from private data to reduce the impact of gradient noise; End - to - end joint training includes: performing end - to - end training with the remaining budget to make the feature extractor and the classifier adapt to private data.

5. The cross-language information data acquisition and structured processing method based on adaptive learning according to claim 1, characterized in that Among them, The optimization process of the meta - learning strategy includes: Pseudo - gradient descent update: , Among them, is the classification parameter during the round of optimization; is the number of steps for optimization iteration; is the learning rate; is the gradient of the current implicit semantic data augmentation loss function with respect to the classification parameter ; is the classification parameter after performing the pseudo-update; Meta - update: , Among them, is the covariance matrix prediction network parameter during the round of optimization; is the learning rate of the covariance matrix prediction network; is the gradient of the cross-entropy loss function with respect to the covariance matrix prediction network parameter ; is the covariance matrix prediction network parameter after performing the meta-update; Real - update: , Among them, is the gradient of the implicit semantic data augmentation loss function with respect to ; is the finally updated classification parameter.

6. The cross-language information data acquisition and structured processing method based on adaptive learning according to claim 4, characterized in that Among them, The privacy budget allocation rule of the three - stage training framework is: under the condition of low privacy budget, a higher budget ratio is allocated in Stage 2; as the total privacy budget increases, the budget ratio of Stage 2 is reduced and the budget ratio of Stage 3 is increased.

7. The cross-language information data acquisition and structured processing method based on adaptive learning according to claim 6, wherein Among them, The cross-lingual entity alignment is achieved through cross-lingual embedding mapping, which maps entities in different languages to a unified semantic space and performs normalization based on similarity.

Citation Information

Patent Citations

  • Cross-language abstract abstract method based on robust self-learning strategy in low-resource scene

    CN117271761A

  • Denoising semantic communication system optimization method based on semantic compression

    CN119210652A