Computational method and system for unstructured text data

By performing multi-level processing and feature extraction on unstructured text data, combined with adversarial training and clustering technology, the problem of insufficient utilization of complex semantic relationships and text structure information in the existing technology is solved, and more efficient and accurate text data calculation results are achieved.

CN119474383BActive Publication Date: 2025-05-16北京科杰科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510076252.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-16
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deal with complex semantic relationships, lacks effective utilization of text structure information, and is poorly robust, resulting in low accuracy and efficiency of calculation results.

Method used

By performing multi-level processing of unstructured text data, dynamically adjusting word-partition granularity, constructing part-of-speech co-occurrence matrix, identifying and extracting entity information, generating hierarchical semantic label sequences, and using technologies such as convolution kernel feature extraction, adversarial training and clustering to improve the robustness and classification accuracy of feature representation.

Benefits of technology

It realizes a more accurate understanding of semantic information in unstructured text data, improves the robustness of feature representation, and improves the accuracy and efficiency of classification, especially when processing complex and noisy text data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119474383B_ABST
    Figure CN119474383B_ABST
Patent Text Reader

Abstract

The present invention provides a calculation method and system for unstructured text data, which relates to the technical field of natural language processing, including multi-level processing of input unstructured text data, dynamically adjusting the word segmentation granularity according to the word frequency distribution, and constructing a part-of-speech co-occurrence matrix in combination with contextual semantic information, extracting entity information, and fusing the part-of-speech co-occurrence matrix and entity information to generate a hierarchical semantic label sequence. Feature extraction units with different convolution kernel sizes are used to extract feature representations, and the cosine similarity between different semantic levels is calculated to establish an association weight matrix. Semantic enhancement vectors are constructed based on entity information, and adversarial training is performed to obtain a multimodal semantic feature matrix. The semantic similarity between fused feature vectors is calculated for clustering, and the difficulty weights are set according to the complexity, consistency and fuzziness of the clusters and then input into a classifier for sorting, and the classification results are iteratively optimized to obtain calculation results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to natural language processing technology, and in particular to a calculation method and system for unstructured text data. Background Art

[0002] Natural language processing is an important branch of artificial intelligence, which aims to enable computers to understand, interpret and generate human language. Unstructured text data is growing explosively. Traditional text computing methods are usually based on bag-of-words models or simple statistical methods, which are difficult to capture the rich semantic information and complex structural relationships in text data. For example, in tasks such as sentiment analysis, text classification and information retrieval, it is difficult to accurately understand the meaning of the text by relying solely on word frequency statistics, resulting in low accuracy and efficiency of the calculation results. The existing technology mainly has the following defects and deficiencies:

[0003] Difficulty in effectively processing complex semantic relationships: Traditional text computing methods have difficulty capturing complex semantic relationships in text, such as the polysemy of words, the hierarchy of semantics, and the influence of context, etc. This results in the calculation results being often inaccurate when processing complex text data, making it difficult to meet the needs of practical applications.

[0004] Lack of effective use of text structure information: Most existing methods ignore the structural information of the text, such as paragraph structure, sentence relations, and connections between entities. This structural information is crucial for understanding the overall meaning of the text. The lack of use of structural information limits the depth and breadth of text computing.

[0005] Poor robustness: Traditional text computing methods are easily affected by noisy data, such as spelling errors, grammatical errors, and semantic ambiguity, which leads to poor stability and reliability of computing results when processing real-world text data. Summary of the invention

[0006] The embodiments of the present invention provide a method and system for calculating unstructured text data, which can solve the problems in the prior art.

[0007] According to a first aspect of the embodiments of the present invention,

[0008] Provides computational methods for unstructured text data, including:

[0009] Perform multi-level processing on the input unstructured text data, dynamically adjust the word segmentation granularity according to the word frequency distribution, build a part-of-speech co-occurrence matrix for part-of-speech tagging in combination with contextual semantic information, identify and extract entity information from the text data, fuse the part-of-speech co-occurrence matrix and the entity information to generate a hierarchical semantic label sequence, and unify the case and replace special characters on the hierarchical semantic label sequence;

[0010] A feature extraction unit with different convolution kernel sizes is used to extract features from the hierarchical semantic tag sequence to obtain a feature representation, an association weight matrix is ​​established by calculating the cosine similarity between different semantic levels in the feature representation, a semantic enhancement vector is constructed based on the category attributes and hierarchical relationships of the entity information, a preset proportion of random noise is inserted into the feature representation to generate adversarial samples for adversarial training, and the trained feature representation, the association weight matrix and the semantic enhancement vector are combined in series on the feature dimension to obtain a multimodal semantic feature matrix;

[0011] The semantic similarity of adjacent fused feature vectors in the multimodal semantic feature matrix is ​​calculated to obtain a similarity matrix, the fused feature vectors are clustered based on the similarity matrix to obtain feature clusters, difficulty weights are set according to the structural complexity, semantic consistency and boundary fuzziness of the feature clusters, the feature clusters are sorted according to the difficulty weights and input into the classifier in sequence, and the category probability distribution and cluster center distance are calculated for each feature cluster; the contrast loss of sample pairs within the cluster is used to construct an optimization target, and the prediction results of the classifier are iteratively optimized based on the optimization target to obtain the calculation results of unstructured text data.

[0012] In an optional embodiment,

[0013] The steps of performing multi-level processing on the input unstructured text data, dynamically adjusting the word segmentation granularity according to the word frequency distribution, constructing a part-of-speech co-occurrence matrix for part-of-speech tagging in combination with contextual semantic information, identifying and extracting entity information from the text data, and fusing the part-of-speech co-occurrence matrix and the entity information to generate a hierarchical semantic label sequence include:

[0014] Scan the input text using a sliding window, calculate the frequency and the normalized mutual information value of the character sequence in the window; construct a window scoring function based on the frequency of the character sequence and the normalized mutual information value, the window scoring function includes a normalized mutual information value component, a frequency component and a sequence length component, and adaptively adjust the size of the sliding window according to the continuous calculation results of the window scoring function and the corresponding word frequency distribution to obtain an initial word segmentation unit;

[0015] Obtaining the vector representation of the initial word segmentation unit, calculating the cosine similarity of adjacent word segmentation unit vectors to obtain context relevance, merging adjacent word segmentation units that meet the relevance merging threshold and whose combined frequency exceeds a preset minimum word frequency threshold, calculating the standardized mutual information value of the internal segmentation position for the merged word segmentation unit, performing segmentation at the position when the ratio of the standardized mutual information values ​​before and after segmentation exceeds a preset mutual information segmentation threshold, and dynamically adjusting the word segmentation granularity to obtain an optimized word segmentation unit;

[0016] Based on the optimized word segmentation unit, a word frequency statistics window is constructed to record the word frequency information of the processed document, the word frequency distribution is recalculated and the mutual information segmentation threshold is adjusted based on the documents in the word frequency statistics window, and the word segmentation effect evaluation index is calculated. When the word segmentation effect evaluation index decreases, the weight of each component in the window scoring function is dynamically adjusted, and the optimized word segmentation sequence and feature information of each word segmentation unit are output, wherein the feature information includes the unit combination mutual information value, context relevance information and part of speech information;

[0017] A part-of-speech co-occurrence matrix is ​​constructed based on the part-of-speech information in the feature information, and a part-of-speech transfer probability matrix is ​​generated using the transfer relationship between parts of speech. Word-level features and character features are extracted in combination with the unit combination mutual information value and context relevance information to perform entity information recognition. The part-of-speech combination pattern represented by the part-of-speech co-occurrence matrix and the part-of-speech sequence predicted by the part-of-speech transfer probability matrix are fused with the entity information recognition result to construct a multi-layer semantic label, and a semantic label sequence is generated according to the hierarchical structure of basic parts of speech, entity type and attribute information.

[0018] In an optional embodiment,

[0019] The step of inserting a preset proportion of random noise into the feature representation to generate adversarial samples for adversarial training includes:

[0020] Generate random noise based on Gaussian distribution, and superimpose the random noise with the input feature vector to obtain an adversarial sample;

[0021] During the training process, the gradient of the feature extraction loss to the feature vector is calculated to obtain a feature importance score, and the noise intensity coefficient is adjusted inversely proportionally according to the relative size of the feature importance score, where the noise intensity coefficient is inversely proportional to the feature importance score;

[0022] Calculate the cosine similarity between the adversarial sample and the original feature vector to obtain feature consistency, calculate the Euclidean distance ratio between the adversarial sample and the original feature vector to obtain the perturbation amplitude, calculate the predicted output deviation between the adversarial sample and the original feature vector to obtain the predicted deviation, and include adversarial samples that simultaneously meet the feature consistency threshold, the perturbation amplitude threshold, and the predicted deviation threshold into training;

[0023] Constructing a dynamically weighted adversarial training loss function, wherein the dynamically weighted adversarial training loss function includes an original sample loss term and an adversarial sample loss term, and dynamically adjusting weight coefficients of the original sample loss term and the adversarial sample loss term based on the training round and the accuracy of the validation set of feature extraction;

[0024] Build an adversarial sample library, calculate the effectiveness score of the adversarial samples based on feature consistency, perturbation amplitude and prediction deviation, add adversarial samples with a score higher than the threshold to the adversarial sample library, remove the lowest-scoring samples when the adversarial sample library exceeds the preset capacity, regularly re-evaluate the effectiveness scores of samples in the adversarial sample library, and adaptively adjust the screening threshold based on the accuracy of the feature extraction verification set to update the adversarial sample library.

[0025] In an optional embodiment,

[0026] During the training process, the gradient of the feature extraction loss to the feature vector is calculated to obtain the feature importance score, and the noise intensity coefficient is adjusted inversely proportionally according to the relative size of the feature importance score. The step of adjusting the noise intensity coefficient inversely proportional to the feature importance score includes:

[0027] Obtaining the gradient value of the feature vector to the feature extraction loss during the training process, and calculating the local importance according to the gradient value of the feature extraction loss, wherein the local importance is obtained by multiplying the absolute value of the feature gradient by the local sensitivity function, and the local sensitivity function is calculated based on the gradient difference before and after the feature perturbation;

[0028] Calculating the global importance based on the local importance, the global importance is obtained by multiplying the expected value of the local importance by the feature dimension contribution coefficient, the feature dimension contribution coefficient is obtained by calculating the prediction accuracy change gradient of the feature in multiple training batches, and the prediction accuracy change gradient is determined based on the degree of influence of the fluctuation of the feature between training batches on the performance of feature extraction;

[0029] Constructing a feature interaction graph to model the dynamic correlation between features, concatenating the local importance and the global importance and mapping them through a learnable projection matrix to obtain a feature interaction matrix, and updating the local importance and the global importance based on the feature interaction matrix to obtain a feature importance score;

[0030] The feature importance score is input into a nonlinear mapping function to obtain a noise intensity coefficient. The nonlinear mapping function includes a product of a hyperbolic tangent term and a sigmoid term. The hyperbolic tangent term is used to control the overall noise range, and the sigmoid term is used to construct an inverse relationship between the feature importance score and the noise intensity coefficient. An oscillation term is introduced into the noise intensity coefficient to obtain a final noise intensity. The oscillation term modulates the feature importance score through a sine function to increase the diversity of adversarial samples.

[0031] In an optional embodiment,

[0032] The steps of calculating the semantic similarity of adjacent fused feature vectors in the multimodal semantic feature matrix to obtain a similarity matrix, clustering the fused feature vectors based on the similarity matrix to obtain feature clusters, setting difficulty weights according to the structural complexity, semantic consistency and boundary fuzziness of the feature clusters, sorting the feature clusters according to the difficulty weights and sequentially inputting the feature clusters into the classifier, and calculating the category probability distribution and cluster center distance for each feature cluster include:

[0033] Obtaining the local structural distance and the global semantic distance of adjacent fused feature vectors, wherein the local structural distance is calculated by an attention-weighted k-nearest neighbor graph, and the global semantic distance is calculated by a contrastive learning framework, and dynamically fusing the local structural distance and the global semantic distance by a time-varying weight coefficient to obtain a similarity matrix;

[0034] Based on the similarity matrix, local neighborhood density estimation is performed on the fused feature vector, wherein the local neighborhood density estimation adopts an adaptive bandwidth parameter, and the adaptive bandwidth parameter is obtained by multiplying the median of the neighborhood similarity by a density-dependent adjustment factor;

[0035] Calculate the node representative score based on the local neighborhood density estimation result, the node representative score is obtained by multiplying the node density by the minimum similarity of the density advantage neighbor, the density advantage neighbor is the adjacent node with a density value greater than the current node density value, determine the node with a node representative score greater than the time-varying split threshold as a split node, and recursively construct a subclass cluster for the split node to obtain a hierarchical feature cluster;

[0036] The difficulty indexes of structural complexity, semantic consistency and boundary fuzziness are calculated for the hierarchical feature clusters respectively. The structural complexity is calculated based on the information entropy of the distance distribution between samples in the cluster, the semantic consistency is calculated based on the cosine similarity of the sample to the cluster center, and the boundary fuzziness is calculated based on the minimum similarity of samples inside and outside the cluster;

[0037] The difficulty index is input into the meta-learning module to obtain the difficulty weight, the difficulty index is weighted by the difficulty weight to obtain the comprehensive difficulty score of the cluster, the clusters are sorted based on the comprehensive difficulty score to construct a difficulty sequence, classifiers are trained for clusters of different difficulty levels in the difficulty sequence, the membership of the input sample and the difficulty level is calculated, and the membership is used as the combined weight of the prediction results of classifiers at each level to obtain the category probability distribution; the difference between the sample feature vector and the cluster center vector is calculated, and the difference is multiplied by the inverse matrix of the cluster covariance matrix to obtain the cluster center distance; the final classification result is obtained based on the weighted combination of the category probability distribution and the cluster center distance.

[0038] In an optional embodiment,

[0039] The difficulty index is input into a meta-learning module to obtain a difficulty weight, the difficulty index is weighted by the difficulty weight to obtain a comprehensive difficulty score of the cluster, the clusters are sorted based on the comprehensive difficulty score to construct a difficulty sequence, and the steps of training classifiers for clusters of different difficulty levels in the difficulty sequence include:

[0040] A meta-learning network with two-layer transformation is designed, wherein the difficulty feature vector corresponding to the difficulty index is mapped to the hidden feature space through the first layer transformation to obtain the hidden layer feature, and the first layer transformation adopts the hyperbolic tangent activation function; the hidden layer feature is mapped to the weight space through the second layer transformation to obtain the difficulty weight, and the second layer transformation adopts the softmax normalization function; the dimension of the weight matrix of the first layer transformation corresponds to the dimension of the hidden layer feature, and the dimension of the weight matrix of the second layer transformation corresponds to the dimension of the difficulty index;

[0041] A first loss term is constructed based on the cross entropy loss of the classifier on the validation set of feature clustering; a regularization loss term is constructed based on the second norm of the difficulty weight and the difficulty weight difference, and the difficulty weight difference is calculated by the absolute value difference between the difficulty weights; the first loss term and the regularization loss term are weighted and combined to obtain a meta-learning loss function; the meta-learning network is trained using an alternating optimization strategy, the classifier parameters are fixed, and the transformation parameters of the meta-learning network are gradient updated based on the meta-learning loss function; the classifier parameters are retrained based on the updated difficulty weight; the meta-learning network parameter update and the classifier parameter update process are alternately executed until the performance of the validation set of feature clustering converges to obtain the final difficulty weight; the difficulty feature vector is weighted using the difficulty weight obtained by training to obtain a comprehensive difficulty score of the cluster; a smoothing process is performed based on the comprehensive difficulty scores of the cluster and the adjacent clusters to obtain a smoothed difficulty score, the clusters are sorted according to the smoothed difficulty score, and the clusters are divided into multiple difficulty levels by the difficulty threshold determined adaptively;

[0042] Classifiers are trained for different difficulty levels. For the first difficulty level, a difficulty-dependent loss weight function is introduced, and the loss weight function is calculated based on the difference between the sample difficulty and the average difficulty. For the second difficulty level, a constraint term based on the mean square error between the input features and the reconstructed features is introduced. For the third difficulty level, a comparative learning term consisting of the feature similarity of the same type of sample pairs and the feature difference of the different type of sample pairs is introduced. The integration of the classifier output results adopts a joint weighting method of difficulty membership based on a Gaussian kernel and confidence of the prediction entropy to obtain a classifier corresponding to the difficulty level.

[0043] In an optional embodiment,

[0044] The steps of constructing an optimization target by using the contrast loss of sample pairs within a cluster, iteratively optimizing the prediction result of the classifier based on the optimization target, and obtaining the calculation result of the unstructured text data include:

[0045] Preprocess the prediction results of the classifier, use the local sensitive hashing method to construct the neighbor index of the samples in the cluster, calculate the cosine similarity between the sample pairs based on the neighbor index, set an adaptive similarity threshold according to the mean and standard deviation of the cosine similarity of the sample pairs in the cluster, and use the sample pairs greater than the adaptive similarity threshold as the candidate positive sample pair set;

[0046] Construct a multi-layer heterogeneous graph structure based on the candidate positive sample pair set, use samples as nodes to calculate edge weights through an attention mechanism to construct a sample-level graph, perform spectral clustering on the edge weight matrix of the sample-level graph to obtain subclass clusters and construct a subclass cluster-level graph based on the member sample interaction intensity, construct class cluster nodes based on classification labels, and construct a class cluster-level graph through the subclass cluster affiliation relationship;

[0047] Designing a multi-granularity contrast loss function for the multi-layer heterogeneous graph structure, including an intra-layer contrast loss for enhancing the consistency of representation of nodes in the same layer, an inter-layer contrast loss for cross-layer distribution alignment, and a graph structure consistency loss for ensuring hierarchical synergy, and automatically adjusting the weight coefficients of different loss items through a reinforcement learning method;

[0048] The multi-granularity contrast loss function is combined with the cross entropy loss output by the classifier to form an optimization target, and an adaptive learning rate is calculated based on the current batch sample distribution; a three-stage alternating optimization strategy is adopted, in which fixed classifier parameters optimize feature extractor parameters to minimize multi-granularity contrast loss, fixed feature extractor parameters optimize classifier parameters to minimize classification loss, and finally, parameters of feature extractor and classifier are simultaneously optimized to minimize the optimization target, and an early stopping strategy is introduced; wherein the feature representation output by the feature extractor is transferred between different optimization stages to construct a multi-layer heterogeneous graph structure and calculate a multi-granularity contrast loss function; unstructured text data is classified based on the optimized feature extractor and classifier to obtain a calculation result.

[0049] According to a second aspect of the embodiments of the present invention,

[0050] Computational systems that provide unstructured text data, including:

[0051] The first unit is used to perform multi-level processing on the input unstructured text data, dynamically adjust the word segmentation granularity according to the word frequency distribution, construct a part-of-speech co-occurrence matrix for part-of-speech tagging in combination with contextual semantic information, identify and extract entity information from the text data, fuse the part-of-speech co-occurrence matrix and the entity information to generate a hierarchical semantic label sequence, and unify the uppercase and lowercase letters and replace special characters on the hierarchical semantic label sequence;

[0052] The second unit is used to extract features from the hierarchical semantic tag sequence using feature extraction units with different convolution kernel sizes to obtain feature representation, establish an association weight matrix by calculating the cosine similarity between different semantic levels in the feature representation, construct a semantic enhancement vector based on the category attributes and hierarchical relationships of the entity information, insert a preset proportion of random noise into the feature representation to generate adversarial samples for adversarial training, and combine the trained feature representation, association weight matrix and semantic enhancement vector in series on the feature dimension to obtain a multimodal semantic feature matrix;

[0053] The third unit is used to calculate the semantic similarity of adjacent fused feature vectors in the multimodal semantic feature matrix to obtain a similarity matrix, cluster the fused feature vectors based on the similarity matrix to obtain feature clusters, set difficulty weights according to the structural complexity, semantic consistency and boundary fuzziness of the feature clusters, sort the feature clusters according to the difficulty weights and input them into the classifier in sequence, calculate the category probability distribution and cluster center distance for each feature cluster; use the contrast loss of sample pairs within the cluster to construct an optimization target, iteratively optimize the prediction results of the classifier based on the optimization target, and obtain the calculation results of unstructured text data.

[0054] According to a third aspect of the embodiments of the present invention,

[0055] An electronic device is provided, comprising:

[0056] processor;

[0057] a memory for storing processor-executable instructions;

[0058] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0059] A fourth aspect of the embodiments of the present invention is:

[0060] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0061] The present invention can more accurately understand the semantic information in unstructured text data through technical means such as multi-level processing, part-of-speech co-occurrence matrix, and entity information fusion, especially has significant advantages in processing complex sentences and ambiguous expressions.

[0062] The present invention adopts strategies such as multimodal feature fusion and adversarial training to effectively improve the robustness of feature representation, reduce the impact of noise and abnormal data on the calculation results, and make the model more stable when facing different types of text data.

[0063] The present invention optimizes the training process of the classifier based on mechanisms such as difficulty weight ranking of feature clusters and sample contrast loss within clusters, improves the accuracy and efficiency of classification, and has better effects when processing data sets with unbalanced categories and blurred boundaries. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 A schematic diagram of a flow chart of a method for calculating unstructured text data according to an embodiment of the present invention;

[0065] Figure 2 A schematic diagram of the structure of a computing system for unstructured text data according to an embodiment of the present invention. DETAILED DESCRIPTION

[0066] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0067] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0068] Figure 1 FIG. 1 is a flow chart of a method for calculating unstructured text data according to an embodiment of the present invention. Figure 1 As shown, the method includes:

[0069] S1. Perform multi-level processing on the input unstructured text data, dynamically adjust the word segmentation granularity according to the word frequency distribution, construct a part-of-speech co-occurrence matrix for part-of-speech tagging in combination with contextual semantic information, identify and extract entity information from the text data, fuse the part-of-speech co-occurrence matrix and the entity information to generate a hierarchical semantic label sequence, and unify the case and replace special characters on the hierarchical semantic label sequence;

[0070] S2. Use feature extraction units with different convolution kernel sizes to extract features from the hierarchical semantic tag sequence to obtain feature representation, establish an association weight matrix by calculating the cosine similarity between different semantic levels in the feature representation, construct a semantic enhancement vector based on the category attributes and hierarchical relationships of the entity information, insert a preset proportion of random noise into the feature representation to generate adversarial samples for adversarial training, and combine the trained feature representation, association weight matrix and semantic enhancement vector in series on the feature dimension to obtain a multimodal semantic feature matrix;

[0071] S3. Calculate the semantic similarity of adjacent fused feature vectors in the multimodal semantic feature matrix to obtain a similarity matrix, cluster the fused feature vectors based on the similarity matrix to obtain feature clusters, set difficulty weights according to the structural complexity, semantic consistency and boundary fuzziness of the feature clusters, sort the feature clusters according to the difficulty weights and input them into the classifier in sequence, calculate the category probability distribution and cluster center distance for each feature cluster; use the contrast loss of sample pairs within the cluster to construct an optimization target, iteratively optimize the prediction results of the classifier based on the optimization target, and obtain the calculation results of unstructured text data.

[0072] In an optional embodiment,

[0073] The steps of performing multi-level processing on the input unstructured text data, dynamically adjusting the word segmentation granularity according to the word frequency distribution, constructing a part-of-speech co-occurrence matrix for part-of-speech tagging in combination with contextual semantic information, identifying and extracting entity information from the text data, and fusing the part-of-speech co-occurrence matrix and the entity information to generate a hierarchical semantic label sequence include:

[0074] Scan the input text using a sliding window, calculate the frequency and the normalized mutual information value of the character sequence in the window; construct a window scoring function based on the frequency of the character sequence and the normalized mutual information value, the window scoring function includes a normalized mutual information value component, a frequency component and a sequence length component, and adaptively adjust the size of the sliding window according to the continuous calculation results of the window scoring function and the corresponding word frequency distribution to obtain an initial word segmentation unit;

[0075] Obtaining the vector representation of the initial word segmentation unit, calculating the cosine similarity of adjacent word segmentation unit vectors to obtain context relevance, merging adjacent word segmentation units that meet the relevance merging threshold and whose combined frequency exceeds a preset minimum word frequency threshold, calculating the standardized mutual information value of the internal segmentation position for the merged word segmentation unit, performing segmentation at the position when the ratio of the standardized mutual information values ​​before and after segmentation exceeds a preset mutual information segmentation threshold, and dynamically adjusting the word segmentation granularity to obtain an optimized word segmentation unit;

[0076] Based on the optimized word segmentation unit, a word frequency statistics window is constructed to record the word frequency information of the processed document, the word frequency distribution is recalculated and the mutual information segmentation threshold is adjusted based on the documents in the word frequency statistics window, and the word segmentation effect evaluation index is calculated. When the word segmentation effect evaluation index decreases, the weight of each component in the window scoring function is dynamically adjusted, and the optimized word segmentation sequence and feature information of each word segmentation unit are output, wherein the feature information includes the unit combination mutual information value, context relevance information and part of speech information;

[0077] A part-of-speech co-occurrence matrix is ​​constructed based on the part-of-speech information in the feature information, and a part-of-speech transfer probability matrix is ​​generated using the transfer relationship between parts of speech. Word-level features and character features are extracted in combination with the unit combination mutual information value and context relevance information to perform entity information recognition. The part-of-speech combination pattern represented by the part-of-speech co-occurrence matrix and the part-of-speech sequence predicted by the part-of-speech transfer probability matrix are fused with the entity information recognition result to construct a multi-layer semantic label, and a semantic label sequence is generated according to the hierarchical structure of basic parts of speech, entity type and attribute information.

[0078] Exemplarily, first, the input text is segmented initially. The sliding window mechanism is used to scan the text character by character. For example, the initial window size is set to 3 characters. During the scanning process, the frequency of occurrence of the character sequence in the window is counted, and its standardized mutual information value is calculated. For example, for the text "I like to eat apples", the window scans "I like", "like", "enjoy eating", "eat apples", and "apples" in sequence. The number of occurrences of "I like" and the number of occurrences of "I" and "like" are counted, and the standardized mutual information value of "I like" is calculated. The window scoring function comprehensively considers the standardized mutual information value, frequency, and sequence length, for example, giving a higher weight to the standardized mutual information value. According to the continuous calculation results of the window scoring function and the word frequency distribution, the size of the sliding window is dynamically adjusted. For example, if the frequency of "like" is very high and the standardized mutual information value is also high, the window size is expanded; if the frequency and standardized mutual information value of "enjoy eating" are very low, the window size is reduced. In this way, initial segmentation units such as "I", "like", "eat", and "apple" can be obtained.

[0079] Next, the initial word segmentation units are merged and split, and the word segmentation granularity is adjusted dynamically. Each word segmentation unit is converted into a vector representation, for example, using Word2Vec. The cosine similarity of the vectors of adjacent word segmentation units is calculated, such as the cosine similarity of "like" and "eat". If the similarity exceeds the preset merging threshold and the combined frequency (the frequency of "like to eat") exceeds the preset minimum threshold of word frequency, the adjacent word segmentation units are merged, for example, "like" and "eat" are merged into "like to eat". For the merged word segmentation units, the standardized mutual information value of the internal segmentation position is calculated, for example, the standardized mutual information value of "like / eat" is calculated. If the ratio of the standardized mutual information values ​​before and after the segmentation exceeds the preset mutual information segmentation threshold, the segmentation is performed at this position. For example, if the ratio of the standardized mutual information value of "like / eat" to the standardized mutual information value of "like to eat" is large, "like to eat" is split into "like" and "eat". After dynamic adjustment of merging and segmentation, optimized word segmentation units are obtained, such as "I", "like", "eat", and "apple".

[0080] Then, a word frequency statistics window is constructed based on the optimized word segmentation unit to record the word frequency information of the processed documents. For example, the frequency of occurrence of words such as "I", "like", "eat", and "apple" in the processed documents is counted. The word frequency distribution is recalculated and the mutual information segmentation threshold is adjusted based on the documents in the word frequency statistics window. The word segmentation effect evaluation index is calculated, for example, using the F1 value. If the word segmentation effect evaluation index decreases, the weight of each component in the window scoring function is dynamically adjusted, such as reducing the weight of the frequency component and increasing the weight of the standardized mutual information value component. Finally, the optimized word segmentation sequence and the feature information of each word segmentation unit are output, including the unit combination mutual information value (for example, the mutual information value of "like to eat"), context relevance information (for example, the cosine similarity between "like" and "eat"), and part of speech information (for example, "like" is a verb and "apple" is a noun).

[0081] Finally, a hierarchical semantic label sequence is generated based on the feature information of the word segmentation unit. A part-of-speech co-occurrence matrix is ​​constructed based on the part-of-speech information, for example, the number of times a verb and a noun appear together is counted. A part-of-speech transfer probability matrix is ​​generated using the transfer relationship between parts of speech, for example, the probability of a verb being followed by a noun is calculated. Word-level features and character features are extracted by combining the unit combination mutual information value and context relevance information for entity information recognition, for example, recognizing that "apple" is a fruit entity. The part-of-speech combination pattern represented by the part-of-speech co-occurrence matrix and the part-of-speech sequence predicted by the part-of-speech transfer probability matrix are fused with the entity information recognition results to construct a multi-layer semantic label, for example, "like" is marked as a verb, "eat" is marked as a verb, and "apple" is marked as a noun and a fruit entity. A semantic label sequence is generated according to the hierarchical structure of basic part-of-speech, entity type, and attribute information.

[0082] By dynamically adjusting the word segmentation granularity, the present invention can more accurately identify word boundaries in a text, especially has advantages in processing unregistered words and ambiguous words; by constructing a part-of-speech co-occurrence matrix and a part-of-speech transition probability matrix, it can extract richer part-of-speech information and contextual semantic information, so as to better understand the meaning of the text; the generated hierarchical semantic label sequence can support more sophisticated semantic analysis, such as entity recognition, relationship extraction, etc., and provide richer semantic information for downstream applications.

[0083] In an optional embodiment,

[0084] The step of inserting a preset proportion of random noise into the feature representation to generate adversarial samples for adversarial training includes:

[0085] Generate random noise based on Gaussian distribution, and superimpose the random noise with the input feature vector to obtain an adversarial sample;

[0086] During the training process, the gradient of the feature extraction loss to the feature vector is calculated to obtain a feature importance score, and the noise intensity coefficient is adjusted inversely proportionally according to the relative size of the feature importance score, where the noise intensity coefficient is inversely proportional to the feature importance score;

[0087] Calculate the cosine similarity between the adversarial sample and the original feature vector to obtain feature consistency, calculate the Euclidean distance ratio between the adversarial sample and the original feature vector to obtain the perturbation amplitude, calculate the predicted output deviation between the adversarial sample and the original feature vector to obtain the predicted deviation, and include adversarial samples that simultaneously meet the feature consistency threshold, the perturbation amplitude threshold, and the predicted deviation threshold into training;

[0088] Constructing a dynamically weighted adversarial training loss function, wherein the dynamically weighted adversarial training loss function includes an original sample loss term and an adversarial sample loss term, and dynamically adjusting weight coefficients of the original sample loss term and the adversarial sample loss term based on the training round and the accuracy of the validation set of feature extraction;

[0089] Build an adversarial sample library, calculate the effectiveness score of the adversarial samples based on feature consistency, perturbation amplitude and prediction deviation, add adversarial samples with a score higher than the threshold to the adversarial sample library, remove the lowest-scoring samples when the adversarial sample library exceeds the preset capacity, regularly re-evaluate the effectiveness scores of samples in the adversarial sample library, and adaptively adjust the screening threshold based on the accuracy of the feature extraction verification set to update the adversarial sample library.

[0090] Exemplarily, first, prepare a training dataset and a pre-trained model. The training dataset contains a large number of labeled samples for model training and verification. The pre-trained model can be any deep learning model suitable for a specific task, such as a convolutional neural network (CNN) in an image classification task or a recurrent neural network (RNN) in a natural language processing task.

[0091] Next, perform feature extraction. Use the pre-trained model to extract features from the input sample and obtain a feature vector representation of the sample. For example, in an image classification task, if you input a picture of a cat, the pre-trained model will extract a feature vector representing the cat, such as color, texture, shape, etc.

[0092] Then, generate adversarial samples. Generate random noise based on Gaussian distribution and add the noise to the extracted feature vector to generate adversarial samples. For example, assuming the feature vector is [0.1, 0.2, 0.3] and the generated random noise is [0.01, 0.02, 0.03], then the superimposed adversarial sample feature vector is [0.11, 0.22, 0.33]. The initial value of the noise intensity coefficient can be set to 0.1.

[0093] In order to make adversarial examples more effective, the noise intensity needs to be adjusted according to the feature importance. During the training process, the gradient of the feature extraction loss with respect to the feature vector is calculated to obtain the feature importance score. Assuming the feature importance score is [0.5, 0.3, 0.2], the noise intensity coefficient is adjusted in the opposite proportional dimension. The higher the feature importance score, the lower the corresponding noise intensity coefficient. For example, the noise intensity coefficient can be adjusted to [0.02, 0.033, 0.05].

[0094] Next, screen effective adversarial samples. Calculate the cosine similarity between the adversarial sample and the original feature vector to obtain feature consistency. Calculate the ratio of the Euclidean distance between the adversarial sample and the original feature vector to obtain the perturbation amplitude. Calculate the predicted output deviation between the adversarial sample and the original feature vector to obtain the prediction deviation. For example, assume that the feature consistency threshold is 0.9, the perturbation amplitude threshold is 0.1, and the prediction deviation threshold is 0.05. Only adversarial samples that meet all three thresholds will be included in the training.

[0095] Construct a dynamically weighted adversarial training loss function. This loss function contains the original sample loss term and the adversarial sample loss term. For example, in the initial state, the weight coefficient of the original sample loss term is 0.8, and the weight coefficient of the adversarial sample loss term is 0.2. As the number of training rounds increases and the accuracy of the feature extraction validation set changes, these two weight coefficients are dynamically adjusted. For example, as training progresses, if the accuracy of the validation set improves, the weight of the adversarial sample loss term is gradually increased, such as adjusted to 0.5, while the weight of the original sample loss term is reduced to 0.5.

[0096] Build an adversarial sample library. Calculate the effectiveness score of adversarial samples based on feature consistency, perturbation amplitude, and prediction bias. Add adversarial samples with a score higher than the threshold to the adversarial sample library. For example, if the score threshold is set to 0.8, only adversarial samples with a score higher than 0.8 will be added to the sample library. When the adversarial sample library exceeds the preset capacity, for example, the capacity is 1000, remove the sample with the lowest score. Regularly re-evaluate the effectiveness score of samples in the adversarial sample library, and adaptively adjust the screening threshold based on the accuracy of the feature extraction validation set to update the adversarial sample library.

[0097] Through adversarial training, the model of the present invention can better resist the attack of adversarial samples and improve the stability and reliability of the model under malicious attacks. Adversarial training can help the model learn more essential features and reduce dependence on noise and bias in training data, thereby improving the generalization performance of the model on unseen data. The noise intensity is adjusted according to the importance of the features, making the generated adversarial samples more effective and more targeted to improve the robustness of the model on key features.

[0098] In an optional embodiment,

[0099] During the training process, the gradient of the feature extraction loss to the feature vector is calculated to obtain the feature importance score, and the noise intensity coefficient is adjusted inversely proportionally according to the relative size of the feature importance score. The step of adjusting the noise intensity coefficient inversely proportional to the feature importance score includes:

[0100] Obtaining the gradient value of the feature vector to the feature extraction loss during the training process, and calculating the local importance according to the gradient value of the feature extraction loss, wherein the local importance is obtained by multiplying the absolute value of the feature gradient by the local sensitivity function, and the local sensitivity function is calculated based on the gradient difference before and after the feature perturbation;

[0101] Calculating the global importance based on the local importance, the global importance is obtained by multiplying the expected value of the local importance by the feature dimension contribution coefficient, the feature dimension contribution coefficient is obtained by calculating the prediction accuracy change gradient of the feature in multiple training batches, and the prediction accuracy change gradient is determined based on the degree of influence of the fluctuation of the feature between training batches on the performance of feature extraction;

[0102] Constructing a feature interaction graph to model the dynamic correlation between features, concatenating the local importance and the global importance and mapping them through a learnable projection matrix to obtain a feature interaction matrix, and updating the local importance and the global importance based on the feature interaction matrix to obtain a feature importance score;

[0103] The feature importance score is input into a nonlinear mapping function to obtain a noise intensity coefficient. The nonlinear mapping function includes a product of a hyperbolic tangent term and a sigmoid term. The hyperbolic tangent term is used to control the overall noise range, and the sigmoid term is used to construct an inverse relationship between the feature importance score and the noise intensity coefficient. An oscillation term is introduced into the noise intensity coefficient to obtain a final noise intensity. The oscillation term modulates the feature importance score through a sine function to increase the diversity of adversarial samples.

[0104] Exemplarily, first, obtain the gradient value of the feature extraction loss. Feed the input data into the feature extractor to obtain a feature vector. Calculate the feature extraction loss and obtain the gradient value of the feature vector to the loss. These gradient values ​​reflect the degree of influence of each dimension of the feature vector on the change of the loss. For example, assuming that the feature vector dimension is 5, the obtained gradient value is [0.1, 0.5, 0.2, 0.05, 0.15].

[0105] Then, calculate the local importance. The local importance of each feature dimension is calculated based on the gradient value of the feature extraction loss. The local importance is obtained by multiplying the absolute value of the feature gradient by the local sensitivity function. The local sensitivity function is calculated based on the gradient difference before and after the feature perturbation. For example, add a small perturbation to each dimension of the feature vector and observe the change in the gradient before and after the perturbation. If the gradient changes greatly after the perturbation, it means that the dimension has a greater impact on the model prediction and has a higher sensitivity. Assuming that the calculated local sensitivity function value is [0.8, 0.9, 0.7, 0.6, 0.8], the local importance is [0.08, 0.45, 0.14, 0.03, 0.12].

[0106] Next, calculate the global importance. Calculate the global importance based on the local importance. The global importance is obtained by multiplying the expected value of the local importance by the contribution coefficient of the feature dimension. The contribution coefficient of the feature dimension is obtained by calculating the gradient of the prediction accuracy change of the feature in multiple training batches. The gradient of the prediction accuracy change is determined based on the degree of influence of the fluctuation of the feature between training batches on the performance of feature extraction. For example, by counting the fluctuation of the feature vectors in multiple training batches and the corresponding changes in prediction accuracy, the contribution coefficient of each feature dimension can be obtained. Assuming that the calculated contribution coefficient is [0.9, 0.8, 0.7, 0.6, 0.9], the global importance is [0.072, 0.36, 0.098, 0.018, 0.108].

[0107] Subsequently, a feature interaction graph is constructed to model the dynamic correlation between features. The local importance and the global importance are concatenated and mapped through a learnable projection matrix to obtain a feature interaction matrix. The local importance and the global importance are updated based on the feature interaction matrix to obtain a feature importance score. For example, the local importance and the global importance are concatenated to obtain [0.08, 0.45, 0.14, 0.03, 0.12, 0.072, 0.36, 0.098, 0.018, 0.108], and then mapped through a trained projection matrix to obtain a feature interaction matrix. The local importance and the global importance are then updated according to the feature interaction matrix, and finally the feature importance score is obtained, for example, [0.075, 0.42, 0.13, 0.025, 0.115].

[0108] After that, the feature importance score is input into the nonlinear mapping function to obtain the noise intensity coefficient. The nonlinear mapping function contains the product form of the hyperbolic tangent term and the sigmoid term. The hyperbolic tangent term is used to control the overall noise range, and the sigmoid term is used to construct an inverse relationship between the feature importance score and the noise intensity coefficient. For example, the feature importance score is input into the nonlinear mapping function containing the hyperbolic tangent and sigmoid functions to obtain the noise intensity coefficient, such as [0.05, 0.02, 0.06, 0.08, 0.055].

[0109] Finally, an oscillating term is introduced into the noise intensity coefficient to obtain the final noise intensity. The oscillating term modulates the feature importance score through a sine function to increase the diversity of adversarial samples. For example, the feature importance score is modulated using a sine function and added to the noise intensity coefficient as an oscillating term to obtain the final noise intensity, such as [0.052, 0.021, 0.063, 0.082, 0.057]. The generated noise is added to the original feature vector to obtain the adversarial sample.

[0110] The present invention considers the importance of features and applies noise of different intensities to different feature dimensions, thereby more effectively attacking the weak points of the model and improving the success rate of adversarial attacks. The introduction of oscillation terms increases the diversity of adversarial samples, making it more difficult for the model to defend against adversarial attacks.

[0111] In an optional embodiment,

[0112] The steps of calculating the semantic similarity of adjacent fused feature vectors in the multimodal semantic feature matrix to obtain a similarity matrix, clustering the fused feature vectors based on the similarity matrix to obtain feature clusters, setting difficulty weights according to the structural complexity, semantic consistency and boundary fuzziness of the feature clusters, sorting the feature clusters according to the difficulty weights and sequentially inputting the feature clusters into the classifier, and calculating the category probability distribution and cluster center distance for each feature cluster include:

[0113] Obtaining the local structural distance and the global semantic distance of adjacent fused feature vectors, wherein the local structural distance is calculated by an attention-weighted k-nearest neighbor graph, and the global semantic distance is calculated by a contrastive learning framework, and dynamically fusing the local structural distance and the global semantic distance by a time-varying weight coefficient to obtain a similarity matrix;

[0114] Based on the similarity matrix, local neighborhood density estimation is performed on the fused feature vector, wherein the local neighborhood density estimation adopts an adaptive bandwidth parameter, and the adaptive bandwidth parameter is obtained by multiplying the median of the neighborhood similarity by a density-dependent adjustment factor;

[0115] Calculate the node representative score based on the local neighborhood density estimation result, the node representative score is obtained by multiplying the node density by the minimum similarity of the density advantage neighbor, the density advantage neighbor is the adjacent node with a density value greater than the current node density value, determine the node with a node representative score greater than the time-varying split threshold as a split node, and recursively construct a subclass cluster for the split node to obtain a hierarchical feature cluster;

[0116] The difficulty indexes of structural complexity, semantic consistency and boundary fuzziness are calculated for the hierarchical feature clusters respectively. The structural complexity is calculated based on the information entropy of the distance distribution between samples in the cluster, the semantic consistency is calculated based on the cosine similarity of the sample to the cluster center, and the boundary fuzziness is calculated based on the minimum similarity of samples inside and outside the cluster;

[0117] The difficulty index is input into the meta-learning module to obtain the difficulty weight, the difficulty index is weighted by the difficulty weight to obtain the comprehensive difficulty score of the cluster, the clusters are sorted based on the comprehensive difficulty score to construct a difficulty sequence, classifiers are trained for clusters of different difficulty levels in the difficulty sequence, the membership of the input sample and the difficulty level is calculated, and the membership is used as the combined weight of the prediction results of classifiers at each level to obtain the category probability distribution; the difference between the sample feature vector and the cluster center vector is calculated, and the difference is multiplied by the inverse matrix of the cluster covariance matrix to obtain the cluster center distance; the final classification result is obtained based on the weighted combination of the category probability distribution and the cluster center distance.

[0118] Exemplarily, first, the similarity between the fused feature vectors is calculated and a similarity matrix is ​​constructed. Specifically, for each pair of adjacent fused feature vectors, their local structural distance and global semantic distance are calculated. The local structural distance is calculated by weighting the nodes in the K nearest neighbor graph using an attention mechanism to capture local structural information. The global semantic distance is calculated by using a contrastive learning framework to learn the representation of feature vectors in the global semantic space and calculate the distance between them. Then, the local structural distance and the global semantic distance are dynamically fused through a weight coefficient that changes over time to obtain the final similarity value. For example, in a video classification task, the visual features and audio features of a video frame can be fused and the similarity between them can be calculated.

[0119] Next, the fused feature vectors are clustered based on the similarity matrix to obtain feature clusters. Specifically, a clustering method based on local neighborhood density estimation is adopted. First, for each feature vector, its local neighborhood density is calculated. In order to adaptively adjust the bandwidth parameter, the bandwidth parameter is determined by multiplying the median of the neighborhood similarity with a density-dependent adjustment factor. For example, if the neighborhood similarity of a feature vector is high, its bandwidth parameter will be smaller, thereby capturing the local density information more finely. Then, the node representative score of each feature vector is calculated. The node representative score is calculated by multiplying the node density with the minimum similarity of its density-dominant neighbor. The density-dominant neighbor refers to the neighboring node whose density value is greater than the current node density value. Finally, nodes whose node representative scores are greater than a time-varying split threshold are determined as split nodes, and subclass clusters are recursively constructed for these split nodes to finally obtain hierarchical feature clusters.

[0120] Then, the difficulty weight is set according to the structural complexity, semantic consistency and boundary fuzziness of the feature cluster. Specifically, for each feature cluster, the three indicators of structural complexity, semantic consistency and boundary fuzziness are calculated respectively. The calculation method of structural complexity is based on the information entropy of the distance distribution between samples in the cluster. For example, if the distance distribution between samples in the cluster is more dispersed, its structural complexity is higher. The calculation method of semantic consistency is based on the cosine similarity between the sample and the cluster center. For example, if the cosine similarity between the sample in the cluster and the cluster center is high, its semantic consistency is high. The calculation method of boundary fuzziness is based on the minimum similarity between samples inside and outside the cluster. For example, if the minimum similarity between samples inside and outside the cluster is small, its boundary fuzziness is high. These indicators are input into a meta-learning module to obtain the difficulty weight of each cluster. Then, the difficulty indicators are weighted by the difficulty weight to obtain the comprehensive difficulty score of the cluster. Finally, the clusters are sorted based on the comprehensive difficulty score to construct the difficulty sequence.

[0121] Finally, the feature clusters are sorted according to the difficulty weights and input into the classifier for classification. For each feature cluster, the category probability distribution and the cluster center distance are calculated. The category probability distribution is calculated by training classifiers for clusters of different difficulty levels, calculating the membership of the input sample to each difficulty level, and using the membership as the combined weight of the prediction results of the classifiers at each level. The cluster center distance is calculated by calculating the difference between the sample feature vector and the cluster center vector and multiplying it by the inverse matrix of the cluster covariance matrix. The final classification result is obtained based on the weighted combination of the category probability distribution and the cluster center distance. For example, in an image classification task, suppose there are two clusters, representing "cat" and "dog". For an input image, if its center distance to the "cat" cluster is closer and the classifier prediction probability of the "cat" cluster is higher, then the image will be classified as "cat".

[0122] The present invention can better capture the inherent semantic structure of multimodal data by clustering based on semantic similarity, thereby improving the accuracy of classification; by adaptively adjusting the classification strategy, it can better handle clusters of different difficulty levels, thereby enhancing the robustness of classification; by sorting by difficulty weight, it can give priority to clusters with lower difficulty, thereby improving the efficiency of classification. For example, when processing large-scale data sets, this method can significantly reduce the time cost of classification.

[0123] In an optional embodiment,

[0124] The difficulty index is input into a meta-learning module to obtain a difficulty weight, the difficulty index is weighted by the difficulty weight to obtain a comprehensive difficulty score of the cluster, the clusters are sorted based on the comprehensive difficulty score to construct a difficulty sequence, and the steps of training classifiers for clusters of different difficulty levels in the difficulty sequence include:

[0125] A meta-learning network with two-layer transformation is designed, wherein the difficulty feature vector corresponding to the difficulty index is mapped to the hidden feature space through the first layer transformation to obtain the hidden layer feature, and the first layer transformation adopts the hyperbolic tangent activation function; the hidden layer feature is mapped to the weight space through the second layer transformation to obtain the difficulty weight, and the second layer transformation adopts the softmax normalization function; the dimension of the weight matrix of the first layer transformation corresponds to the dimension of the hidden layer feature, and the dimension of the weight matrix of the second layer transformation corresponds to the dimension of the difficulty index;

[0126] A first loss term is constructed based on the cross entropy loss of the classifier on the validation set of feature clustering; a regularization loss term is constructed based on the second norm of the difficulty weight and the difficulty weight difference, and the difficulty weight difference is calculated by the absolute value difference between the difficulty weights; the first loss term and the regularization loss term are weighted and combined to obtain a meta-learning loss function; the meta-learning network is trained using an alternating optimization strategy, the classifier parameters are fixed, and the transformation parameters of the meta-learning network are gradient updated based on the meta-learning loss function; the classifier parameters are retrained based on the updated difficulty weight; the meta-learning network parameter update and the classifier parameter update process are alternately executed until the performance of the validation set of feature clustering converges to obtain the final difficulty weight; the difficulty feature vector is weighted using the difficulty weight obtained by training to obtain a comprehensive difficulty score of the cluster; a smoothing process is performed based on the comprehensive difficulty scores of the cluster and the adjacent clusters to obtain a smoothed difficulty score, the clusters are sorted according to the smoothed difficulty score, and the clusters are divided into multiple difficulty levels by the difficulty threshold determined adaptively;

[0127] Classifiers are trained for different difficulty levels. For the first difficulty level, a difficulty-dependent loss weight function is introduced, and the loss weight function is calculated based on the difference between the sample difficulty and the average difficulty. For the second difficulty level, a constraint term based on the mean square error between the input features and the reconstructed features is introduced. For the third difficulty level, a comparative learning term consisting of the feature similarity of the same type of sample pairs and the feature difference of the different type of sample pairs is introduced. The integration of the classifier output results adopts a joint weighting method of difficulty membership based on a Gaussian kernel and confidence of the prediction entropy to obtain a classifier corresponding to the difficulty level.

[0128] Exemplarily, the core idea of ​​this embodiment is to use a meta-learning module to learn the difficulty weights of different feature clusters, sort the clusters according to difficulty, train classifiers for clusters of different difficulty levels separately, and finally combine the output results of multiple classifiers through an integration strategy.

[0129] First, the input data needs to be preprocessed, such as data cleaning, normalization, and other operations. Suppose there is a data set containing 1,000 samples, each with 10-dimensional features, and divided into 10 feature clusters according to pre-defined rules or domain knowledge. Then, it is necessary to define difficulty indicators, such as the variance of samples in each cluster, entropy, or cross entropy loss of classifiers. Taking variance as an example, the variance of the 10-dimensional features in each cluster is calculated to obtain 10 difficulty indicator values.

[0130] Next, a two-layer transformation meta-learning network is designed to learn difficulty weights. The first layer of transformation maps the 10-dimensional difficulty index vector to a 5-dimensional hidden feature space, using the hyperbolic tangent function as the activation function. Assuming that the dimension of the weight matrix of the first layer of transformation is 10x5, the difficulty index vector is multiplied by the weight matrix, and then the 5-dimensional hidden features are obtained through the hyperbolic tangent activation function. The second layer of transformation maps the 5-dimensional hidden features to a 10-dimensional weight space, and the softmax normalization function is used to ensure that the sum of the weights is 1. Assuming that the dimension of the weight matrix of the second layer of transformation is 5x10, the hidden features are multiplied by the weight matrix, and then the 10-dimensional difficulty weights are obtained through the softmax normalization function.

[0131] In order to train the meta-learning network, a meta-learning loss function is constructed. The loss function consists of two parts: the first part is the cross entropy loss of the classifier on the validation set divided by the feature clusters, for example, 20% of the data set is used as the validation set, and the cross entropy loss of the classifier on the validation set is calculated; the second part is the regularization loss term, which is composed of the second norm of the difficulty weight and the difference in difficulty weights, for example, the square root of the sum of the squares of the 10-dimensional difficulty weights is calculated as the second norm, and the sum of the absolute value differences between the difficulty weights is calculated as the difference in difficulty weights. The weighted combination of these two loss terms gives the final meta-learning loss function, for example, the cross entropy loss and regularization loss terms are multiplied by 0.8 and 0.2 respectively and then added.

[0132] The meta-learning network is trained using an alternating optimization strategy. First, the classifier parameters are fixed, and the transformation parameters of the meta-learning network are gradient updated based on the meta-learning loss function, such as using the gradient descent method to update the weight matrix. Then, the classifier parameters are retrained based on the updated difficulty weights, such as using the updated difficulty weights as sample weights to train the classifier. These two processes are performed alternately until the performance of the validation set of feature clustering converges, such as the accuracy of the validation set no longer improves, and the final difficulty weight is obtained.

[0133] The difficulty feature vector is weighted by the difficulty weight obtained by training to obtain the comprehensive difficulty score of the cluster. For example, the 10-dimensional difficulty index vector is multiplied by the 10-dimensional difficulty weight to obtain the comprehensive difficulty score of each cluster. Then, a smoothing process is performed based on the comprehensive difficulty scores of the cluster and the adjacent clusters, for example, the comprehensive difficulty score of each cluster is averaged with the comprehensive difficulty scores of its adjacent clusters to obtain a smoothed difficulty score. The clusters are sorted according to the smoothed difficulty scores, and the clusters are divided into multiple difficulty levels by the difficulty threshold determined adaptively, for example, the sorted clusters are divided into three difficulty levels.

[0134] Train classifiers for different difficulty levels. For the first difficulty level, introduce a difficulty-dependent loss weight function, such as calculating the loss weight of each sample based on the difference between the sample difficulty and the average difficulty. For the second difficulty level, introduce a constraint based on the mean square error between the input features and the reconstructed features, such as using an autoencoder to reconstruct the input features and adding the mean square error as a constraint to the classifier loss function. For the third difficulty level, introduce a contrastive learning term consisting of feature similarities of pairs of samples of the same class and feature differences of pairs of samples of different classes, such as using a contrastive learning loss function to train the classifier.

[0135] Finally, the output results of classifiers at different difficulty levels are integrated by using a joint weighting method of difficulty membership based on Gaussian kernel and confidence of predicted entropy. For example, the Gaussian membership of each sample belonging to each difficulty level and the predicted entropy of each classifier for each sample are calculated, and the two are multiplied as weights to perform weighted average on the output results of different classifiers to obtain the final classification result.

[0136] The present invention can effectively improve the classification performance of the model on complex data sets by training classifiers for different difficulty levels respectively and combining difficulty-aware integration strategies, especially when processing difficult samples. It can better adapt to samples of different difficulty levels in the data set, thereby enhancing the robustness and generalization ability of the model, so that it can maintain high performance even when facing new and unknown data. By sorting and grading the difficulty of clusters, more refined classification can be achieved, and subtle differences in the data can be better captured, thereby improving the accuracy and reliability of classification.

[0137] In an optional embodiment,

[0138] The steps of constructing an optimization target by using the contrast loss of sample pairs within a cluster, iteratively optimizing the prediction result of the classifier based on the optimization target, and obtaining the calculation result of the unstructured text data include:

[0139] Preprocess the prediction results of the classifier, use the local sensitive hashing method to construct the neighbor index of the samples in the cluster, calculate the cosine similarity between the sample pairs based on the neighbor index, set an adaptive similarity threshold according to the mean and standard deviation of the cosine similarity of the sample pairs in the cluster, and use the sample pairs greater than the adaptive similarity threshold as the candidate positive sample pair set;

[0140] Construct a multi-layer heterogeneous graph structure based on the candidate positive sample pair set, use samples as nodes to calculate edge weights through an attention mechanism to construct a sample-level graph, perform spectral clustering on the edge weight matrix of the sample-level graph to obtain subclass clusters and construct a subclass cluster-level graph based on the member sample interaction intensity, construct class cluster nodes based on classification labels, and construct a class cluster-level graph through the subclass cluster affiliation relationship;

[0141] Designing a multi-granularity contrast loss function for the multi-layer heterogeneous graph structure, including an intra-layer contrast loss for enhancing the consistency of representation of nodes in the same layer, an inter-layer contrast loss for cross-layer distribution alignment, and a graph structure consistency loss for ensuring hierarchical synergy, and automatically adjusting the weight coefficients of different loss items through a reinforcement learning method;

[0142] The multi-granularity contrast loss function is combined with the cross entropy loss output by the classifier to form an optimization target, and an adaptive learning rate is calculated based on the current batch sample distribution; a three-stage alternating optimization strategy is adopted, in which fixed classifier parameters optimize feature extractor parameters to minimize multi-granularity contrast loss, fixed feature extractor parameters optimize classifier parameters to minimize classification loss, and finally, parameters of feature extractor and classifier are simultaneously optimized to minimize the optimization target, and an early stopping strategy is introduced; wherein the feature representation output by the feature extractor is transferred between different optimization stages to construct a multi-layer heterogeneous graph structure and calculate a multi-granularity contrast loss function; unstructured text data is classified based on the optimized feature extractor and classifier to obtain a calculation result.

[0143] Exemplarily, first, the prediction results of the classifier are preprocessed. The unstructured text data is input into a pre-trained classifier to obtain the predicted label of each text data. For example, for a movie review dataset, the classifier can predict the sentiment tendency of each review, such as positive, negative or neutral.

[0144] Then, a local sensitive hashing method is used to construct a neighbor index of samples within the cluster. Each text data is represented as a vector, for example, using a pre-trained word embedding model. Then, a local sensitive hashing algorithm is used to map similar text vectors into the same hash bucket to construct a neighbor index. For example, the SimHash algorithm can be used to convert text vectors into binary hash codes and group them according to the first few bits of the hash code.

[0145] Next, the cosine similarity between sample pairs is calculated based on the neighbor index. For each text data, the cosine similarity between it and other text vectors in the hash bucket to which it belongs is calculated. The larger the cosine similarity value, the more similar the two text vectors are.

[0146] Set an adaptive similarity threshold based on the mean and standard deviation of the cosine similarity of pairs of samples within the cluster. Calculate the mean and standard deviation of the cosine similarity of all pairs of samples within each cluster. Then, set an adaptive similarity threshold based on the mean and standard deviation. For example, you can set the threshold to a multiple of the mean plus the standard deviation.

[0147] The sample pairs with a cosine similarity greater than the adaptive similarity threshold are selected as candidate positive sample pairs. The sample pairs with a cosine similarity greater than the adaptive similarity threshold are selected and added to the candidate positive sample pair set.

[0148] A multi-layer heterogeneous graph structure is constructed based on the set of candidate positive sample pairs. A three-layer graph structure is constructed: sample-level graph, subclass cluster-level graph, and class cluster-level graph. In the sample-level graph, each sample is taken as a node, and the edge weights between sample pairs are calculated based on the attention mechanism. The attention mechanism can consider the semantic similarity between samples and the consistency of predicted labels. In the subclass cluster-level graph, spectral clustering is performed on the edge weight matrix of the sample-level graph to obtain subclass clusters. Then, each subclass cluster is taken as a node, and the edge weights between subclass clusters are calculated based on the interaction strength between member samples, such as the average cosine similarity. In the class cluster-level graph, the class cluster corresponding to each classification label is taken as a node, and the edge weights between class clusters are calculated based on the affiliation relationship of the subclass clusters.

[0149] A multi-granularity contrast loss function is designed for multi-layer heterogeneous graph structures. Three loss functions are designed: intra-layer contrast loss, inter-layer contrast loss, and graph structure consistency loss. Intra-layer contrast loss is used to enhance the consistency of node representations in the same layer, for example, samples in the same subclass cluster are encouraged to have similar vector representations. Inter-layer contrast loss is used for cross-layer distribution alignment, for example, consistency is encouraged between node representations of sample-level graphs and subclass cluster-level graphs. Graph structure consistency loss is used to ensure hierarchical coordination, for example, the division of subclass clusters is encouraged to be consistent with the structure of class clusters. The weight coefficients of different loss terms are automatically adjusted through reinforcement learning methods. The reward function of reinforcement learning can take into account the performance of the classifier and the quality of the graph structure.

[0150] The multi-granularity contrast loss function is combined with the cross entropy loss of the classifier output to form the optimization target. The multi-granularity contrast loss and the cross entropy loss are combined in the form of a weighted sum.

[0151] Calculate the adaptive learning rate based on the distribution of the current batch of samples. For example, you can use a periodic learning rate strategy to dynamically adjust the learning rate based on the loss value of the current batch of samples.

[0152] A three-stage alternating optimization strategy is adopted. In the first stage, the classifier parameters are fixed and the feature extractor parameters are optimized to minimize the multi-granularity contrast loss. In the second stage, the feature extractor parameters are fixed and the classifier parameters are optimized to minimize the classification loss. In the third stage, the parameters of the feature extractor and the classifier are optimized simultaneously to minimize the optimization target. An early stopping strategy is introduced to stop training when the performance on the validation set no longer improves. The feature representation output by the feature extractor is passed between different optimization stages to build a multi-layer heterogeneous graph structure and calculate the multi-granularity contrast loss function.

[0153] The present invention can learn more discriminative feature representations through multi-granularity contrastive learning, thereby improving classification accuracy, especially when processing complex and noisy unstructured text data; the multi-layer heterogeneous graph structure can capture information of different granularities and enhance the robustness of the model to noise and data distribution changes through contrastive learning; the multi-layer heterogeneous graph structure can provide interpretability of classification results, for example, the contribution of different subclass clusters to the classification results can be analyzed.

[0154] In an optional embodiment,

[0155] Use the programming language Java / Scala to write ETL processing programs (data extraction, transformation and loading processing programs), use regular expressions to parse and extract unstructured text data, and convert the data into standardized semi-structured text data; store the processed semi-structured text data in a distributed storage system; use the big data computing framework to map the text data into a virtual table in memory, so that it has a structure similar to a relational database table; use the standard SQL language to calculate the data on the virtual table, and after the calculation is completed, save the results to a distributed storage system or database for subsequent use.

[0156] The present invention realizes efficient extraction and standardized processing of unstructured text data through the combined application of ETL processing program and regular expressions; solves the storage problem of large-scale unstructured data by using a distributed storage system; reduces the computational complexity of unstructured data and improves computational efficiency by mapping text data into virtual tables and using SQL for calculations; the entire solution makes full use of the advantages of the big data computing framework and is suitable for processing scenarios of large-scale unstructured text data.

[0157] Figure 2 FIG. 1 is a schematic diagram of a computing system for unstructured text data according to an embodiment of the present invention. Figure 2 As shown, the system comprises:

[0158] The first unit is used to perform multi-level processing on the input unstructured text data, dynamically adjust the word segmentation granularity according to the word frequency distribution, construct a part-of-speech co-occurrence matrix for part-of-speech tagging in combination with contextual semantic information, identify and extract entity information from the text data, fuse the part-of-speech co-occurrence matrix and the entity information to generate a hierarchical semantic label sequence, and unify the uppercase and lowercase letters and replace special characters on the hierarchical semantic label sequence;

[0159] The second unit is used to extract features from the hierarchical semantic tag sequence using feature extraction units with different convolution kernel sizes to obtain feature representation, establish an association weight matrix by calculating the cosine similarity between different semantic levels in the feature representation, construct a semantic enhancement vector based on the category attributes and hierarchical relationships of the entity information, insert a preset proportion of random noise into the feature representation to generate adversarial samples for adversarial training, and combine the trained feature representation, association weight matrix and semantic enhancement vector in series on the feature dimension to obtain a multimodal semantic feature matrix;

[0160] The third unit is used to calculate the semantic similarity of adjacent fused feature vectors in the multimodal semantic feature matrix to obtain a similarity matrix, cluster the fused feature vectors based on the similarity matrix to obtain feature clusters, set difficulty weights according to the structural complexity, semantic consistency and boundary fuzziness of the feature clusters, sort the feature clusters according to the difficulty weights and input them into the classifier in sequence, calculate the category probability distribution and cluster center distance for each feature cluster; use the contrast loss of sample pairs within the cluster to construct an optimization target, iteratively optimize the prediction results of the classifier based on the optimization target, and obtain the calculation results of unstructured text data.

[0161] According to a third aspect of the embodiments of the present invention,

[0162] An electronic device is provided, comprising:

[0163] processor;

[0164] a memory for storing processor-executable instructions;

[0165] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0166] A fourth aspect of the embodiments of the present invention is:

[0167] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0168] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.

[0169] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for calculating unstructured text data, characterized in that: include: Perform multi-level processing on the input unstructured text data, dynamically adjust the word segmentation granularity according to the word frequency distribution, build a part-of-speech co-occurrence matrix for part-of-speech tagging in combination with contextual semantic information, identify and extract entity information from the text data, fuse the part-of-speech co-occurrence matrix and the entity information to generate a hierarchical semantic label sequence, and unify the case and replace special characters on the hierarchical semantic label sequence; A feature extraction unit with different convolution kernel sizes is used to extract features from the hierarchical semantic tag sequence to obtain a feature representation, an association weight matrix is ​​established by calculating the cosine similarity between different semantic levels in the feature representation, a semantic enhancement vector is constructed based on the category attributes and hierarchical relationships of the entity information, a preset proportion of random noise is inserted into the feature representation to generate adversarial samples for adversarial training, and the trained feature representation, the association weight matrix and the semantic enhancement vector are combined in series on the feature dimension to obtain a multimodal semantic feature matrix; The semantic similarity of adjacent fused feature vectors in the multimodal semantic feature matrix is ​​calculated to obtain a similarity matrix, the fused feature vectors are clustered based on the similarity matrix to obtain feature clusters, difficulty weights are set according to the structural complexity, semantic consistency and boundary fuzziness of the feature clusters, the feature clusters are sorted according to the difficulty weights and input into the classifier in sequence, and the category probability distribution and cluster center distance are calculated for each feature cluster; the contrast loss of sample pairs within the cluster is used to construct an optimization target, and the prediction results of the classifier are iteratively optimized based on the optimization target to obtain the calculation results of unstructured text data.

2. The method according to claim 1, characterized in that The steps of performing multi-level processing on the input unstructured text data, dynamically adjusting the word segmentation granularity according to the word frequency distribution, constructing a part-of-speech co-occurrence matrix for part-of-speech tagging in combination with contextual semantic information, identifying and extracting entity information from the text data, and fusing the part-of-speech co-occurrence matrix and the entity information to generate a hierarchical semantic label sequence include: Scan the input text using a sliding window, calculate the frequency and the normalized mutual information value of the character sequence in the window; construct a window scoring function based on the frequency of the character sequence and the normalized mutual information value, the window scoring function includes a normalized mutual information value component, a frequency component and a sequence length component, and adaptively adjust the size of the sliding window according to the continuous calculation results of the window scoring function and the corresponding word frequency distribution to obtain an initial word segmentation unit; Obtaining the vector representation of the initial word segmentation unit, calculating the cosine similarity of adjacent word segmentation unit vectors to obtain context relevance, merging adjacent word segmentation units that meet the relevance merging threshold and whose combined frequency exceeds a preset minimum word frequency threshold, calculating the standardized mutual information value of the internal segmentation position for the merged word segmentation unit, performing segmentation at the position when the ratio of the standardized mutual information values ​​before and after segmentation exceeds a preset mutual information segmentation threshold, and dynamically adjusting the word segmentation granularity to obtain an optimized word segmentation unit; Based on the optimized word segmentation unit, a word frequency statistics window is constructed to record the word frequency information of the processed document, the word frequency distribution is recalculated and the mutual information segmentation threshold is adjusted based on the documents in the word frequency statistics window, and the word segmentation effect evaluation index is calculated. When the word segmentation effect evaluation index decreases, the weight of each component in the window scoring function is dynamically adjusted, and the optimized word segmentation sequence and feature information of each word segmentation unit are output, wherein the feature information includes the unit combination mutual information value, context relevance information and part of speech information; A part-of-speech co-occurrence matrix is ​​constructed based on the part-of-speech information in the feature information, and a part-of-speech transfer probability matrix is ​​generated using the transfer relationship between parts of speech. Word-level features and character features are extracted in combination with the unit combination mutual information value and context relevance information to perform entity information recognition. The part-of-speech combination pattern represented by the part-of-speech co-occurrence matrix and the part-of-speech sequence predicted by the part-of-speech transfer probability matrix are fused with the entity information recognition result to construct a multi-layer semantic label, and a semantic label sequence is generated according to the hierarchical structure of basic parts of speech, entity type and attribute information.

3. The method according to claim 1, characterized in that The step of inserting a preset proportion of random noise into the feature representation to generate adversarial samples for adversarial training includes: Generate random noise based on Gaussian distribution, and superimpose the random noise with the input feature vector to obtain an adversarial sample; During the training process, the gradient of the feature extraction loss to the feature vector is calculated to obtain a feature importance score, and the noise intensity coefficient is adjusted inversely proportionally according to the relative size of the feature importance score, where the noise intensity coefficient is inversely proportional to the feature importance score; Calculate the cosine similarity between the adversarial sample and the original feature vector to obtain feature consistency, calculate the Euclidean distance ratio between the adversarial sample and the original feature vector to obtain the perturbation amplitude, calculate the predicted output deviation between the adversarial sample and the original feature vector to obtain the predicted deviation, and include adversarial samples that simultaneously meet the feature consistency threshold, the perturbation amplitude threshold, and the predicted deviation threshold into training; Constructing a dynamically weighted adversarial training loss function, wherein the dynamically weighted adversarial training loss function includes an original sample loss term and an adversarial sample loss term, and dynamically adjusting weight coefficients of the original sample loss term and the adversarial sample loss term based on the training round and the accuracy of the validation set of feature extraction; Build an adversarial sample library, calculate the effectiveness score of the adversarial samples based on feature consistency, perturbation amplitude and prediction deviation, add adversarial samples with a score higher than the threshold to the adversarial sample library, remove the lowest-scoring samples when the adversarial sample library exceeds the preset capacity, regularly re-evaluate the effectiveness scores of samples in the adversarial sample library, and adaptively adjust the screening threshold based on the accuracy of the feature extraction verification set to update the adversarial sample library.

4. The method according to claim 3, characterized in that During the training process, the gradient of the feature extraction loss to the feature vector is calculated to obtain the feature importance score, and the noise intensity coefficient is adjusted inversely proportionally according to the relative size of the feature importance score. The step of adjusting the noise intensity coefficient inversely proportional to the feature importance score includes: Obtaining the gradient value of the feature vector to the feature extraction loss during the training process, and calculating the local importance according to the gradient value of the feature extraction loss, wherein the local importance is obtained by multiplying the absolute value of the feature gradient by the local sensitivity function, and the local sensitivity function is calculated based on the gradient difference before and after the feature perturbation; Calculating the global importance based on the local importance, the global importance is obtained by multiplying the expected value of the local importance by the feature dimension contribution coefficient, the feature dimension contribution coefficient is obtained by calculating the prediction accuracy change gradient of the feature in multiple training batches, and the prediction accuracy change gradient is determined based on the degree of influence of the fluctuation of the feature between training batches on the performance of feature extraction; Constructing a feature interaction graph to model the dynamic correlation between features, concatenating the local importance and the global importance and mapping them through a learnable projection matrix to obtain a feature interaction matrix, and updating the local importance and the global importance based on the feature interaction matrix to obtain a feature importance score; The feature importance score is input into a nonlinear mapping function to obtain a noise intensity coefficient. The nonlinear mapping function includes a product of a hyperbolic tangent term and a sigmoid term. The hyperbolic tangent term is used to control the overall noise range, and the sigmoid term is used to construct an inverse relationship between the feature importance score and the noise intensity coefficient. An oscillation term is introduced into the noise intensity coefficient to obtain a final noise intensity. The oscillation term modulates the feature importance score through a sine function to increase the diversity of adversarial samples.

5. The method according to claim 1, characterized in that The steps of calculating the semantic similarity of adjacent fused feature vectors in the multimodal semantic feature matrix to obtain a similarity matrix, clustering the fused feature vectors based on the similarity matrix to obtain feature clusters, setting difficulty weights according to the structural complexity, semantic consistency and boundary fuzziness of the feature clusters, sorting the feature clusters according to the difficulty weights and sequentially inputting the feature clusters into the classifier, and calculating the category probability distribution and cluster center distance for each feature cluster include: Obtaining the local structural distance and the global semantic distance of adjacent fused feature vectors, wherein the local structural distance is calculated by an attention-weighted k-nearest neighbor graph, and the global semantic distance is calculated by a contrastive learning framework, and dynamically fusing the local structural distance and the global semantic distance by a time-varying weight coefficient to obtain a similarity matrix; Based on the similarity matrix, local neighborhood density estimation is performed on the fused feature vector, wherein the local neighborhood density estimation adopts an adaptive bandwidth parameter, and the adaptive bandwidth parameter is obtained by multiplying the median of the neighborhood similarity by a density-dependent adjustment factor; Calculate the node representative score based on the local neighborhood density estimation result, the node representative score is obtained by multiplying the node density by the minimum similarity of the density advantage neighbor, the density advantage neighbor is the adjacent node with a density value greater than the current node density value, determine the node with a node representative score greater than the time-varying split threshold as a split node, and recursively construct a subclass cluster for the split node to obtain a hierarchical feature cluster; The difficulty indexes of structural complexity, semantic consistency and boundary fuzziness are calculated for the hierarchical feature clusters respectively. The structural complexity is calculated based on the information entropy of the distance distribution between samples in the cluster, the semantic consistency is calculated based on the cosine similarity of the sample to the cluster center, and the boundary fuzziness is calculated based on the minimum similarity of samples inside and outside the cluster; The difficulty index is input into the meta-learning module to obtain the difficulty weight, the difficulty index is weighted by the difficulty weight to obtain the comprehensive difficulty score of the cluster, the clusters are sorted based on the comprehensive difficulty score to construct a difficulty sequence, classifiers are trained for clusters of different difficulty levels in the difficulty sequence, the membership of the input sample and the difficulty level is calculated, and the membership is used as the combined weight of the prediction results of classifiers at each level to obtain the category probability distribution; the difference between the sample feature vector and the cluster center vector is calculated, and the difference is multiplied by the inverse matrix of the cluster covariance matrix to obtain the cluster center distance; the final classification result is obtained based on the weighted combination of the category probability distribution and the cluster center distance.

6. The method according to claim 5, characterized in that The difficulty index is input into a meta-learning module to obtain a difficulty weight, the difficulty index is weighted by the difficulty weight to obtain a comprehensive difficulty score of the cluster, the clusters are sorted based on the comprehensive difficulty score to construct a difficulty sequence, and the steps of training classifiers for clusters of different difficulty levels in the difficulty sequence include: A meta-learning network with two-layer transformation is designed, wherein the difficulty feature vector corresponding to the difficulty index is mapped to the hidden feature space through the first layer transformation to obtain the hidden layer feature, and the first layer transformation adopts the hyperbolic tangent activation function; the hidden layer feature is mapped to the weight space through the second layer transformation to obtain the difficulty weight, and the second layer transformation adopts the softmax normalization function; the dimension of the weight matrix of the first layer transformation corresponds to the dimension of the hidden layer feature, and the dimension of the weight matrix of the second layer transformation corresponds to the dimension of the difficulty index; A first loss term is constructed based on the cross entropy loss of the classifier on the validation set of feature clustering; a regularization loss term is constructed based on the second norm of the difficulty weight and the difficulty weight difference, and the difficulty weight difference is calculated by the absolute value difference between the difficulty weights; the first loss term and the regularization loss term are weighted and combined to obtain a meta-learning loss function; the meta-learning network is trained using an alternating optimization strategy, the classifier parameters are fixed, and the transformation parameters of the meta-learning network are gradient updated based on the meta-learning loss function; the classifier parameters are retrained based on the updated difficulty weight; the meta-learning network parameter update and the classifier parameter update process are alternately executed until the performance of the validation set of feature clustering converges to obtain the final difficulty weight; the difficulty feature vector is weighted using the difficulty weight obtained by training to obtain a comprehensive difficulty score of the cluster; a smoothing process is performed based on the comprehensive difficulty scores of the cluster and the adjacent clusters to obtain a smoothed difficulty score, the clusters are sorted according to the smoothed difficulty score, and the clusters are divided into multiple difficulty levels by the difficulty threshold determined adaptively; Classifiers are trained for different difficulty levels. For the first difficulty level, a difficulty-dependent loss weight function is introduced, and the loss weight function is calculated based on the difference between the sample difficulty and the average difficulty. For the second difficulty level, a constraint term based on the mean square error between the input features and the reconstructed features is introduced. For the third difficulty level, a comparative learning term consisting of the feature similarity of the same type of sample pairs and the feature difference of the different type of sample pairs is introduced. The integration of the classifier output results adopts a joint weighting method of difficulty membership based on a Gaussian kernel and confidence of the prediction entropy to obtain a classifier corresponding to the difficulty level.

7. The method according to claim 1, characterized in that The steps of constructing an optimization target by using the contrast loss of sample pairs within a cluster, iteratively optimizing the prediction result of the classifier based on the optimization target, and obtaining the calculation result of the unstructured text data include: Preprocess the prediction results of the classifier, use the local sensitive hashing method to construct the neighbor index of the samples in the cluster, calculate the cosine similarity between the sample pairs based on the neighbor index, set an adaptive similarity threshold according to the mean and standard deviation of the cosine similarity of the sample pairs in the cluster, and use the sample pairs greater than the adaptive similarity threshold as the candidate positive sample pair set; Construct a multi-layer heterogeneous graph structure based on the candidate positive sample pair set, use samples as nodes to calculate edge weights through an attention mechanism to construct a sample-level graph, perform spectral clustering on the edge weight matrix of the sample-level graph to obtain subclass clusters and construct a subclass cluster-level graph based on the member sample interaction intensity, construct class cluster nodes based on classification labels, and construct a class cluster-level graph through the subclass cluster affiliation relationship; Designing a multi-granularity contrast loss function for the multi-layer heterogeneous graph structure, including an intra-layer contrast loss for enhancing the consistency of representation of nodes in the same layer, an inter-layer contrast loss for cross-layer distribution alignment, and a graph structure consistency loss for ensuring hierarchical synergy, and automatically adjusting the weight coefficients of different loss items through a reinforcement learning method; The multi-granularity contrast loss function is combined with the cross entropy loss output by the classifier to form an optimization target, and an adaptive learning rate is calculated based on the current batch sample distribution; a three-stage alternating optimization strategy is adopted, in which fixed classifier parameters optimize feature extractor parameters to minimize multi-granularity contrast loss, fixed feature extractor parameters optimize classifier parameters to minimize classification loss, and finally, parameters of feature extractor and classifier are simultaneously optimized to minimize the optimization target, and an early stopping strategy is introduced; wherein the feature representation output by the feature extractor is transferred between different optimization stages to construct a multi-layer heterogeneous graph structure and calculate a multi-granularity contrast loss function; unstructured text data is classified based on the optimized feature extractor and classifier to obtain a calculation result.

8. A computing system for unstructured text data, used to implement the method according to any one of claims 1 to 7, characterized in that: include: The first unit is used to perform multi-level processing on the input unstructured text data, dynamically adjust the word segmentation granularity according to the word frequency distribution, construct a part-of-speech co-occurrence matrix for part-of-speech tagging in combination with contextual semantic information, identify and extract entity information from the text data, fuse the part-of-speech co-occurrence matrix and the entity information to generate a hierarchical semantic label sequence, and unify the uppercase and lowercase letters and replace special characters on the hierarchical semantic label sequence; The second unit is used to extract features from the hierarchical semantic tag sequence using feature extraction units with different convolution kernel sizes to obtain feature representation, establish an association weight matrix by calculating the cosine similarity between different semantic levels in the feature representation, construct a semantic enhancement vector based on the category attributes and hierarchical relationships of the entity information, insert a preset proportion of random noise into the feature representation to generate adversarial samples for adversarial training, and combine the trained feature representation, association weight matrix and semantic enhancement vector in series on the feature dimension to obtain a multimodal semantic feature matrix; The third unit is used to calculate the semantic similarity of adjacent fused feature vectors in the multimodal semantic feature matrix to obtain a similarity matrix, cluster the fused feature vectors based on the similarity matrix to obtain feature clusters, set difficulty weights according to the structural complexity, semantic consistency and boundary fuzziness of the feature clusters, sort the feature clusters according to the difficulty weights and input them into the classifier in sequence, calculate the category probability distribution and cluster center distance for each feature cluster; use the contrast loss of sample pairs within the cluster to construct an optimization target, iteratively optimize the prediction results of the classifier based on the optimization target, and obtain the calculation results of unstructured text data.

9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Science and technology project text semantic extraction and representation analysis method

    CN114254653A

  • Unstructured text data security attribute mining method and system based on multi-model collaboration

    CN117993008A