Large model vector optimization analysis system
Through HanLP cleaning, attention mechanism, MASK matrix optimization, model compression and knowledge distillation, the problem of low multi-task processing efficiency is solved, and an efficient student model adapted to different tasks is generated, which improves the performance and adaptability of the large-model vector analysis system.
Patent Information
- Application Number
- CN202510437930.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-25
AI Technical Summary
The prior art lacks work optimization during multitasking, affects model efficiency and is difficult to adapt to the needs of different types of tasks.
HanLP is used for text cleaning, attention mechanism and MASK matrix optimization are introduced, model compression and knowledge distillation are combined, and the big model vector analysis system is optimized through cycle update strategies and teaching assistant mechanisms.
The prediction accuracy and operation efficiency of the model are improved, and independent student models are generated that are adapted to the needs of different tasks, reducing computational costs and complexity.
Smart Images

Figure CN120373390A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of vector model optimization, and in particular to a large model vector optimization analysis system. Background Art
[0002] Large model vector optimization analysis is a quantitative analysis of deep learning models with huge parameter quantities and complex structures using mathematical programming and cybernetics methods. This analysis aims to find the optimal vector parameter combination through systematic analysis and mathematical modeling to improve the model's performance, reduce costs, or increase reliability. It is commonly used in the processing of natural language data models. Through large model vector optimization analysis, the prediction accuracy, generalization ability, and computational efficiency of the model are significantly improved, providing stronger support for various applications in the field of artificial intelligence, improving the user experience, and preventing users from wasting time by repeatedly asking the same question.
[0003] After retrieval, a patent with the Chinese patent publication number CN117494784A discloses a large model vector optimization analysis method, including the following steps: data preprocessing, normalizing the vector and removing noise; optimization analysis, using an optimization analysis algorithm to optimize the vector to obtain the optimal solution; post-processing, performing denormalization processing and restoring noise processing on the optimal solution, and the optimization analysis algorithm adopts any one of the stochastic gradient descent algorithm, conjugate gradient method, Newton method, and quasi-Newton method.
[0004] Although the data used by the model in the above patent is well optimized, there is a lack of work optimization during multi-task processing, and the problem that processing multiple different types of tasks will affect the model's working efficiency. Summary of the Invention
[0005] The purpose of the present invention is to solve the shortcomings existing in the prior art, and to propose a large model vector optimization analysis system.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions:
[0007] A large model vector optimization analysis system includes the following steps:
[0008] S1: Data collection step, select the optimization direction and import data from multiple sources into the system;
[0009] S2: Text feature extraction and vectorization step, use HanLP to clean the text, then use a tokenizer tool to filter out useless words, remove HTML tags, special characters, and punctuation marks to obtain standard text. If processing images or electromyogram signals, use a supervised autoencoder:
[0010] The encoder output needs to reconstruct the input signal and predict the gesture category (such as 'thank you' or 'hello') simultaneously;
[0011] The loss function includes reconstruction loss and classification loss to ensure that the features are strongly relevant to the downstream tasks, and obtain feature vectors;
[0012] S3: Attention mechanism optimization step. Without changing the total dimension of the attention mechanism, increase the number of attention heads, and apply the attention mechanism layer at multiple levels to build a deeper model;
[0013] S4: MASK matrix optimization step. Set the MASK matrix to determine which words in the text are related to each other and can "see" each other, and use the "bidirectional" feature of the MASK matrix to improve the model's understanding ability of the text when the large model generates text;
[0014] S5: Model compression step. Prune the optimized large model to reduce the number of parameters and computational complexity, and then use the TextBrewer tool for knowledge distillation to generate a BERT model with reduced number of parameters and improved speed;
[0015] S6: Feature vector concatenation step. Concatenate the feature vectors of the new text data generated by the student model with the feature vectors of the original data of the teacher model, and pass the concatenated feature vectors to the teacher model for utilization.
[0016] Preferably: In step S2, the HanLP is used for lexical analysis such as word segmentation, part-of-speech tagging, and named entity recognition, as well as sentence grammar analysis and text classification functions.
[0017] Furthermore: In step S2, the BERT model is the pre-trained BERT-base-chinese model. The processed text is input into the model. The input text is tokenized using the BERT tokenizer and converted into an integer ID sequence accepted by the model. The sequence is padded or truncated, special tokens are added, and it is converted into the Tensor form accepted by the BERT model. Then the tokenized tokens are converted into indices in the BERT model vocabulary and re-input into the BERT model to obtain the hidden state of the last layer as the feature vector of the text.
[0018] Furthermore: In step S3, an adaptive time window is added in the cross-modal attention:
[0019] Dynamically adjust the time steps of the image and electromyogram signals to align with the text sequence;
[0020] Example: The timing signals of the 'thank you' gesture with the left hand and the 'goodbye' gesture with the right hand are precisely matched with the text 'thank you' after being segmented by a time window. And through the cross-layer attention mechanism, it allows information exchange between different levels of Transformer layers.
[0021] As a preferred solution of the present invention: In step S4, the MASK matrix optimization further introduces a substitute word detection task, enabling the model to learn to distinguish between words in the original sentence and words generated by the language model.
[0022] As a further solution of the present invention: In step S5, the knowledge distillation includes using the pruned large model as the teacher model to output soft labels, and the newly generated student model fits the teacher's output. Distillation is performed in the pre-training stage to reduce the size and improve the speed. Intermediate layer knowledge is distilled to avoid overfitting.
[0023] As a still further solution of the present invention: In step S5, the system further introduces a teaching assistant mechanism. First, the teaching assistant model is used for distillation, and then the knowledge of the teaching assistant is passed to the student model.
[0024] Based on the foregoing solutions: In step S5, the system collects the running data of the student model and the text data of the user's questions during the operation of the student model, and after processing these data, re-imports them into the teacher model for cyclic update.
[0025] Based on the foregoing solutions: During the cyclic update process in step S5, the new text generated by the student model is regularized. L1 and L2 regularization are used to prevent overfitting between the newly generated text and the original data.
[0026] Based on the foregoing solutions: The system can design different student models for different tasks to meet their respective needs. The student models have different architectures, numbers of parameters, and complexities.
[0027] The beneficial effects of the present invention are as follows:
[0028] 1. A large model vector optimization analysis system, by introducing a cyclic update strategy and a teaching assistant mechanism, continuously learns and optimizes to adapt to different tasks and requirements. This mechanism enables the system to continuously update and optimize the model, improve prediction accuracy, and at the same time generate multiple independently operable student models in different fields to meet the needs of different tasks.
[0029] 2. A large model vector optimization analysis system, through model compression, including pruning and optimizing the number of parameters. At the same time, the TextBrewer tool is used for knowledge distillation to generate a BERT model with fewer parameters and faster speed. This optimization enables the system to reduce the computational cost and improve the operation efficiency while maintaining high performance.
[0030] 3. A large model vector optimization analysis system, by increasing the number of attention heads and the number of Transformer layers in cooperation with MASK matrix optimization, the system significantly improves the model's representation ability and semantic understanding ability. The introduction of the cross-layer attention mechanism further enhances the model's hierarchical representation ability, enabling the model to more accurately capture complex semantic information in the text, reducing the computational cost and improving the operation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 is a schematic diagram of the optimization process of a large model vector optimization analysis system proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0032] The technical solutions of this patent will be further described in detail below in conjunction with the specific embodiments.
[0033] The embodiments of this patent are described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below by referring to the drawings are exemplary and are only used to explain this patent and should not be construed as a limitation of this patent.
[0034] Embodiment 1:
[0035] A large model vector optimization analysis system, as Figure 1 shown, includes the following steps:
[0036] S1: Data collection: First, select the optimization direction, and then import the collected data into the system. The data sources for import can be various sources such as publicly available data on the Internet and purchased third-party data;
[0037] S2: Perform feature extraction and vectorization of the text:
[0038] First, use HanLP to clean the natural language text of the large model. HanLP can perform lexical analysis such as word segmentation, part-of-speech tagging, and named entity recognition on the natural language text in the large model, and can also provide sentence grammar analysis and text classification functions. First, install HanLP through pip, and then import the HanLP library into the Python code for text cleaning;
[0039] Subsequently, a tokenizer tool is used to filter out useless words such as "in", "is", and "of" that frequently appear in the text but have no specific meaning, reducing noise and improving analysis efficiency; subsequently, regular expressions are used to remove HTML tags, special characters, and punctuation marks from the text, and the text is normalized to obtain a standard text that is convenient for subsequent processing;
[0040] Then, feature extraction is performed on the normalized text. Information that can represent the essential characteristics of the data can be extracted. These information are usually represented in the form of feature vectors. The task of feature extraction is to find key features that have an important impact on the model performance, so as to better utilize these features in the subsequent model training and optimization processes; preferably, the BERT model of the Transformer architecture is used here to perform feature extraction on the cleaned standard text;
[0041] Preferably, the pre-trained BERT-base-chinese model is used. Subsequently, the processed text is input into the model. The tokenizer of BERT is used to tokenize the input text and convert it into an integer ID sequence that the model can accept, and this sequence is padded or truncated; subsequently, special tokens (such as [CLS] and [SEP]) are added to the tokenized data. The [CLS] token is used to represent the global features of the entire sentence. After being processed by the BERT encoder, the vector corresponding to the [CLS] token can be regarded as the vector representation of the entire sentence; and it is converted into the Tensor form that the BERT model can accept; then the tokenized tokens are converted into indices in the BERT model vocabulary, and then the converted indices are re-input back into the BERT model to obtain the hidden state of the last layer as the feature vector of the text, realizing the vectorization of the text features;
[0042] S3: Introduce the attention mechanism to optimize the model's representation ability and semantic understanding ability. Without changing the total dimension of the attention mechanism, increase the number of attention heads, and apply the attention mechanism layer at multiple levels to prevent too many parallel attention heads from affecting the model performance and computational efficiency. By increasing the number of Transformer layers, a deeper model can be constructed to capture more complex semantic information. The cross-layer attention mechanism allows information exchange between different levels of Transformer layers, thereby enhancing the model's hierarchical representation ability;
[0043] S4: Perform MASK matrix optimization: The original BERT model performs whole-word masking on word segmentation, phrases, and named entities in the MLM task and replaces them with [MASK] tokens. By introducing a substitute word detection task, the model can learn to distinguish between words in the original sentence and words generated by the language model. MASK matrix optimization can set which words in the text are related to each other and can "see" each other. These words can be related to each other under the self-attention mechanism and can utilize the "bidirectional" feature of the MASK matrix when the large model generates text, enabling BERT to consider context information simultaneously, thereby improving the model's text understanding ability.
[0044] S5: Perform model compression. First, prune the optimized large model for attention heads and remove unimportant attention heads. This method can reduce the number of model parameters and computational complexity while maintaining good performance.
[0045] Subsequently, use the TextBrewer tool for knowledge distillation. Take the pruned large model as the teacher model to output soft labels, and the newly generated student model fits the teacher's output. Perform distillation during the pre-training stage to reduce the size and improve the speed. Distill the knowledge of the intermediate layer to avoid overfitting. Combine the distillation in the pre-training and fine-tuning stages to obtain a BERT model with reduced number of parameters and improved speed, whose effect is close to BERT-base, and reduce the dimension of each layer, retain 24 layers, reduce parameters, and improve speed. The effect is better than TinyBERT and DistillBERT. Utilize the distillation Value-Value matrix to enable the student model to imitate the teacher model more deeply. Introduce a teaching assistant mechanism. First, distill with the teaching assistant model and then transfer the knowledge of the teaching assistant to the student. The student model learns the generalization ability of the teacher model. Theoretically, the result is better than that of a student model that simply fits the training data, and multiple independent student models that can run in different fields can be generated. For different tasks, different student models can be designed to meet their respective needs. The student models can have different architectures, numbers of parameters, and complexities to adapt to the characteristics and resource limitations of different tasks. Each student model can be optimized for its specific task and reduce the model size while maintaining high performance.
[0046] During the operation of the student model, collect the running data of the student model and the text data of the questions asked by users. After collecting these data, repeat the steps of S1 - S2 for processing, and then re-import them into the teacher model. After processing and optimization, use the new teacher model to train and fit the corresponding student model for cyclic update. During this process, perform regularization on the new text generated by the student model. Use L1 and L2 regularization to prevent overfitting between the newly generated text and the original data, resulting in data hallucination, and manually adjust the problems with too high frequency of repeated questions from users to improve the overall prediction accuracy.
[0047] S6: Feature vector concatenation: Concatenate the vectors generated from the new text data produced by the student model and the original data vectors of the teacher model, and pass the concatenated feature vectors to the teacher model for utilization.
[0048] In the text feature extraction and vectorization stage, the system uses HanLP for text cleaning, including key steps such as word segmentation, part-of-speech tagging, and named entity recognition, to ensure the accuracy and standardization of the text data. Subsequently, the system filters out meaningless words, removes HTML tags, special characters, and punctuation marks to obtain standard text, and uses the advanced BERT model to extract features from the standard text, efficiently converting the text information into feature vectors to provide strong support for subsequent analysis and processing.
[0049] To further optimize the model, the system introduces an attention mechanism. By increasing the number of attention heads and the number of Transformer layers, the representational ability and semantic understanding ability of the model are significantly improved. At the same time, the introduction of the cross-layer attention mechanism enhances the hierarchical representation ability of the model. In addition, through MASK matrix optimization and the introduction of substitute word detection tasks, BERT can consider context information simultaneously, further improving the text understanding ability.
[0050] At the same time, to reduce the complexity and computational cost of the model, the system performs model compression, including pruning and optimizing the number of parameters. At the same time, the TextBrewer tool is used for knowledge distillation to generate a BERT model with fewer parameters and faster speed. In addition, the system also introduces a teaching assistant mechanism and a cyclic update strategy, and adapts to different tasks and requirements through continuous learning and optimization.
[0051] The above is a preferred specific implementation manner of the present invention. The protection scope of the present invention is not limited thereto. Any modifications, equivalent substitutions, and improvements made by those skilled in the art within the technical scope disclosed by the present invention in combination with the prior art or common knowledge, within the spirit and principles of the present invention, shall be covered by the protection scope of the present invention.
Claims
1. A large model vector optimization analysis system, characterized in that, It includes the following steps: S1: Data collection step, select the optimization direction and import data from multiple sources into the system; S2: Text feature extraction and vectorization step, use HanLP to clean the text, then use a tokenizer tool to filter out useless words, remove HTML tags, special characters and punctuation marks to obtain standard text, and then use the BERT model to extract features from the standard text to obtain feature vectors; S3: Attention mechanism optimization step, without changing the total dimension of the attention mechanism, increase the number of attention heads, and apply the attention mechanism layer at multiple levels to build a deeper model; S4: MASK matrix optimization step, set the MASK matrix to determine which words in the text are related to each other and "see" each other, and use the "bidirectional" feature of the MASK matrix to improve the model's understanding ability of the text when the large model generates text; S5: Model compression step, prune the optimized large model to reduce the number of parameters and computational complexity, and then use the TextBrewer tool for knowledge distillation to generate a BERT model with reduced number of parameters and improved speed; S6: Feature vector concatenation step, non-linearly fusing the feature vectors of the student model and the teacher model through a cross-modal Transformer, and dynamically aligning the dependencies of the image, EMG signal, and text modalities using the attention mechanism system , and passing the concatenated feature vectors to the teacher model for utilization.
2. The large model vector optimization analysis system according to claim 1, characterized in that, In step S2, the HanLP is used for lexical analysis such as word segmentation, part-of-speech tagging, and named entity recognition, as well as sentence grammar analysis and text classification functions.
3. The large model vector optimization analysis system according to claim 2, wherein In step S2, the BERT model is the pre-trained BERT-base-chinese model. The processed text is input into the model, and the input text is tokenized using the BERT tokenizer and converted into an integer ID sequence accepted by the model. The sequence is padded or truncated, special tokens are added, and it is converted into the Tensor form accepted by the BERT model. Then the tokenized tokens are converted into indices in the BERT model vocabulary and re-input back into the BERT model to obtain the hidden state of the last layer as the feature vector of the text.
4. A large model vector optimization analysis system according to claim 1, characterized in that, In step S3, the number of Transformer layers is increased to capture more complex semantic information, and contrastive learning (InfoNCE Loss) is introduced in cross-layer attention: Positive sample pair: Text features with the same semantics and the output of the teacher model (such as the gesture of 'thank you' corresponding to the text 'thank you'); Negative sample pair: Text features with different semantics and random noise (such as the gesture of 'thank you' and the text'reject'); By maximizing the similarity of positive samples and minimizing the correlation of negative samples.
5. The large model vector optimization analysis system according to claim 4, characterized in that In step S4, the MASK matrix optimization also introduces a substitute word detection task, enabling the model to learn to distinguish between the words in the original sentence and the words generated by the language model.
6. The large model vector optimization analysis system according to claim 4, characterized in that, In step S5, the knowledge distillation includes using the pruned large model as the teacher model to output soft labels, and the newly generated student model fits the teacher output. Distillation is carried out in the pre-training stage to reduce the size and improve the speed, and the knowledge of the intermediate layer is distilled to avoid overfitting.
7. The large model vector optimization analysis system according to claim 6, wherein In step S5, the system also introduces a teaching assistant mechanism, first distill with the teaching assistant model, and then transfer the knowledge of the teaching assistant to the student model.
8. The large model vector optimization analysis system according to claim 7, wherein, In step S5, the system collects the running data of the student model and the text data of the user's questions during the operation of the student model, and after processing these data, re-imports them into the teacher model for cyclic update.
9. The large model vector optimization analysis system according to claim 8, characterized in that, In the cyclic update process in step S5, regularization processing is performed on the new text generated by the student model, and L1 and L2 regularization processing are used to prevent overfitting between the newly generated text and the original data.
10. A large model vector optimization analysis system according to any one of claims 1-9, characterized in that, The system can design different student models for different tasks to meet their respective needs, and the student models have different architectures, numbers of parameters, and complexities.
Citation Information
Patent Citations
Large model vector optimization analysis method
CN117494784A