Natural language recognition processing method and device based on deep learning

By combining natural language processing technology with mixed attention mechanism, dynamic parameter module and shared encoder, the problems of high computational complexity and high training cost of multilingual models are solved, and efficient and low-cost cross-domain semantic understanding is achieved.

CN120297288AInactive Publication Date: 2025-07-11云南迅盛科技有限公司

Patent Information

Application Number
CN202510771869.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-11
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing natural language processing technology has high computational complexity, large memory usage when processing long text, poor vertical field adaptability, and high cost of multilingual model training.

Method used

The natural language processing module that combines a hybrid attention mechanism is adopted, with dynamic parameter modules and knowledge distillation technology, cross-domain knowledge migration is realized through domain adapters, and cross-language semantic understanding is supported based on a shared encoder.

Benefits of technology

Improves computing efficiency, reduces memory usage, enhances vertical field adaptability, and reduces the cost of multilingual model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297288A_ABST
    Figure CN120297288A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field related to semantic comprehension, in particular to a natural language recognition processing method and device based on deep learning, and aims at solving the problems that in the prior art, a model is large in memory occupation, low in calculation efficiency, poor in adaptability in the vertical field and high in multi-language model training cost. Specifically, the natural language processing module is provided with a dynamic parameter module which is used for adaptively adjusting the number of network layers and the number of attention heads according to the complexity of an input text so as to reduce memory occupation as much as possible and improve calculation efficiency; the vertical field adaptability is improved through the field adapter; cross-language semantic understanding is supported by sharing an encoder, and the problem that the multi-language model training cost is high is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of semantic understanding, and specifically relates to a natural language recognition and processing method and device based on deep learning. Background Art

[0002] Natural Language Processing (NLP) is an important branch in the field of artificial intelligence, which aims to enable computers to understand, generate, and process human languages. With the rapid development of the Internet and digital technologies, text data has grown explosively, and there is an urgent need for efficient and accurate natural language processing technologies.

[0003] However, with the expansion of application scenarios, some deficiencies have emerged in existing natural language processing technologies. For example, when the Transformer model processes long texts, the computational complexity increases exponentially, resulting in high memory occupancy and low computational efficiency; pre-trained models have poor adaptability in vertical domains, and the direct application effect is not good, and a large amount of labeled data is required for fine-tuning; models that support multiple languages usually require a large amount of additional corpora, and the training cost is high. Summary of the Invention

[0004] In view of this, embodiments of this application are committed to providing a natural language recognition and processing method and device based on deep learning to solve the problems of high memory occupancy, low computational efficiency, poor adaptability in vertical domains, and high training cost of multi-language models in the prior art.

[0005] The technical solution of this application provides a natural language recognition and processing method based on deep learning, including: Obtain the input text; Input the input text into a preset natural language processing module for semantic recognition; Wherein the natural language processing module is constructed by combining a hybrid attention mechanism; The natural language processing module has a dynamic parameter module for adaptively adjusting the number of network layers and the number of attention heads according to the complexity of the input text; During the training process of the natural language processing module, the knowledge distillation technology is used to compress the number of parameters of the model; The natural language processing module realizes cross-domain knowledge transfer through a domain adapter; The natural language processing module supports cross-language semantic understanding based on a shared encoder.

[0006] In some embodiments, the hybrid attention mechanism is an attention mechanism that combines local window attention and global sparse attention.

[0007] In some embodiments, the dynamic parameter module is used to adaptively adjust the number of network layers and the number of attention heads based on the complexity of the input text, such that the number of network layers and the number of attention heads are positively correlated with the complexity of the input text.

[0008] In some embodiments, the training process of the natural language processing module includes: Training a preset model to obtain a natural language processing teacher model; Constructing a natural language processing student model; the number of parameters and the computational complexity of the natural language processing student module are less than those of the natural language processing teacher model; The training objective of the natural language processing student model includes two parts: one is to fit the true labels of the data, and the other is to fit the output probability distribution of the natural language processing teacher model. The student model is optimized by minimizing the loss functions of these two parts.

[0009] In some embodiments, the natural language processing module includes: a general model and a domain adapter; The domain adapter transfers the knowledge of the general model to the target domain by learning the mapping relationship between the features of the general model and the features of the target domain; During the training process, the parameters of the general model are fixed, and only the parameters of the domain adapter are trained, so that the domain adapter can convert the features extracted by the general model into features applicable to the target domain.

[0010] In some embodiments, the shared encoder is used to map texts in different languages to a unified semantic space and obtain a feature representation that integrates cross - language information.

[0011] This application provides a natural language recognition and processing device based on deep learning, including: An acquisition module, used to acquire the input text; An identification module, used to input the input text into a preset natural language processing module for semantic identification; Wherein the natural language processing module is constructed by combining a hybrid attention mechanism; The natural language processing module has a dynamic parameter module, which is used to adaptively adjust the number of network layers and the number of attention heads according to the complexity of the input text; During the training process of the natural language processing module, the knowledge distillation technology is used to compress the number of parameters of the model; The natural language processing module realizes cross - domain knowledge transfer through a domain adapter; The natural language processing module supports cross - language semantic understanding based on a shared encoder.

[0012] This application provides an electronic device, including: A processor and a memory for storing programs executable by the processor; The processor is configured to implement the natural language recognition processing method based on deep learning as described above by running the program in the memory.

[0013] This application provides a computer-readable storage medium with a computer program stored thereon. When the computer program is run by a processor, the processor is caused to execute the natural language recognition processing method based on deep learning as described above.

[0014] This application provides a computer program product, including a computer program that implements the natural language recognition processing method based on deep learning as described above when executed by a processor.

[0015] A natural language recognition processing method based on deep learning provided by this application first obtains an input text; inputs the input text into a preset natural language processing module for semantic recognition; wherein the natural language processing module is constructed by combining a hybrid attention mechanism; the natural language processing module has a dynamic parameter module for adaptively adjusting the number of network layers and the number of attention heads according to the complexity of the input text; during the training process of the natural language processing module, a knowledge distillation technique is used to compress the number of model parameters; the natural language processing module realizes cross-domain knowledge transfer through a domain adapter; the natural language processing module supports cross-language semantic understanding based on a shared encoder. With such settings, this application has the following beneficial effects: Compared with the solutions in the prior art, the solution in this application uses a hybrid attention mechanism to better extract features; the dynamic parameter module adaptively adjusts the number of network layers and the number of attention heads according to the complexity of the input text; the complexity of the model can be adjusted based on actual needs to improve efficiency, so as to minimize memory occupation and improve computing efficiency as much as possible; the vertical domain adaptability is improved through the domain adapter; cross-language semantic understanding is supported through the shared encoder, solving the problem of high model training costs for multiple languages. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present application will become more obvious. The accompanying drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation to the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.

[0017] Figure 1 It is a flowchart of a natural language recognition processing method based on deep learning provided by an embodiment of the present application.

[0018] Figure 2It is a schematic structural diagram of a natural language recognition and processing device based on deep learning provided by an embodiment of the present application.

[0019] Figure 3 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0020] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0021] Figure 1 It is a schematic flowchart of a natural language recognition and processing method based on deep learning provided by an embodiment of the present application. As Figure 1 shown, the method includes the following contents.

[0022] Step S110, obtain the input text; This step aims to collect the text information to be processed. Such texts can come from various channels, such as queries input by users in the software interface, document content read by the system from files, or web page texts obtained by web crawlers. Through corresponding program interfaces or input modules, for example, setting a text input box in the application for users to input, or using file reading functions to read the content of text files from local storage, and text data on a remote server can also be obtained through network requests. It provides the original data input for the subsequent natural language processing process and is the starting point of the entire semantic recognition process. The quality and accuracy of the input text will directly affect the subsequent processing effect.

[0023] Step S120, input the input text into a preset natural language processing module for semantic recognition; the natural language processing module is constructed by combining a hybrid attention mechanism; the natural language processing module has a dynamic parameter module for adaptively adjusting the number of network layers and the number of attention heads according to the complexity of the input text; during the training process of the natural language processing module, the knowledge distillation technology is used to compress the number of model parameters; the natural language processing module realizes cross-domain knowledge transfer through a domain adapter; the natural language processing module supports cross-language semantic understanding based on a shared encoder.

[0024] The obtained input text is sent into a specially designed natural language processing module, which is responsible for analyzing and understanding the text at the semantic level to identify information such as the meaning and intention expressed by the text. A pre-constructed and trained natural language processing model is called, and the input text is passed to the model in a suitable format (such as a string, text vector, etc.). Inside the model, a series of complex operations and algorithms are used to perform semantic recognition on the text. This is a key step in realizing text semantic understanding. Through the processing of the natural language processing module, the input text can be converted into semantic information that can be understood and processed by a computer, thus providing a basis for subsequent various applications (such as intelligent dialogue, text classification, information extraction, etc.).

[0025] Specifically, the hybrid attention mechanism is an attention mechanism that combines local window attention and global sparse attention.

[0026] Hybrid attention mechanism: An attention mechanism that combines local window attention and global sparse attention. Principle of local window attention (Local Window Attention): The local window attention mechanism divides the input text into multiple local windows, and each window contains a fixed number of consecutive words or phrases. Inside each window, the attention mechanism calculates the correlation between words to capture local semantic information. Implementation method: The input sequence is divided according to a fixed step size and window size. For example, the window size is 10 and the step size is 5, so that each window has partial overlap with the previous window to avoid information loss. Inside each window, the attention weights between words are calculated. To improve computational efficiency and capture local semantics. Principle of global sparse attention (Global Sparse Attention): The global sparse attention mechanism selectively focuses on key words or phrases globally, reduces the computational amount through a sparse connection method, and at the same time retains important semantic information. Implementation method: Globally, selectively calculate the attention weights between some words instead of performing a fully connected calculation on all words. Function: Capture long-range dependencies and reduce computational complexity.

[0027] Combination principle of the hybrid attention mechanism: Combine local window attention and global sparse attention, both utilize local window attention to capture local semantic information and grasp the overall structure and long-range dependency relationship of the text through global sparse attention. Implementation method: In the same model layer, simultaneously use local window attention and global sparse attention, and fuse the results of the two through the multi-head attention mechanism. Function: Comprehensive semantic understanding and adaptation to diverse texts.

[0028] Application Example: Long Text Processing: When processing long news articles or academic papers, the hybrid attention mechanism can quickly capture the semantics within paragraphs and grasp the overall structure of the full text. Multilingual Text Processing: When processing multilingual texts, the hybrid attention mechanism can adapt to different language structures and unify the semantic space. Dialogue System: In a dialogue system, the hybrid attention mechanism can understand the context and grasp the dialogue topic. Through this hybrid attention mechanism that combines local window attention and global sparse attention, natural language processing models can achieve better performance in various tasks and improve the efficiency and accuracy of semantic understanding and processing. The dynamic parameter module is used to adaptively adjust the number of network layers and the number of attention heads based on the complexity of the input text, such that the number of network layers and the number of attention heads are positively correlated with the complexity of the input text.

[0029] In some embodiments, the dynamic parameter module is used to adaptively adjust the number of network layers and the number of attention heads based on the complexity of the input text, such that the number of network layers and the number of attention heads are positively correlated with the complexity of the input text.

[0030] In a natural language processing model, the dynamic parameter module is a mechanism that can automatically adjust the model structure according to the complexity of the input text. This module adaptively adjusts the number of network layers and the number of attention heads by evaluating the complexity of the input text in real time, thereby optimizing the performance and resource utilization of the model. The following will elaborate in detail from three aspects: text complexity evaluation, dynamic adjustment mechanism, and advantages.

[0031] Text Complexity Evaluation: Syntactic Structure Analysis: By analyzing the syntactic structure of the text, such as sentence length, number of clause levels, lexical diversity, etc., the syntactic complexity of the text is evaluated. For example, a text containing long sentences and multiple nested clauses usually has a high syntactic complexity.

[0032] Semantic Information Extraction: Using semantic analysis techniques, information such as keywords, entity relationships, and semantic roles in the text are extracted to evaluate the semantic richness of the text. The richer the semantic information, the higher the text complexity.

[0033] Dynamic adjustment mechanism: Adjustment of the number of network layers: According to the evaluation result of text complexity, dynamically increase or decrease the number of layers of the neural network. For complex texts, increase the number of network layers to enhance the model's expressive ability and learning ability; for simple texts, reduce the number of layers to improve computational efficiency and avoid overfitting. Adjustment of the number of attention heads: Similarly, adjust the number of attention heads according to text complexity. Complex texts require more attention heads to capture semantic associations in different aspects; simple texts can use fewer attention heads to reduce the consumption of computational resources. Adjustment algorithm: Adopt specific algorithms to achieve this dynamic adjustment, such as rule-based threshold judgment, reinforcement learning, etc. Through these algorithms, the model can respond in real time to changes in text complexity and make reasonable structural adjustments.

[0034] By adaptively adjusting the network structure, the model can reasonably allocate computational resources in the processing of texts with different complexities, avoid wasting too many resources on simple texts, and at the same time ensure that complex texts are fully processed. This enables the model to better adapt to diverse text inputs, improve semantic understanding and the accuracy of task completion. Especially when processing long texts and texts in professional fields, dynamic adjustment can enhance the model's expressive ability and learning ability, thereby improving performance. It endows the model with stronger generalization ability, enabling it to quickly adapt to different fields and types of texts without the need for retraining for each text type, improving the practicality and scalability of the model.

[0035] Through this adaptive adjustment mechanism based on the complexity of the input text, the dynamic parameter module can effectively optimize the performance and resource utilization of the natural language processing model, making it more intelligent and efficient in the face of diverse text tasks.

[0036] In some embodiments, the training process of the natural language processing module includes: training a preset model to obtain a natural language processing teacher model; constructing a natural language processing student model; the number of parameters and computational complexity of the natural language processing student module are less than those of the natural language processing teacher model; the training objective of the natural language processing student model includes two parts: one is to fit the true labels of the data, and the other is to fit the output probability distribution of the natural language processing teacher model. The student model is optimized by minimizing the loss functions of these two parts.

[0037] In the training process of the natural language processing module, knowledge distillation technology is used to compress the number of parameters of the model, thereby reducing computational complexity and resource occupancy while maintaining the model's performance. The following are the specific training steps: Train the natural language processing teacher model 1. Model Selection: Construct a large-scale pre-trained model as the teacher model, such as BERT-large, GPT-3, etc. These models are pre-trained on a large corpus and have rich semantic understanding and language generation capabilities.

[0038] Training Data Preparation: Collect and organize a large-scale labeled dataset to ensure that the data covers a wide range of fields and language styles for fully training the teacher model.

[0039] Training Process: Fine-tune the teacher model using the labeled data to optimize its performance on specific natural language processing tasks (such as semantic recognition, text classification, etc.). During the training process, the teacher model learns rich semantic representations and language knowledge.

[0040] 2. Construct a Natural Language Processing Student Model Model Design: Design a student model with a relatively simple structure to reduce its number of parameters and computational complexity. For example, smaller Transformer layers can be used, the number of attention heads can be reduced, or a lightweight neural network architecture can be adopted.

[0041] Parameter Initialization: Initialize the parameters of the student model, which can be randomly initialized or partially initialized based on the parameters of the teacher model.

[0042] 3. Training Objectives of the Natural Language Processing Student Model The training objectives of the natural language processing student model include two parts: Fitting the True Labels of the Data: The student model needs to learn the true labels of the training data to master the basic knowledge and rules of semantic recognition.

[0043] Fitting the Output Probability Distribution of the Natural Language Processing Teacher Model: The student model not only needs to learn the true labels of the data but also the output probability distribution of the teacher model, so as to inherit the semantic understanding and language knowledge of the teacher model.

[0044] 4. Optimize the Student Model Loss Function Design: Define a comprehensive loss function that combines the two parts of the training objectives. The loss function can be expressed as: L = λLtrue+(1−λ)Lteacher where Ltrue is the loss between the output of the student model and the true labels, usually using cross-entropy loss; Lteacher is the loss between the output of the student model and the output of the teacher model, which can use KL divergence loss; λ is the weight coefficient that balances the two parts of the loss and is usually determined by experimental adjustment.

[0045] Training process: The student model is trained using training data, and the parameters of the student model are optimized by minimizing the comprehensive loss function. During the training process, the student model continuously learns the true labels and the knowledge of the teacher model, gradually improving its semantic recognition performance.

[0046] 5. Evaluation and adjustment Performance evaluation: Evaluate the semantic recognition performance of the student model on the validation set, including metrics such as accuracy, recall, and F1 score. At the same time, evaluate the inference speed and resource occupancy of the model.

[0047] Model adjustment: According to the evaluation results, adjust the structure, parameters, or training strategy of the student model to further improve performance. For example, the value of λ can be adjusted to change the weight of the loss function; or the number of layers and the number of parameters of the student model can be adjusted to optimize the model complexity.

[0048] Through the above training process, the natural language processing student model can inherit the semantic understanding and language knowledge of the teacher model while maintaining a small number of parameters and low computational complexity, thus achieving efficient and accurate semantic recognition in practical applications.

[0049] In some embodiments, the natural language processing module includes: a general model and a domain adapter; the domain adapter transfers the knowledge of the general model to the target domain by learning the mapping relationship between the general model features and the target domain features; during the training process, the parameters of the general model are fixed, and only the parameters of the domain adapter are trained, enabling the domain adapter to convert the features extracted by the general model into features applicable to the target domain.

[0050] Specifically, the natural language processing module consists of a general model and a domain adapter, aiming to achieve cross-domain semantic understanding and knowledge transfer. The following is the detailed architecture and training process: 1. General model function description: The general model is a pre-trained large language model with strong semantic understanding and language generation capabilities, suitable for a variety of natural language processing tasks. Implementation method: Usually based on the Transformer architecture, through unsupervised pre-training on a large-scale corpus to learn the general features and semantic representations of the language. Role: Provide general semantic understanding and language knowledge as the basis for knowledge transfer.

[0051] 2. Field Adapter Function Description: The field adapter transfers the knowledge of the general model to a specific field by learning the mapping relationship between the general model features and the target field features, enabling the model to adapt to semantic recognition tasks in different fields. Implementation Method: The field adapter is a relatively small neural network module connected after the output layer of the general model. It learns to convert the general features extracted by the general model into features applicable to the target field. Function: Achieve cross-field knowledge transfer and reduce the adaptation cost when applying in a new field.

[0052] 3. Training Process Fix the parameters of the general model: When training the field adapter, the parameters of the general model remain fixed and are not updated. This can preserve the general semantic knowledge of the general model. Train the parameters of the field adapter: Refers to training the parameters of the field adapter so that it can convert the output features of the general model into features applicable to the target field. The training data is the labeled data of the target field. Optimization Method: Optimize the parameters by minimizing the loss function between the output of the field adapter and the true labels in the target field. Commonly used loss functions include cross-entropy loss, etc.

[0053] 4. Application Scenarios: Medical Field: Transfer the knowledge of the general model to the medical field to process medical text data, such as medical record analysis, medical literature retrieval, etc. Legal Field: Adapt to the semantic characteristics of legal texts and process tasks such as legal case analysis and legal literature retrieval. Financial Field: Conduct semantic understanding and analysis of financial texts, such as financial news classification, financial report generation, etc.

[0054] Through this architecture and training method, the natural language processing module can achieve efficient semantic understanding and knowledge transfer in different fields, reduce the development cost, and improve the generality and practicality of the model.

[0055] In some embodiments, the shared encoder is used to map texts in different languages to a unified semantic space and obtain a feature representation that integrates cross-language information.

[0056] Specifically, multi-language joint modeling aims to construct a unified model architecture that can understand and process multiple languages. Among them, the shared encoder is one of the key components to achieve this goal. The following will elaborate in detail from three aspects: the role of the shared encoder, the implementation method, and the application scenario.

[0057] Function of the shared encoder: The main function of the shared encoder is to map texts in different languages into a unified semantic space. In this space, texts in different languages with similar semantics will be represented as similar feature vectors. In this way, the model can capture the semantic associations between different languages, thus achieving cross-lingual semantic understanding. For example, when processing Chinese-English bilingual texts, the shared encoder can map the Chinese sentence "The weather is nice today" and the English sentence "The weather is nice today" to similar feature representations, enabling the model to recognize that they express the same semantic information.

[0058] Implementation method: Multi-language data collection and preprocessing: First, large-scale multi-language corpus data needs to be collected, including parallel corpora (such as bilingual sentence pairs) and non-parallel corpora (such as document collections in different languages). This data needs to go through preprocessing steps such as cleaning, tokenization, and normalization to ensure the quality and consistency of the data. Neural network architecture design: The shared encoder is usually based on deep neural network architectures such as Transformer, LSTM, etc. These architectures can effectively capture the semantic features of texts and achieve unified mapping of texts in different languages through parameter sharing. For example, the self-attention mechanism in the Transformer architecture can process texts in different languages simultaneously and learn the semantic associations between different languages. Joint training strategy: During the training process, the shared encoder is jointly trained using multi-language data. By optimizing the objective function, the model can accurately represent the semantic of texts in different languages in the unified semantic space. Supervised learning, unsupervised learning, or semi-supervised learning methods can be used during training, depending on the annotation situation of the data and the task requirements.

[0059] Application scenarios: For multi-language text classification tasks, the shared encoder can map texts in different languages into a unified semantic space, and then use a unified classifier for classification, improving the accuracy and efficiency of classification. For example, in tasks such as news classification and sentiment analysis, text data in multiple languages can be processed simultaneously. Through the shared encoder, the natural language processing model can effectively process and understand texts in multiple languages, achieve cross-lingual semantic understanding and information fusion, and greatly expand the application scope and practicality of the model.

[0060] The device embodiments of this application can be used to execute the method embodiments of this application. For details not disclosed in the device embodiments of this application, please refer to the method embodiments of this application.

[0061] Figure 2 The following is a block diagram of a natural language recognition and processing device based on deep learning provided by an embodiment of this application. As Figure 2 shown, the device includes: An acquisition module 21 for acquiring input text; An identification module 22 for inputting the input text into a preset natural language processing module for semantic recognition; Wherein, the natural language processing module is constructed by combining a hybrid attention mechanism; The natural language processing module has a dynamic parameter module for adaptively adjusting the number of network layers and the number of attention heads according to the complexity of the input text; During the training process of the natural language processing module, a knowledge distillation technique is used to compress the number of model parameters; The natural language processing module realizes cross-domain knowledge transfer through a domain adapter; The natural language processing module supports cross-language semantic understanding based on a shared encoder.

[0062] Next, refer to Figure 3 to describe an electronic device according to an embodiment of the present application. Figure 3 The block diagram of an electronic device according to an embodiment of the present application is illustrated.

[0063] As Figure 3 shown, the electronic device 300 includes one or more processors 310 and a memory 320.

[0064] The processor 310 may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 300 to perform desired functions.

[0065] The memory 320 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage media, and the processor 310 may run the program instructions to implement the natural language recognition processing method based on deep learning and / or other desired functions of various embodiments of the present application described above. Various contents such as category correspondence relationships may also be stored in the computer-readable storage media.

[0066] In one example, the electronic device 300 may further include: an input device 330 and an output device 340, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).

[0067] In addition, the input device 330 may further include, for example, a keyboard, a mouse, an interface, etc. The output device 340 can output various information to the outside, including analysis results and the like. The output device 340 may include, for example, a display, a speaker, a printer, a communication network, and a remote output device connected thereto, etc.

[0068] Of course, for the sake of simplicity, Figure 3 only some of the components related to the present application in the electronic device are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, according to specific application scenarios, the electronic device may further include any other appropriate components.

[0069] In addition to the above methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions. When the computer program instructions are run by a processor, the processor is caused to execute the steps in the deep learning-based natural language recognition processing method according to various embodiments of the present application described in the above "Exemplary Method" section of this specification.

[0070] The computer program product can be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present application. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0071] In addition, an embodiment of the present application may also be a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are run by a processor, the processor is caused to execute the steps in the deep learning-based natural language recognition processing method according to various embodiments of the present application described in the above "Exemplary Method" section of this specification.

[0072] The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0073] The foregoing description has been presented for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of the present application to the form disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.

Claims

1. A natural language recognition and processing method based on deep learning, characterized in that, including: Obtain the input text; Input the input text into a preset natural language processing module for semantic recognition; wherein the natural language processing module is constructed by combining a hybrid attention mechanism; The natural language processing module has a dynamic parameter module for adaptively adjusting the number of network layers and the number of attention heads according to the complexity of the input text; During the training process of the natural language processing module, knowledge distillation technology is used to compress the number of model parameters; The natural language processing module realizes cross-domain knowledge transfer through a domain adapter; The natural language processing module supports cross-language semantic understanding based on a shared encoder.

2. The natural language recognition processing method based on deep learning according to claim 1, wherein The hybrid attention mechanism is an attention mechanism that combines local window attention and global sparse attention.

3. The natural language recognition and processing method based on deep learning according to claim 1, characterized in that The dynamic parameter module is used to adaptively adjust the number of network layers and the number of attention heads based on the complexity of the input text, so that the number of network layers and the number of attention heads are positively correlated with the complexity of the input text.

4. The natural language recognition and processing method based on deep learning according to claim 1, characterized in that The training process of the natural language processing module includes: Train a preset model to obtain a natural language processing teacher model; Construct a natural language processing student model; the number of parameters and the computational complexity of the natural language processing student module are less than those of the natural language processing teacher model; The training objectives of the natural language processing student model include two parts: one is to fit the true label of the data, and the other is to fit the output probability distribution of the natural language processing teacher model. The student model is optimized by minimizing the loss functions of these two parts.

5. The natural language recognition processing method based on deep learning according to claim 1, characterized in that, The natural language processing module includes: a general model and a domain adapter; The domain adapter transfers the knowledge of the general model to the target domain by learning the mapping relationship between the features of the general model and the features of the target domain; During the training process, the parameters of the general model are fixed, and only the parameters of the domain adapter are trained, so that the domain adapter can convert the features extracted by the general model into features applicable to the target domain.

6. The natural language recognition and processing method based on deep learning according to claim 1, characterized in that The shared encoder is used to map texts in different languages to a unified semantic space and obtain a feature representation that integrates cross-language information.

7. A natural language recognition and processing device based on deep learning, characterized in that, including: An acquisition module for acquiring the input text; An identification module for inputting the input text into a preset natural language processing module for semantic recognition; wherein the natural language processing module is constructed by combining a hybrid attention mechanism; The natural language processing module has a dynamic parameter module for adaptively adjusting the number of network layers and the number of attention heads according to the complexity of the input text; During the training process of the natural language processing module, knowledge distillation technology is used to compress the number of model parameters; The natural language processing module realizes cross-domain knowledge transfer through a domain adapter; The natural language processing module supports cross-language semantic understanding based on a shared encoder.

8. An electronic device, characterized in that, including: A processor and a memory for storing the executable program of the processor; The processor is used to implement the deep learning-based natural language recognition processing method according to any one of claims 1 to 6 by running the program in the memory.

9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, the processor is caused to execute the deep learning-based natural language recognition processing method according to any one of claims 1 to 6.

10. A computer program product, characterized in that, It includes a computer program which, when executed by a processor, implements the deep learning-based natural language recognition processing method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Layered iteration-based long text extraction type abstract generation method and device

    CN118332101A

  • APP software defect identification method based on user comments

    CN118966167A

  • Tibetan language pre-training language model training method, system and device and medium

    CN119026608A

Cited By

  • Large language model dynamic calculation path optimization method and system

    CN121684017A