Method and device for converting natural language into command and dispatch instruction
By combining the Bert model and CNN encoder, efficient, accurate, intelligent and real-time conversion of natural language text to command and dispatch instructions is achieved, and the problem of traditional command and dispatch systems relying on manual operations is solved, automated conversion is realized, and efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202411992256.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-06-06
AI Technical Summary
Traditional command and dispatch systems rely on manual operations, are time-consuming and labor-intensive, prone to errors, and are difficult to meet the needs of modern efficient and accurate command and dispatch.
By combining the advantages of the Bert model and CNN encoder, efficient, accurate, intelligent and real-time conversion of natural language text to command and dispatch instructions is achieved. The specific steps include: using the Bert model to extract text information in natural language text, using the CNN encoder to extract feature vectors, and mapping them into the command and dispatch instructions set to determine the target command and dispatch instructions.
It realizes automatic conversion from natural language text to command and dispatch instructions, reduces manual intervention and error rates, improves processing efficiency and accuracy, meets real-time requirements, and is suitable for various scenarios that require command conversion.
Smart Images

Figure CN120106018A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data processing, and in particular to a method and device for converting natural language into command and dispatch instructions. Background Art
[0002] With the rapid development of information technology, natural language processing technology has been widely used in various fields. Especially in command and dispatch systems, natural language processing technology can effectively improve work efficiency and decision-making accuracy.
[0003] Traditional command and dispatch systems usually rely on manual operation, requiring professionals to manually input instructions or complete tasks through complex menu selections. This model is not only time-consuming and labor-intensive, but also prone to errors, and it is difficult to meet the needs of modern efficient and accurate command and dispatch.
[0004] Therefore, there is an urgent need for a more flexible and efficient method for converting natural language into command and dispatch instructions. Summary of the invention
[0005] The present application provides a method and device for converting natural language into command and dispatch instructions, which realizes efficient, accurate, intelligent and real-time conversion from natural language text to command and dispatch instructions by combining the advantages of the Bert model and the CNN encoder.
[0006] In a first aspect of the present application, a method for converting natural language into command and dispatch instructions is provided, characterized in that it is applied to an instruction conversion platform, and the method comprises: When receiving a natural language text, processing the natural language text by using a preset Bert model to extract text information from the natural language text; Processing the text information by a preset CNN encoder to extract a feature vector, and mapping the feature vector to a command and dispatch instruction set to determine a target command and dispatch instruction; The target command and dispatch instruction is output to the target object.
[0007] Optionally, the processing the natural language text by using a preset Bert model to extract text information from the natural language text includes: Performing word segmentation processing on the natural language text by using the preset Bert model to obtain a plurality of text information units, wherein the text information units include words, phrases or punctuation marks; Each of the text information units is converted into a corresponding numerical vector, and the numerical vector is processed using a multi-layer Transformer encoder in the preset Bert model to extract text information in the natural language text, wherein the text information includes entity names, entity relationships, contextual information, and semantic roles.
[0008] Optionally, the processing the text information by using a preset CNN encoder to extract a feature vector includes: Inputting the text information into the preset CNN encoder, and performing a convolution operation on the text information through the convolution layer of the preset CNN encoder to extract local features; The local features after the convolution operation are downsampled using a pooling layer to obtain a sampling result, and the sampling result is processed through a fully connected layer to obtain a feature vector.
[0009] Optionally, processing the sampling result through a fully connected layer to obtain a feature vector includes: Flatten the feature maps output by multiple pooling layers into one-dimensional vectors, and concatenate the one-dimensional vectors together to form a high-dimensional feature representation; The high-dimensional feature representation is input into one or more fully connected layers, and linear transformation and nonlinear activation are performed through weight matrices and bias items to obtain feature vectors.
[0010] Optionally, the processing the text information by using a preset CNN encoder to extract a feature vector further includes: Calculating the attention weight of each text information unit, and performing weighted summation of the attention weight and the corresponding text information unit to obtain a new text information representation; The new text information representation is processed through the fully connected layer of the preset CNN encoder to obtain a feature vector.
[0011] Optionally, mapping the feature vector to a command and dispatch instruction set to determine a target command and dispatch instruction includes: Inputting the feature vector into a softmax layer to map the feature vector into a probability distribution, each probability value in the probability distribution corresponds to an instruction in the command and dispatch instruction set; The instruction with the largest probability value in the probability distribution is selected as the target command and dispatch instruction.
[0012] Optionally, inputting the feature vector into a softmax layer to map the feature vector into a probability distribution includes: Using the feature vector as an input to a softmax layer, wherein the number of input nodes of the softmax layer matches the dimension of the feature vector; Performing a linear transformation on the feature vector using the weight matrix of the softmax layer to obtain a transformed vector; Apply the softmax function to the transformed vector to convert each component of the transformed vector into a probability value.
[0013] In a second aspect of the present application, a natural language to command and dispatch instruction conversion system is provided, characterized in that it includes a text module, an instruction module and an output module, wherein: A text module configured to, when receiving a natural language text, process the natural language text by using a preset Bert model to extract text information from the natural language text; An instruction module is configured to process the text information through a preset CNN encoder to extract a feature vector, and map the feature vector to a command and dispatch instruction set to determine a target command and dispatch instruction; An output module is configured to output the target command and dispatch instruction to the target object.
[0014] In the third aspect of the present application, an electronic device is provided, including a processor, a memory, a user interface and a network interface, the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes any one of the methods described above.
[0015] In a fourth aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores instructions, and when the instructions are executed, any of the methods described above is executed.
[0016] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. Process the natural language text through the preset Bert model. With its powerful natural language processing capabilities, the Bert model can deeply understand the semantic content of the text and accurately extract the key information in the text. This provides a solid foundation for subsequent instruction conversion and ensures the accuracy and efficiency of the conversion results; 2. The extracted text information is processed through the preset CNN (Convolutional Neural Networks) encoder to further extract feature vectors. CNN has excellent performance in processing both structured and unstructured data, and can effectively capture key features in the text. Afterwards, these feature vectors are mapped to the command and dispatch instruction set, and the target command and dispatch instruction is determined through the intelligent matching algorithm. This process not only improves the intelligence level of instruction conversion, but also makes the conversion results more in line with actual needs; 3. It realizes the automatic conversion from natural language text to command and dispatch instructions, greatly reducing manual intervention and error rate. At the same time, due to the fast processing speed of the Bert model and CNN encoder, the instruction conversion can be completed in a short time to meet the real-time requirements. This is especially important for command and dispatch scenarios that require rapid response; 4. It can be applied to various scenarios that require command conversion, such as military command, emergency dispatch, logistics management, etc. By optimizing and adjusting the preset Bert model and CNN encoder, it can adapt to the command conversion needs of different fields and scenarios, and has broad application prospects and market value. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a flowchart of the method for converting natural language into command and dispatch instructions disclosed in the embodiment of the present application; Figure 2 It is a schematic diagram of the principle of the method for converting natural language into command and dispatch instructions disclosed in the embodiment of the present application; Figure 3 It is a flow chart of model training of the method for converting natural language into command and dispatch instructions disclosed in the embodiment of the present application; Figure 4 It is another flowchart diagram of the method for converting natural language into command and dispatch instructions disclosed in the embodiment of the present application; Figure 5 It is a module schematic diagram of the natural language to command and dispatch instruction conversion system disclosed in the embodiment of the present application; Figure 6 It is a structural schematic diagram of an electronic device disclosed in an embodiment of the present application.
[0018] Explanation of the reference numerals: 501, text module; 502, instruction module; 503, output module; 601, processor; 602, communication bus; 603, user interface; 604, network interface; 605, memory. DETAILED DESCRIPTION
[0019] In order to enable technicians in this field to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments.
[0020] In the description of the embodiments of the present application, words such as "for example" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "for example" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "for example" or "for example" is intended to present related concepts in a specific way.
[0021] In the description of the embodiments of the present application, the meaning of the term "multiple" refers to two or more. For example, multiple systems refer to two or more systems, and multiple screen terminals refer to two or more screen terminals. In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. The terms "include", "comprise", "have" and their variations all mean "including but not limited to", unless otherwise specifically emphasized.
[0022] This embodiment discloses a method for converting natural language into command and dispatch instructions, which is applied to an instruction conversion platform. Figure 1 is a flow chart of the method for converting natural language into command and dispatch instructions disclosed in the embodiment of the present application, such as Figure 1 As shown, the method comprises the following steps: S101, when receiving a natural language text, processing the natural language text by using a preset Bert model to extract text information from the natural language text; S102, processing the text information by a preset CNN encoder to extract a feature vector, and mapping the feature vector to a command and dispatch instruction set to determine a target command and dispatch instruction; S103: Output the target command and dispatch instruction to the target object.
[0023] The command conversion platform includes the command and dispatch platform client, the Bert model server and the CNN encoder. The command and dispatch platform client transmits information with the Bert model server and the CNN encoder in sequence to realize the conversion from natural language to command and dispatch instructions. The data transmission between the command and dispatch platform client and the Bert model server is realized through the http interface. Before the natural language text processing on the Bert model server, it is necessary to ensure that the http interface between the command and dispatch client and the Bert model server is interoperable. The natural language text can be successfully transmitted to the Bert model server through the http interface, which is the prerequisite for the natural language text to convert command and dispatch instructions.
[0024] The platform receives a piece of natural language text. Natural language text may come from different sources, such as user input, voice conversion, or sensor data, and it contains information that needs to be understood and processed. The natural language text is fed into a pre-trained Bert (Bidirectional Encoder Representations from Transformers) model. The Bert model is a powerful language representation model that can accurately capture the semantic information in the text by understanding the contextual relationships in the text. In this step, the Bert model converts the text into a series of vector representations that capture the meaning and contextual relationships of each word or phrase in the text. Through the processing of the Bert model, the vector representation of the text is obtained, and these vectors contain important information about the text. This information is extracted and prepared for the next step of processing. The extracted text information (i.e., the vector output by the Bert model) is fed into a pre-trained CNN (Convolutional Neural Network) encoder. The CNN encoder is good at extracting local features from the data and combining these features through convolution operations to form a higher-level feature representation. In this step, the CNN encoder converts the vector output by the Bert model into a series of feature vectors that better represent the key information and patterns in the text. After obtaining the feature vectors, the system maps these vectors to a predefined set of command and dispatch instructions. This set contains all possible command and dispatch instructions, each of which is associated with a specific feature vector pattern. By comparing the feature vectors and the patterns in the instruction set, the system can determine the target command and dispatch instruction that best matches the current text information. Finally, the system outputs the determined target command and dispatch instruction to the target object. This target object may be a person (such as a commander, operator, etc.) or another automated system (such as a drone, robot, etc.). In this way, the system can quickly generate and output corresponding command and dispatch instructions based on the received natural language text.
[0025] By processing natural language texts through the preset Bert model, key information in the text can be automatically extracted without human intervention. The CNN encoder is used to extract features from the extracted text information, further improving the intelligent level of information processing. The Bert model performs well in natural language understanding tasks and can accurately capture semantic information in the text, providing a solid foundation for subsequent instruction determination. The CNN encoder effectively extracts feature vectors through convolution operations, converts complex text information into a form that is easy to classify and identify, and improves the accuracy of instruction determination. Mapping the feature vectors to the command and dispatch instruction set can quickly determine the instruction that best matches the target text, improving processing efficiency. The parameters of the Bert model and CNN encoder can be adjusted as needed to adapt to different application scenarios and text types. The command and dispatch instruction set can be expanded and updated according to actual needs to adapt to the ever-changing command and dispatch needs. Through automated and intelligent processing procedures, target command and dispatch instructions can be quickly generated, reducing the time and cost of manual decision-making. Accurate instruction generation helps improve decision-making quality and ensure the accuracy and effectiveness of command and dispatch. In an emergency, it can quickly respond and generate corresponding command and dispatch instructions, providing strong support for emergency response and crisis management.
[0026] Figure 2 It is a schematic diagram of the principle of the method for converting natural language into command and dispatch instructions disclosed in the embodiment of the present application, such as Figure 2 As shown, the conversion method includes steps S201, inputting natural language text on the client; S202, performing text segmentation processing in the Bert model; S203, capturing text information and contextual semantics; S204, performing feature capture and semantic extraction in the CNN encoder; S205, generating specific instructions and sending them to the client.
[0027] Figure 3 is a flow chart of the model training of the method for converting natural language into command and dispatch instructions disclosed in the embodiment of the present application, such as Figure 3 As shown, model training includes the following steps: S301, data set generation; S302, model training; S303, model testing; S304, tokens conversion; S305, tokens serialization; S306, transformers extract key information; S307, CNN encoder further extracts features and spatial relationships; S308, mapping specific command and dispatch instructions; when instruction recognition does not meet expectations, the model is fine-tuned and step S302 is re-executed. When instruction recognition meets expectations, the training ends.
[0028] Specifically, comprehensively count all relevant instructions that may be used in the command and dispatch platform to ensure that these instructions cover all the functions required by the platform. This can be achieved by analyzing the operation logs, user manuals, and communication with platform users of the existing system. Based on the statistically obtained instruction set, manually or automatically generate an instruction data set that can be used by the Bert model using a script. The data set should include various possible natural language expressions, which should be able to accurately map to specific instructions in the instruction set. At the same time, in order to enhance the generalization ability of the model, the data set should also contain some negative samples that are related to the instructions but not in the instruction set. Before training the Bert model, the data set needs to be preprocessed, including text cleaning, word segmentation, and stop word removal to ensure the quality of the text data input into the model. The Bert model is trained using the preprocessed data set. The training goal is to enable the model to accurately convert natural language text into corresponding command and dispatch instructions. During the training process, strategies such as cross-validation and early stopping can be used to prevent the model from overfitting. Manually input natural language text on the Bert model server to test the conversion of a single instruction to verify the basic functions of the model. Write multiple natural language texts into files for batch conversion of natural language texts to command and dispatch instructions to evaluate the performance and accuracy of the model when processing large amounts of data. After the test is completed, carefully check whether the results of the instruction conversion meet expectations. The performance of the model can be quantified by calculating indicators such as accuracy, recall, and F1 score. If the accuracy of instruction recognition does not meet expectations, the Bert model needs to be fine-tuned. This can be achieved by adding more manual data sets, enhancing data sets (such as using data enhancement technology to generate more diverse training samples), adjusting model parameters, etc. After fine-tuning, continue to train the Bert model until the expected effect of the instruction recognition result is achieved. This process may require multiple iterations, and the performance of the model needs to be re-evaluated after each iteration. Deploy the trained Bert model to the server of the command and dispatch platform for use in actual applications. During the actual operation, continuously monitor the performance of the model, including indicators such as processing speed and accuracy. If the model performance is found to be degraded or other problems occur, it is necessary to troubleshoot and repair them in a timely manner. With the development of the command and dispatch platform and the changes in user needs, the Bert model needs to be continuously updated and optimized to adapt to new application scenarios and instruction requirements.
[0029] Figure 4 is another flow chart of the method for converting natural language into command and dispatch instructions disclosed in the embodiment of the present application, such as Figure 4As shown, the method includes the following steps: S401, inputting natural language text on the client; S402, dividing the text into tokens through the Bert model; S403, token conversion; S404, token serialization; S405, transformers extracting key information; S406, CNN encoder further extracting features and spatial relationships; S407, mapping specific command and dispatch instructions, and sending them to the client.
[0030] Optionally, the processing the natural language text by using a preset Bert model to extract text information from the natural language text includes: Performing word segmentation processing on the natural language text by using the preset Bert model to obtain a plurality of text information units, wherein the text information units include words, phrases or punctuation marks; Each of the text information units is converted into a corresponding numerical vector, and the numerical vector is processed using a multi-layer Transformer encoder in the preset Bert model to extract text information in the natural language text, wherein the text information includes entity names, entity relationships, contextual information, and semantic roles.
[0031] The natural language text is segmented through the preset Bert model. Segmentation is the process of dividing continuous natural language text into independent text information units (tokens). These tokens can be words, phrases or punctuation marks, depending on the tokenizer used by the Bert model. Words: are the basic units that make up sentences, such as "cat", "dog", etc. Phrase: is a combination of multiple characters or words with a specific meaning, such as "natural language processing" or "computer". Punctuation marks: are used to separate sentences, phrases or words, such as periods, commas, quotation marks, etc. The result of word segmentation is a list of multiple tokens, each of which represents an independent information unit in the original text. Convert each token to a corresponding numerical vector. This step is achieved through the word embedding layer of the Bert model. Word embedding is a technique that maps words or phrases in a vocabulary to a high-dimensional vector space, so that semantically similar words are closer in the vector space. Each token is converted into a fixed-length numerical vector that captures the semantic information of the token. Numeric vectors can be directly input into machine learning models for calculation and reasoning. These numerical vectors are processed using the multi-layer Transformer encoder in the Bert model. Transformer is a neural network architecture based on the self-attention mechanism that can efficiently process sequence data. By stacking multiple layers of Transformer encoders, deep features in the text can be gradually extracted. The Bert model uses a bidirectional Transformer encoder that can simultaneously consider the left and right context information of each token in the text, thereby more accurately understanding the meaning of the text. The key text information in the natural language text is extracted from the processed numerical vectors. This information includes entity names, entity relationships, contextual information, and semantic roles. Entity name: a specific object or concept mentioned in the text, such as a person's name, a place name, or an institution name. Entity relationship: an association or relationship between entities, such as "Zhang San" is a friend of "Li Si". Contextual information: the context of each token in the text, which helps to understand the specific meaning of the token. Semantic role: the semantic role played by each token in the text in the sentence, such as subject, predicate, object, etc.
[0032] The Bert model can accurately segment natural language text into multiple text information units (such as words, phrases or punctuation marks), which is the basis for subsequent information extraction. The accuracy of the word segmentation results is crucial for subsequent information extraction. With its powerful context understanding ability, the Bert model can better consider the overall semantics of the text when segmenting, thereby improving the accuracy of word segmentation. Converting each text information unit into a corresponding numerical vector is a common method for machine learning models to process text data. By converting text information into numerical form, the model can perform calculations and reasoning more efficiently. Using the multi-layer Transformer encoder in the Bert model to process these numerical vectors can capture the deep features of the text, including grammatical structure, semantic relations, etc. Through the processing of the Bert model, a variety of text information in natural language text can be extracted, such as entity names, entity relations, contextual information, and semantic roles. This information is of great value for tasks such as understanding text content, performing text analysis, and building knowledge graphs. The extraction of entity names and entity relations helps to identify key entities in the text and their relationships with each other; the extraction of contextual information helps to understand the overall semantics and background of the text; and the extraction of semantic roles helps to reveal the grammatical and semantic relationships between various components in the text.
[0033] Optionally, the processing the text information by using a preset CNN encoder to extract a feature vector includes: Inputting the text information into the preset CNN encoder, and performing a convolution operation on the text information through the convolution layer of the preset CNN encoder to extract local features; The local features after the convolution operation are downsampled using a pooling layer to obtain a sampling result, and the sampling result is processed through a fully connected layer to obtain a feature vector.
[0034] The text information (such as entity names, entity relationships, contextual information, and semantic roles) processed by the BERT model is input into the preset CNN encoder. This text information has usually been converted into a numerical vector form for processing by the CNN encoder. The convolution layer of the CNN encoder is the key part of extracting text features. The convolution layer convolves the input numerical vector through multiple convolution kernels to extract local features in the text. These local features usually correspond to keywords, phrases, or specific semantic patterns in the text. In the convolution operation, each convolution kernel covers a part of the input vector and calculates the weighted sum of this part and the convolution kernel. By sliding the convolution kernel and repeating this process, local features at different positions in the input vector can be extracted. The convolution layer usually uses multiple convolution kernels of different sizes to capture features of different scales. The pooling layer is located after the convolution layer and is used to downsample the local features after the convolution operation. The purpose of downsampling is to reduce the dimension of the feature, thereby reducing the amount of calculation and improving the generalization ability of the model. The pooling layer usually uses methods such as maximum pooling or average pooling to divide the feature map output by the convolution layer into multiple small areas, and selects the maximum value or average value from each small area as the representative feature of the area. Through the downsampling operation of the pooling layer, a more compact and robust feature representation can be obtained, which is of great value for subsequent tasks such as text classification, sentiment analysis, and information extraction. After downsampling in the pooling layer, the sampled result is flattened into a one-dimensional vector and input into the fully connected layer for processing. The fully connected layer usually contains multiple neurons, each of which is connected to each element in the sampled result. The function of the fully connected layer is to combine and transform the features in the sampled result to obtain the final feature vector. In the process of extracting the feature vector, the fully connected layer can optimize the feature representation by learning weights and bias parameters. These parameters are updated and adjusted by the back propagation algorithm during the training process so that the model can better adapt to the distribution and characteristics of the input data. The feature vector processed by the fully connected layer will be output as the output result of the CNN encoder. This feature vector contains a rich feature representation of the input text information and can be used in subsequent tasks such as text classification, sentiment analysis, and information extraction.
[0035] The convolution layer of the CNN encoder can efficiently extract local features of text information through convolution operations. These local features usually correspond to specific patterns or structures in the text, such as keywords, phrases, or specific grammatical structures. The convolution operation applies multiple convolution kernels to the text information in a sliding window manner, and each convolution kernel can capture different features. This local feature extraction method makes CNN highly efficient and accurate when processing text data. The pooling layer downsamples the local features after the convolution operation to obtain the sampling results. The main purpose of this step is to reduce the number of features (i.e., dimensionality reduction), thereby reducing the complexity of subsequent calculations. Through the pooling operation, CNN can retain the most important feature information while removing redundancy and noise, which helps to improve the generalization ability and robustness of the model. The sampling results are processed through the fully connected layer to obtain a feature vector. The fully connected layer integrates the local features output by the pooling layer into global features to form a comprehensive representation of the text information. As the output of the CNN encoder, the feature vector contains the main features and semantic information of the text information, which can be used for subsequent classification, clustering, retrieval and other tasks. As a general feature extractor, the CNN encoder can adapt to different types of text data. By adjusting parameters such as the size, number, and pooling of convolution kernels, the CNN encoder can be optimized to suit specific tasks and datasets. In addition, the CNN encoder can also be combined with other deep learning models (such as RNN, LSTM, BERT, etc.) to form a more complex hybrid model to further improve the performance of text processing tasks.
[0036] Optionally, processing the sampling result through a fully connected layer to obtain a feature vector includes: Flatten the feature maps output by multiple pooling layers into one-dimensional vectors, and concatenate the one-dimensional vectors together to form a high-dimensional feature representation; The high-dimensional feature representation is input into one or more fully connected layers, and linear transformation and nonlinear activation are performed through weight matrices and bias items to obtain feature vectors.
[0037] The feature map output by each pooling layer needs to be flattened into a one-dimensional vector. This step converts the two-dimensional feature map into a one-dimensional vector form for subsequent processing. The specific method of the flattening operation is to arrange each element in the feature map into a one-dimensional array in a certain order (such as row-first or column-first). Next, the flattened one-dimensional vectors output by multiple pooling layers are spliced together to form a high-dimensional feature representation. The purpose of this step is to fuse the features extracted by different pooling layers to form a more comprehensive and richer feature representation of the input data. The high-dimensional feature representation is input to one or more fully connected layers (FC layers) for processing to obtain a feature vector. In the fully connected layer, the input high-dimensional feature representation is first linearly transformed by the weight matrix and the bias term. Each element of the weight matrix is a parameter obtained through learning, which determines the importance of the input feature to the output feature. The bias term is used to adjust the range and offset of the output feature. After the linear transformation, the output is usually also required to be nonlinearly transformed by a nonlinear activation function. Commonly used nonlinear activation functions include ReLU (Rectified Linear Unit), Sigmoid, Tanh, etc. The introduction of nonlinear activation functions enables neural networks to fit more complex nonlinear patterns and improve the expressiveness of the model. In some complex tasks, it may be necessary to use multiple layers of fully connected layers for feature extraction and transformation. Each fully connected layer receives the output of the previous layer as input and outputs a new feature representation. Through multi-layer processing, more advanced and abstract features can be gradually extracted to better complete tasks such as classification and regression. After processing by one or more layers of fully connected layers, a feature vector is finally output. This feature vector contains the global features and semantic information of the input data, which can be used for subsequent classification, clustering, retrieval and other tasks.
[0038] Flattening the feature maps output by multiple pooling layers into one-dimensional vectors and splicing these one-dimensional vectors together can form a high-dimensional feature representation. This step realizes the integration of local features, brings together the information scattered in different feature maps, and provides comprehensive input for subsequent feature extraction and classification tasks. The high-dimensional feature representation formed by splicing may contain a lot of redundant and noisy information, but the introduction of the fully connected layer can be linearly transformed through the weight matrix and bias term to achieve further dimensionality reduction and screening of features, retain the most important feature information, and improve the generalization ability of the model. The fully connected layer is usually used with nonlinear activation functions (such as ReLU, Sigmoid, Tanh, etc.) to introduce nonlinear factors so that the network can learn more complex and abstract feature combinations. This is crucial for capturing nonlinear features in text data and helps improve the accuracy and robustness of the model. The fully connected layer can adjust its output dimension according to task requirements to obtain feature vectors of different lengths. This flexibility enables the fully connected layer to adapt to different application scenarios and task requirements, such as text classification, sentiment analysis, information retrieval, etc. By inputting high-dimensional feature representations into one or more fully connected layers for processing, we can make full use of the feature integration and learning capabilities of the fully connected layers to extract more comprehensive and effective feature information. This helps improve the performance of the model in text processing tasks, such as accuracy and recall. The feature vectors obtained after processing by the fully connected layer usually have a fixed length and dimension, which facilitates subsequent processing and classification tasks. For example, the feature vector can be input into a classifier (such as a softmax classifier) for text classification, or the feature vector can be used for tasks such as text similarity calculation and clustering.
[0039] Optionally, the processing the text information by using a preset CNN encoder to extract a feature vector further includes: Calculating the attention weight of each text information unit, and performing weighted summation of the attention weight and the corresponding text information unit to obtain a new text information representation; The new text information representation is processed through the fully connected layer of the preset CNN encoder to obtain a feature vector.
[0040] The CNN encoder first extracts feature maps from the input text information. These feature maps usually contain abstract representations of various parts of the text (such as words, phrases, or sentences). Based on the extracted feature maps, the CNN encoder calculates the attention weights for each text information unit. The calculation of the attention weights usually depends on the content of the feature map and possible additional contextual information. The purpose of the attention weights is to determine the importance of different text information units (or features). In text processing, some words or phrases may be more helpful than other parts to understand the overall meaning of the text or perform subsequent classification tasks. The calculated attention weights are weighted and summed with the corresponding text information units to obtain a new text information representation. This step is essentially a re-weighting or "screening" of the original text information, so that more important information receives more attention. The new text information representation obtained by the weighted summation of the attention mechanism will be input into the fully connected layer of the CNN encoder. The main function of the fully connected layer is to further integrate and transform this information to extract more abstract and useful feature vectors. In the fully connected layer, the input new text information representation is linearly transformed through a weight matrix. This weight matrix is a learnable parameter that determines how to map the input information into the new feature space. After the linear transformation, it is usually activated by a nonlinear activation function (such as ReLU, Sigmoid, etc.) to introduce nonlinear factors and increase the expressiveness of the model. After being processed by the fully connected layer, a feature vector is finally output. This feature vector usually has a fixed length and dimension. It contains important features extracted from the original text information and can be used in subsequent classification, clustering or other tasks.
[0041] By calculating the attention weight of each text information unit, it is possible to clarify which text information is more important for feature extraction. This helps the model focus more on key information when extracting features, thereby improving the accuracy and efficiency of feature extraction. The introduction of attention weights enables the model to clarify which text information is given higher weights when extracting features, which helps explain the model's behavior and decision-making process. In practical applications, this can increase users' trust and acceptance of the model output. Since the attention mechanism can dynamically adjust according to the importance of different text information, it helps the model maintain good generalization ability when processing different texts. This means that the model can better adapt to different data sets and task requirements, thereby improving overall performance. By weighted summing the attention weights with the corresponding text information units, a new text information representation can be obtained. This representation focuses more on key information, so it helps to generate more representative feature vectors. These feature vectors can play a better role in subsequent text processing tasks, improving the accuracy and efficiency of tasks. The attention mechanism is usually combined with the convolution operation to complete the processing of text information. Since the convolution operation itself is efficient, the model combined with the attention mechanism can still maintain high computational efficiency when processing large-scale text data.
[0042] Optionally, mapping the feature vector to a command and dispatch instruction set to determine a target command and dispatch instruction includes: Inputting the feature vector into a softmax layer to map the feature vector into a probability distribution, each probability value in the probability distribution corresponds to an instruction in the command and dispatch instruction set; The instruction with the largest probability value in the probability distribution is selected as the target command and dispatch instruction.
[0043] The feature vector previously extracted by the CNN encoder is input into the softmax layer. The Softmax layer is a commonly used multi-classification output layer that can convert the input feature vector into a probability distribution. Each probability value in this probability distribution corresponds to an instruction in the command and dispatch instruction set. The Softmax layer converts the dot product result into a probability value by calculating the dot product between the input feature vector and the weight vector corresponding to each instruction, and applying the softmax function. In this way, each instruction will have a corresponding probability value, and the sum of these probability values is 1. This probability distribution reflects the model's confidence or possibility for each instruction. The instruction with the largest probability value from the probability distribution output by the softmax layer is selected as the target command and dispatch instruction. This instruction is considered to be the most likely correct decision made by the model based on the input text information.
[0044] The feature vector is input to the softmax layer, which can efficiently map the feature vector to a probability distribution. Each probability value in this probability distribution corresponds to an instruction in the command and dispatch instruction set, so that each instruction has a clear probability value corresponding to it. This probability distribution is highly interpretable and can be intuitively understood as the model's confidence or possibility for each instruction. In practical applications, this helps decision makers evaluate the rationality and feasibility of different instructions based on the size of the probability value. By selecting the instruction with the largest probability value in the probability distribution as the target command and dispatch instruction, it can be ensured that the selected instruction is the instruction that the model believes is most likely to be correct. This determination method is based on the principle of probability maximization and can reduce the risk of wrong decisions to a certain extent. At the same time, since the probability distribution output by the softmax layer is continuous, even if there are multiple instructions with high similarity, the model can accurately distinguish and select them based on the size of the probability value. As a general classifier, the Softmax layer can adapt to different command and dispatch instruction sets and feature vector representations. This means that whether it is a change in the instruction set or an update of the feature vector, the model can adapt to new task requirements by adjusting the parameters of the Softmax layer. In addition, the Softmax layer can also be used in combination with other deep learning models (such as CNN, RNN, etc.) to form a more complex hybrid model to further improve the performance of the model in command and dispatch tasks. The calculation process of the Softmax layer is relatively simple and efficient, and can complete the mapping of feature vectors to probability distributions in a short time. This helps to achieve fast response and real-time decision-making in practical applications. At the same time, the softmax layer also has good numerical stability, which can avoid numerical overflow or underflow problems during the calculation process. This helps to ensure the stability and reliability of the model.
[0045] Optionally, inputting the feature vector into a softmax layer to map the feature vector into a probability distribution includes: Using the feature vector as an input to a softmax layer, wherein the number of input nodes of the softmax layer matches the dimension of the feature vector; Performing a linear transformation on the feature vector using the weight matrix of the softmax layer to obtain a transformed vector; Apply the softmax function to the transformed vector to convert each component of the transformed vector into a probability value.
[0046] The dimension of the feature vector determines the number of input nodes of the softmax layer. The number of input nodes of the softmax layer needs to match the dimension of the feature vector. This means that if the feature vector is an n-dimensional vector, then the softmax layer needs to have n input nodes. The softmax layer contains a weight matrix whose dimension is (number of output nodes, number of input nodes). In this scenario, the number of output nodes is usually equal to the number of instructions in the command and dispatch instruction set, and the number of input nodes is equal to the dimension of the feature vector. The feature vector is linearly transformed by the weight matrix of the softmax layer. The formula for linear transformation is: y=Wx+b, where W is the weight matrix, x is the feature vector, and b is the bias term (in some embodiments, the bias term may not be included). The dimension of the transformed vector y is the same as the number of output nodes of the softmax layer. The Softmax function is a function that converts a real vector into a probability distribution. The linearly transformed vector y is used as the input of the softmax function. Apply the softmax function to each component of y to obtain a new vector p, where each component of p is a probability value between 0 and 1. The vector p forms a probability distribution, where each probability value corresponds to an instruction in the command and dispatch instruction set. The sum of the probability distributions is equal to 1. In the obtained probability distribution, the instruction with the largest probability value is selected as the target command and dispatch instruction. This means that the model believes that the instruction is the most likely correct instruction. The determination of the target instruction is based on the principle of probability maximization, which helps to reduce the risk of wrong decisions. At the same time, since the probability distribution output by the softmax function is interpretable, decision makers can evaluate the rationality and feasibility of different instructions based on the size of the probability value.
[0047] This embodiment also discloses a system for converting natural language into command and dispatch instructions. Figure 5 is a module diagram of the natural language to command and dispatch instruction conversion system disclosed in the embodiment of the present application, such as Figure 5 As shown, the system includes a text module 501, an instruction module 502 and an output module 503, wherein: The text module 501 is configured to, when receiving a natural language text, process the natural language text by using a preset Bert model to extract text information from the natural language text; The instruction module 502 is configured to process the text information through a preset CNN encoder to extract a feature vector, and map the feature vector to a command and dispatch instruction set to determine a target command and dispatch instruction; The output module 503 is configured to output the target command and dispatch instruction to the target object.
[0048] Optionally, the text module 501 is configured to: Performing word segmentation processing on the natural language text by using the preset Bert model to obtain a plurality of text information units, wherein the text information units include words, phrases or punctuation marks; Each of the text information units is converted into a corresponding numerical vector, and the numerical vector is processed using a multi-layer Transformer encoder in the preset Bert model to extract text information in the natural language text, wherein the text information includes entity names, entity relationships, contextual information, and semantic roles.
[0049] Optionally, the instruction module 502 is configured to: Inputting the text information into the preset CNN encoder, and performing a convolution operation on the text information through the convolution layer of the preset CNN encoder to extract local features; The local features after the convolution operation are downsampled using a pooling layer to obtain a sampling result, and the sampling result is processed through a fully connected layer to obtain a feature vector.
[0050] Optionally, the instruction module 502 is configured to: Flatten the feature maps output by multiple pooling layers into one-dimensional vectors, and concatenate the one-dimensional vectors together to form a high-dimensional feature representation; The high-dimensional feature representation is input into one or more fully connected layers, and linear transformation and nonlinear activation are performed through weight matrices and bias items to obtain feature vectors.
[0051] Optionally, the instruction module 502 is configured to: Calculating the attention weight of each text information unit, and performing weighted summation of the attention weight and the corresponding text information unit to obtain a new text information representation; The new text information representation is processed through the fully connected layer of the preset CNN encoder to obtain a feature vector.
[0052] Optionally, the instruction module 502 is configured to: Inputting the feature vector into a softmax layer to map the feature vector into a probability distribution, each probability value in the probability distribution corresponds to an instruction in the command and dispatch instruction set; The instruction with the largest probability value in the probability distribution is selected as the target command and dispatch instruction.
[0053] Optionally, the instruction module 502 is configured to: Using the feature vector as an input to a softmax layer, wherein the number of input nodes of the softmax layer matches the dimension of the feature vector; Performing a linear transformation on the feature vector using the weight matrix of the softmax layer to obtain a transformed vector; Apply the softmax function to the transformed vector to convert each component of the transformed vector into a probability value.
[0054] It should be noted that: when the device provided in the above embodiment realizes its function, only the division of the above functional modules is used as an example. In actual application, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0055] This embodiment also discloses an electronic device, referring to Figure 6 The electronic device may include: at least one processor 601 , at least one communication bus 602 , a user interface 603 , a network interface 604 , and at least one memory 605 .
[0056] The communication bus 602 is used to realize the connection and communication between these components.
[0057] The user interface 603 may include a display screen (Display) and a camera (Camera). The optional user interface 603 may also include a standard wired interface and a wireless interface.
[0058] The network interface 604 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface).
[0059] Among them, the processor 601 may include one or more processing cores. The processor 601 uses various interfaces and lines to connect various parts in the entire server, and executes various functions of the server and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 605, and calling data stored in the memory 605. Optionally, the processor 601 can be implemented in at least one hardware form of digital signal processing (Digital Signal Processing, DSP), field programmable gate array (Field-Programmable Gate Array, FPGA), and programmable logic array (Programmable Logic Array, PLA). The processor 601 can integrate one or a combination of a central processing unit (Central Processing Unit, CPU), a graphics processing unit (Graphics Processing Unit, GPU) and a modem. Among them, the CPU mainly processes the operating system, user interface and application programs; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 601, and it can be implemented separately through a chip.
[0060] Among them, the memory 605 may include a random access memory (Random Access Memory, RAM) and may also include a read-only memory (Read-Only Memory). Optionally, the memory 605 includes a non-transitory computer-readable storage medium. The memory 605 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 605 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned method embodiments, etc.; the data storage area may store data involved in the above-mentioned method embodiments, etc. The memory 605 may optionally be at least one storage device located away from the aforementioned processor 601. As Figure 6 As shown, the memory 605 as a computer storage medium may include an operating system, a network communication module, a user interface module, and an application program for a method of converting natural language into command and dispatch instructions.
[0061] exist Figure 6In the electronic device shown, the user interface 603 is mainly used to provide an input interface for the user and obtain data input by the user; and the processor 601 can be used to call the application program for the conversion method from natural language to command and dispatch instructions stored in the memory 605. When executed by one or more processors 601, the electronic device executes one or more methods in the above-mentioned embodiments.
[0062] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the order of the actions described, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required for the present application.
[0063] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0064] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are only schematic, such as the division of units, which is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some service interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.
[0065] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0066] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0067] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a memory 605, including several instructions for a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the various embodiments of the present application. The aforementioned memory 605 includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk.
[0068] The above is only an exemplary embodiment of the present disclosure and cannot be used to limit the scope of the present disclosure. That is, any equivalent changes and modifications made according to the teachings of the present disclosure are still within the scope of the present disclosure. After considering the disclosure of the specification, those skilled in the art will easily think of other embodiments of the present disclosure. This application is intended to cover any modification, use or adaptation of the present disclosure, which follows the general principles of the present disclosure and includes common knowledge or customary technical means in the technical field that are not recorded in the present disclosure. The description and examples are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.
Claims
1. A method for converting natural language into command and dispatch instructions, characterized in that: Applied to an instruction conversion platform, the method comprises: When receiving a natural language text, processing the natural language text by using a preset Bert model to extract text information from the natural language text; Processing the text information by a preset CNN encoder to extract a feature vector, and mapping the feature vector to a command and dispatch instruction set to determine a target command and dispatch instruction; The target command dispatch instruction is output to the target object.
2. The method for converting natural language into command and dispatch instructions according to claim 1, characterized in that: The processing of the natural language text by using a preset Bert model to extract text information from the natural language text includes: Performing word segmentation processing on the natural language text by using the preset Bert model to obtain a plurality of text information units, wherein the text information units include words, phrases or punctuation marks; Each of the text information units is converted into a corresponding numerical vector, and the numerical vector is processed using a multi-layer Transformer encoder in the preset Bert model to extract text information in the natural language text, wherein the text information includes entity names, entity relationships, contextual information, and semantic roles.
3. The method for converting natural language into command and dispatch instructions according to claim 1, characterized in that: The processing of the text information by a preset CNN encoder to extract a feature vector comprises: Inputting the text information into the preset CNN encoder, and performing a convolution operation on the text information through the convolution layer of the preset CNN encoder to extract local features; The local features after the convolution operation are downsampled using a pooling layer to obtain a sampling result, and the sampling result is processed through a fully connected layer to obtain a feature vector.
4. The method for converting natural language into command and dispatch instructions according to claim 3, characterized in that: The step of processing the sampling result through a fully connected layer to obtain a feature vector comprises: Flatten the feature maps output by multiple pooling layers into one-dimensional vectors, and concatenate the one-dimensional vectors together to form a high-dimensional feature representation; The high-dimensional feature representation is input into one or more fully connected layers, and linear transformation and nonlinear activation are performed through weight matrices and bias items to obtain feature vectors.
5. The method for converting natural language into command and dispatch instructions according to claim 1, characterized in that: The processing of the text information by a preset CNN encoder to extract a feature vector also includes: Calculating the attention weight of each text information unit, and performing weighted summation of the attention weight and the corresponding text information unit to obtain a new text information representation; The new text information representation is processed through the fully connected layer of the preset CNN encoder to obtain a feature vector.
6. The method for converting natural language into command and dispatch instructions according to claim 1, characterized in that: Mapping the feature vector to a command and dispatch instruction set to determine a target command and dispatch instruction comprises: Inputting the feature vector into a softmax layer to map the feature vector into a probability distribution, each probability value in the probability distribution corresponds to an instruction in the command and dispatch instruction set; The instruction with the largest probability value in the probability distribution is selected as the target command and dispatch instruction.
7. The method for converting natural language into command and dispatch instructions according to claim 6, characterized in that: Inputting the feature vector into a softmax layer to map the feature vector into a probability distribution comprises: Using the feature vector as an input to a softmax layer, wherein the number of input nodes of the softmax layer matches the dimension of the feature vector; Performing a linear transformation on the feature vector using the weight matrix of the softmax layer to obtain a transformed vector; Apply the softmax function to the transformed vector to convert each component of the transformed vector into a probability value.
8. A natural language to command and dispatch instruction conversion system, characterized in that: It includes text module, instruction module and output module, among which: A text module configured to, when receiving a natural language text, process the natural language text by using a preset Bert model to extract text information from the natural language text; An instruction module is configured to process the text information through a preset CNN encoder to extract a feature vector, and map the feature vector to a command and dispatch instruction set to determine a target command and dispatch instruction; An output module is configured to output the target command and dispatch instruction to the target object.
9. An electronic device, characterized in that: It includes a processor, a memory, a user interface and a network interface, the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed, the method according to any one of claims 1 to 7 is performed.
Citation Information
Cited By
Method and apparatus for converting natural language to command and dispatch instruction
WO2026144085A1