Information Processing Method, Apparatus, Electronic Device, and Computer Readable Storage Medium
By using the word segmentation dictionary and word vector library of the target field in natural language processing, word segmentation and characterization processing of the text information to be processed, the problem of low accuracy of text information processing in the prior art is solved, and higher processing accuracy and effect are achieved.
Patent Information
- Application Number
- CN202110432639.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-21
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-04-21
AI Technical Summary
The prior art uses natural language processing (NLP) technology to process text information to be processed, with low accuracy and poor results.
An information processing method is provided, by obtaining the pending text information of the target field, performing word segmentation processing based on the word segmentation dictionary of the target field, finding the target word vectors in the word vector library of the target field, and processing the text information based on these word vectors.
The accuracy of word segmentation of text information in the target field is improved, and the accuracy and effect of text processing is improved through more accurate word vector representation.
Smart Images

Figure CN113761117B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to an information processing method, apparatus, electronic device, and computer-readable storage medium. Background Art
[0002] With the rapid development of artificial intelligence technology, the application scope of artificial intelligence (AI) technology is becoming wider and wider. Artificial intelligence technology includes natural language processing (NLP) technology, and NLP technology is widely used in the analysis and processing of text information. However, currently, the accuracy of using NLP technology to process the text information to be processed is relatively low, and the effect is poor. Summary of the Invention
[0003] This application provides an information processing method, apparatus, electronic device, and computer-readable storage medium, which can solve the problem of relatively low accuracy and poor effect in processing the text information to be processed. The technical solutions are as follows:
[0004] On the one hand, an information processing method is provided, and the method includes:
[0005] Obtain the text information to be processed in the target field;
[0006] Based on the word segmentation dictionary in the target field, perform word segmentation processing on the text information to be processed to obtain multiple words in the text information to be processed;
[0007] In the word vector library in the target field, search for the target word vectors of the multiple words, and the target word vectors are used to represent the semantics of the corresponding words;
[0008] Based on the target word vectors of the target words found in the multiple words, perform text processing on the text information to be processed;
[0009] Wherein, the word vector library in the target field includes: the target word vectors of multiple target words belonging to the target field; the target word vector of each target word in the word vector library is determined based on the model parameters of the neural network model, and the model parameters are obtained by training the neural network model based on the training text information in the target field to update the initial model parameters of the neural network model; the training text information includes the multiple target words, and the initial model parameters are determined based on the initial word vectors configured for each target word in the multiple target words, and the dimension of the initial word vectors is less than the number of the multiple target words.
[0010] Optionally, the method further includes:
[0011] Obtain the training text information of the target field;
[0012] Based on the word segmentation dictionary of the target field, perform word segmentation on the training text information to obtain the multiple target words;
[0013] Configure the initial word vectors for each target word among the multiple target words;
[0014] Obtain a training sample set based on the training text information, where the training sample set includes multiple training samples and the training labels of each training sample; each training sample is used to reflect the semantics of the target words in a sample text, and the sample text is a sentence in the training text information that includes at least two target words; the training label of the training sample is used to reflect the positions of the target words in the sample text among the multiple target words;
[0015] Train the neural network model based on each training sample and the corresponding training label in the training sample set to update the initial word vectors of the target words in the sample text corresponding to the training sample among the model parameters of the neural network model, and obtain the trained neural network model;
[0016] Based on the model parameters of the trained neural network model, obtain the target word vectors of the multiple target words to construct the word vector library of the target field.
[0017] Optionally, the configuring the initial word vectors for each target word among the multiple target words includes:
[0018] Based on the dimension of the input initial word vectors, randomly initialize the word vectors for each target word among the multiple target words to configure the initial word vectors for each target word.
[0019] Optionally, the training the neural network model based on each training sample and the corresponding training label in the training sample set to update the initial word vectors of the target words in the sample text corresponding to the training sample among the model parameters of the neural network model includes:
[0020] Input each training sample in the training sample set into the neural network model, where each training sample includes the initial word vectors of 2m surrounding words of one target word among the multiple target words; where m≥1, and the 2m surrounding words include: among the sample text corresponding to the training sample, the m target words that are before the target word and closest to the one target word, and the m target words that are after each target word and closest to the one target word;
[0021] Based on the initial word vectors of the 2m surrounding words of the one target word, the reference word vector of the one target word is output through the neural network model; the reference word vector of the target word is used to reflect: the probabilities that the words represented by the initial word vectors of the 2m surrounding words of the target word are each of the multiple target words;
[0022] When it is determined that the representation probabilities corresponding to each of the multiple target words are all greater than or equal to the probability threshold based on the reference word vectors of each of the multiple target words and the training labels in the training set, stop training the neural network model; the representation probability corresponding to the target word is: the probability that the word represented by the initial word vectors of the 2m surrounding words of the target word is the target word;
[0023] When it is determined that the representation probability corresponding to the one target word is less than the probability threshold based on the reference word vector of the one target word and the training label corresponding to the one target word, update the initial word vectors of the 2m surrounding words of the one target word among the model parameters of the neural network model; return to execute the step of outputting the reference word vector of the one target word through the neural network model based on the initial word vectors of the 2m surrounding words of the one target word.
[0024] Optionally, the outputting the reference word vector of the one target word through the neural network model based on the initial word vectors of the 2m surrounding words of the one target word includes:
[0025] Through the neural network model, the vector obtained by adding and then averaging the initial word vectors of the 2m surrounding words of the one target word is determined as the hidden vector of the one target word;
[0026] Through the neural network model, the vector obtained by multiplying the hidden vector of the one target word by the output word matrix is determined as the reference word vector of the one target word, and the output word matrix is used to reflect the mapping relationship from the 2m surrounding words to the one target word.
[0027] Optionally, the training label of the training sample includes a benchmark word vector set for the one target word, and the benchmark word vector is used to reflect the position of the one target word among the multiple target words;
[0028] The training of the neural network model based on each training sample and the corresponding training label in the training sample set to update the initial word vectors of the target words in the sample text corresponding to the training samples among the model parameters of the neural network model further includes:
[0029] Determine the loss value corresponding to each target word based on the reference word vector of the one target word and the benchmark word vector set for the one target word, where the loss value is used to reflect the degree of difference between the reference word vector and the benchmark word vector;
[0030] When the loss value corresponding to the one target word is less than the loss value threshold, determine that the representation probability corresponding to the one target word is less than the probability threshold;
[0031] When determining that the representation probability corresponding to the one target word is less than the probability threshold, update the initial word vectors of the 2m surrounding words of the one target word in the model parameters of the neural network model, including:
[0032] Based on the loss value corresponding to the one target word, update the initial word vectors of the 2m surrounding words of the one target word and the output word matrix.
[0033] Optionally, before performing word segmentation processing on the to-be-processed text information based on the word segmentation dictionary of the target domain, the method further includes:
[0034] Obtain multiple entity names of the target domain;
[0035] Construct the word segmentation dictionary based on the multiple entity names.
[0036] On the other hand, an information processing device is provided, and the information processing device includes:
[0037] An acquisition module, configured to acquire to-be-processed text information of a target domain;
[0038] A first processing module, configured to perform word segmentation processing on the to-be-processed text information based on the word segmentation dictionary of the target domain to obtain multiple words in the to-be-processed text information;
[0039] A lookup module, configured to look up target word vectors of the multiple words in a word vector library of the target domain, where the target word vectors are used to represent the semantics of the corresponding words;
[0040] A second processing module, configured to perform text processing on the to-be-processed text information based on the target word vectors of the target words found in the multiple words;
[0041] Among them, the word vector library of the target domain includes: target word vectors of multiple target words belonging to the target domain; the target word vectors of each target word in the word vector library are determined based on the model parameters of a neural network model, and the model parameters are obtained after training the neural network model based on the training text information of the target domain to update the initial model parameters of the neural network model; the training text information includes the multiple target words, and the initial model parameters are determined based on the initial word vectors configured for each of the multiple target words, and the dimension of the initial word vectors is less than the number of the multiple target words.
[0042] Optionally, the information processing device further includes:
[0043] A second acquisition module, configured to acquire the training text information of the target domain;
[0044] A third processing module, configured to perform word segmentation processing on the training text information based on the word segmentation dictionary of the target domain to obtain the multiple target words;
[0045] A configuration module, configured to configure the initial word vectors for each of the multiple target words;
[0046] A fourth acquisition module, configured to acquire a training sample set based on the training text information, where the training sample set includes multiple training samples and training labels of each training sample; each training sample is used to reflect the semantics of the target words in a sample text, and the sample text is a sentence including at least two target words in the training text information; the training label of the training sample is used to reflect the positions of the target words in the sample text among the multiple target words;
[0047] A training module, configured to train the neural network model based on each training sample and the corresponding training label in the training sample set to update the initial word vectors of the target words in the sample text corresponding to the training samples in the model parameters of the neural network model, so as to obtain a trained neural network model;
[0048] A first construction module, configured to obtain the target word vectors of the multiple target words based on the model parameters of the trained neural network model, so as to construct the word vector library of the target domain.
[0049] Optionally, the configuration module is further configured to:
[0050] Based on the dimension of the input initial word vectors, randomly initialize the word vectors for each of the multiple target words to configure the initial word vectors for each of the target words.
[0051] Optionally, the training module is further configured to:
[0052] Input each training sample in the training sample set into the neural network model, where each training sample includes the initial word vectors of 2m surrounding words of one target word among the multiple target words; where m≥1, and the 2m surrounding words include: among the sample texts corresponding to the training samples, the m target words closest to the one target word before the target word, and the m target words closest to the one target word after each target word;
[0053] Based on the initial word vectors of the 2m surrounding words of the one target word, output a reference word vector of the one target word through the neural network model; the reference word vector of the target word is used to reflect: the probability that the words represented by the initial word vectors of the 2m surrounding words of the target word are each target word among the multiple target words;
[0054] When it is determined that the representation probability corresponding to each target word among the multiple target words is greater than or equal to a probability threshold based on the reference word vectors of each target word among the multiple target words and the training labels in the training set, stop training the neural network model; the representation probability corresponding to the target word is: the probability that the words represented by the initial word vectors of the 2m surrounding words of the target word are the target word;
[0055] When it is determined that the representation probability corresponding to the one target word is less than the probability threshold based on the reference word vector of the one target word and the training label corresponding to the one target word, update the initial word vectors of the 2m surrounding words of the one target word in the model parameters of the neural network model; return to execute the step of outputting a reference word vector of the one target word through the neural network model based on the initial word vectors of the 2m surrounding words of the one target word.
[0056] Optionally, the training module is further configured to:
[0057] Determine, through the neural network model, the vector obtained by adding and then averaging the initial word vectors of the 2m surrounding words of the one target word as the hidden vector of the one target word;
[0058] Determine, through the neural network model, the vector obtained by multiplying the hidden vector of the one target word by an output word matrix as the reference word vector of the one target word, where the output word matrix is used to reflect the mapping relationship from the 2m surrounding words to the one target word.
[0059] Optionally, the training label of the training sample includes a benchmark word vector set for the one target word, and the benchmark word vector is used to reflect the position of the one target word among the multiple target words;
[0060] The training module is further configured to:
[0061] Based on the reference word vector of the one target word and the benchmark word vector set for the one target word, determine the loss value corresponding to each target word, and the loss value is used to reflect the difference degree between the reference word vector and the benchmark word vector;
[0062] When the loss value corresponding to the one target word is less than the loss value threshold, determine that the representation probability corresponding to the one target word is less than the probability threshold;
[0063] When determining that the representation probability corresponding to the one target word is less than the probability threshold, update the initial word vectors of the 2m surrounding words of the one target word in the model parameters of the neural network model, including:
[0064] Based on the loss value corresponding to the one target word, update the initial word vectors of the 2m surrounding words of the one target word and the output word matrix.
[0065] Optionally, the information processing device further includes:
[0066] A fifth acquisition module, configured to acquire multiple entity names in the target domain before performing word segmentation processing on the to-be-processed text information based on the word segmentation dictionary in the target domain;
[0067] A second construction module, configured to construct the word segmentation dictionary based on the multiple entity names.
[0068] On the other hand, an electronic device is provided, and the electronic device includes: a processor and a memory, and at least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the above information processing method.
[0069] On another hand, a computer-readable storage medium is provided, and at least one instruction, at least one program, a code set or an instruction set is stored in the storage medium, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the above information processing method.
[0070] The beneficial effects brought by the technical solution provided by this application at least include:
[0071] In the information processing method provided by this application, for the to-be-processed text information in the target field, the to-be-processed text information can be segmented based on the segmentation dictionary of the target field to obtain multiple words. In this way, the situation of dividing the information of the same word into two words when segmenting the text information in the target field can be avoided, and the accuracy of segmenting the text information in the target field is improved. And the target word vector in the word vector library based on which the target word vector is searched is determined by updating the initial model parameters of the neural network model based on the training text information in the target field. The target word vector determined in this way represents the words in the target field more accurately, and the accuracy of text processing for the to-be-processed text information based on this target word vector is relatively high, and the processing effect is good. And the initial model parameters are determined based on the initial word vectors with dimensions less than the number of target words in the word vector library. In this way, the determination process of the target word vector of the target word is also relatively simple and fast. Description of the Drawings
[0072] Figure 1 is a flowchart of an information processing method provided by an embodiment of this application;
[0073] Figure 2 is a flowchart of another information processing method provided by an embodiment of this application;
[0074] Figure 3 is a schematic structural diagram of a neural network provided by an embodiment of this application;
[0075] Figure 4 is a schematic diagram of the change of accuracy rate during the training process of an information processing model provided by an embodiment of this application;
[0076] Figure 5 is a schematic structural diagram of an information processing device provided by an embodiment of this application;
[0077] Figure 6 is a schematic structural diagram of a server provided by an embodiment of this application. Detailed Embodiments
[0078] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe the embodiments of this application in detail with reference to the accompanying drawings.
[0079] Currently, in order to ensure the efficiency and accuracy of text information analysis and processing, artificial intelligence (AI) technology is increasingly used to analyze and process text information. Artificial intelligence uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems for perceiving the environment, acquiring knowledge, and using knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies.
[0080] AI basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. AI software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing (NLP) technology, and machine learning (ML) / deep learning technology. The solution provided in the following embodiments of this application relates to natural language processing technology and machine learning technology in artificial intelligence.
[0081] Machine learning is an interdisciplinary subject involving multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. Machine learning specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning by teaching.
[0082] Natural language processing is an important direction in the field of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, and it has a close connection with the research of linguistics. NLP technology usually includes technologies such as text processing, semantic understanding, machine translation, robot question answering, and Knowledge Graph.
[0083] A knowledge graph is a series of different graphs that display the development process and structural relationships of knowledge. The knowledge graph uses visualization technology to describe knowledge resources and their carriers, and can mine, analyze, construct, draw, and display knowledge and the relationships between knowledge. Essentially, a knowledge graph is a semantic network that describes the knowledge objectively existing in the real world and the association relationships between knowledge. The knowledge graph has powerful semantic processing capabilities and the ability to interconnect and organize information, which can contribute to the intelligent application of information in the Internet. Knowledge graphs are usually divided into general knowledge graphs and vertical knowledge graphs. The general knowledge graph is not oriented to a specific domain and emphasizes more on the breadth of the covered knowledge. The general knowledge graph only covers the common sense knowledge in multiple domains. The vertical knowledge graph, on the other hand, is oriented to a specific domain and emphasizes more on the depth of the covered knowledge. The vertical knowledge graph covers the more detailed knowledge in that specific domain.
[0084] The embodiments of this application provide an information processing method, which can improve the processing accuracy and processing effect of text information. This method can be applied to an information processing device, such as an electronic device capable of information processing. This electronic device can be a server. For example, this electronic device can be a single server, or a server cluster composed of multiple servers, or a cloud computing service center. Optionally, this electronic device can also be a terminal with strong computing capabilities, such as a desktop computer, a personal computer (PC), a smart phone, a wearable device, or a tablet computer, etc.
[0085] Figure 1 is a flowchart of an information processing method provided by the embodiments of this application. This method can be used in an information processing device, such as Figure 1 As shown, this method can include:
[0086] Step 101, obtain the text information to be processed in the target domain.
[0087] Exemplarily, the text information to be processed in the target domain can include any text information such as articles, web page information, and teaching content in the target domain.
[0088] Step 102, based on the word segmentation dictionary in the target domain, perform word segmentation on the text information to be processed to obtain multiple words in the text information to be processed.
[0089] It should be noted that a word refers to the smallest unit that can be independently used in a language, and a word is the basic element of a natural language. For example, each word can be a string, and each word can be composed of one or more characters, such as composed of at least two characters. Optionally, the word segmentation dictionary in the target domain can include multiple entity names in the target domain.
[0090] Step 103: In the word vector library of the target domain, search for the target word vectors of the multiple words, where the target word vectors are used to represent the semantics of the corresponding words.
[0091] Among them, the word vector library of the target domain includes: the target word vectors of multiple target words belonging to the target domain. It should be noted that a word vector is a word represented by a digital vector, and the target word vector of a word can represent the semantics of the word. The target word vectors of the target words in the word vector library of the target domain can be used in the processing of text information in the target domain. The number of target words in this word vector library is greater than a quantity threshold. For example, the number of target words can be in the thousands or even reach hundreds of thousands or millions.
[0092] All the words in the text to be processed may belong to the word vector library of the target domain, so that the target word vectors of each word in the text information to be processed can be found; the text information to be processed may also include words that do not exist in this word vector library. In this case, after step 103, only the target word vectors of the target words that exist in the word vector library among the multiple words in the text information to be processed can be found. The target words in this word vector library appear more frequently in the target domain than a frequency threshold and play a relatively key role in the text information of the target domain. Based on the target words in the word vector library, it is sufficient to understand the text information of the target domain. Even if there are words in the text information to be processed for which no target word vectors are found, the absence of the word vectors of these words has a minimal impact on the processing of the text information to be processed.
[0093] In the embodiment of the present application, the target word vector of each target word in the word vector library of the target domain is determined based on the model parameters of a neural network model; the model parameters are obtained by training the neural network model based on the training text information of the target domain to update the initial model parameters of the neural network model. The training text information includes the multiple target words, and the initial model parameters are determined based on the initial word vectors configured for each of the multiple target words. The dimension of the initial word vectors can be less than the number of the multiple target words. It should be noted that the target word vectors of the target words in the word vector library can directly be the model parameters of the neural network model, or the model parameters of the neural network model can be adjusted to obtain the target word vectors of the target words in the word vector library. The initial model parameters of the neural network model can include the initial word vectors configured for each of the multiple target words, or can also include the parameters obtained by adjusting the initial word vectors of the multiple target words.
[0094] Step 104: Based on the target word vectors of the target words found among the multiple words, perform text processing on the text information to be processed.
[0095] In the embodiments of the present application, different text processing operations can be performed on the text information to be processed based on the user's requirements. For example, information classification processing can be performed on the text information to be processed to distinguish the information belonging to different sub-domains in the target domain in the text information to be processed; alternatively, information comparison processing can be performed on the text information to be processed to determine the similarity degree between the text information to be processed and other text information; or information extraction processing can be performed on the text information to be processed to obtain an abstract of the text information to be processed. Or other text processing operations can be performed on the text information to be processed, which are not listed in the embodiments of the present application.
[0096] Optionally, the user can set in advance the text processing method to be performed on the text information to be processed, and the information processing device can obtain this text processing method together when obtaining the text information to be processed. Alternatively, the information processing device can have its corresponding information processing method and directly perform information processing on the text information to be processed in this way after obtaining the text information to be processed.
[0097] The processing of text information is actually the processing of the word vectors of the words in the text information. Therefore, when processing the text information, it is necessary to first convert the words in the text information into word vectors and then perform subsequent processing. In the embodiments of the present application, when processing the text information to be processed, if it is determined that there is a word in the text information to be processed that is the same as any of the multiple target words in the word vector library, the target word vector of this word can be directly obtained from the word vector library for subsequent text processing, without the need to perform the conversion from the word to the word vector again. Optionally, the information processing device can process the text information to be processed through an information processing model. In this case, the word vector library of the target domain can be pre-loaded into the information processing model. Furthermore, after inputting the text information to be processed into the information processing model, the information processing model can directly process the text information to be processed based on the target word vectors of the multiple target words in the word vector library to obtain the processing result of the text information to be processed.
[0098] It should be noted that in the related art, word segmentation is performed on the text information to be processed based on a general word segmentation dictionary. Based on general web text information, such as a large amount of various types of text information obtained on the Internet (such as news, information, and encyclopedia text information), a vector conversion model (i.e., a neural network model) is trained to obtain the target word vectors of each word among multiple words in the text information. Subsequently, the text information to be processed is processed based on the target word vectors of the multiple words. The general web text information includes text information in multiple fields, such as text information including encyclopedic knowledge in various fields, text information of news and information, and text information such as user comments. The word vectors obtained in this way are relatively difficult to match the words in the text information of some specific fields, thereby resulting in a poor processing effect on the text information of the specific field. For example, the text information in the medical and health field is filled with a large number of words such as disease names, symptom names, and drug names, and there are usually specific language expression habits in the medical and health field. These words and language expression habits are different from the words and language expression habits used in other fields and people's daily communication. Thus, if the target word vectors of words obtained by training based on general web text information in the related art are used to process the text information in the medical and health field, the accuracy of the processing result is poor.
[0099] In the embodiment of the present application, word segmentation can be performed on the text information to be processed in the target field based on the word segmentation dictionary of the target field to obtain multiple words. In this way, the multiple words are adapted to the word segmentation method and language expression habits in the target field. And the neural network model can be trained based on the training text information of the target field to update the model parameters of the neural network model, obtain the target word vectors of multiple target words in the training text information, and then construct a word vector library for the target field. For example, the target field can be the medical and health field. In this way, the target word vectors of the multiple target words in the word vector library can accurately represent the words in the text information of the target field, and thus the processing effect of the text information of the target field can be better based on the target word vectors of the multiple target words.
[0100] In summary, in the information processing method provided by the embodiments of the present application, for the to-be-processed text information in the target field, the to-be-processed text information can be segmented based on the segmentation dictionary of the target field to obtain multiple words. In this way, the situation where the information of the same word is divided into two words when segmenting the text information in the target field can be avoided, and the accuracy of segmenting the text information in the target field is improved. And the target word vector in the word vector library based on which the target word vector is searched is determined by updating the initial model parameters of the neural network model based on the training text information of the target field. The target word vector determined in this way represents the words in the target field more accurately, and the accuracy of processing the to-be-processed text information based on the target word vector is relatively high, and the processing effect is good. And the initial model parameters are determined based on the initial word vectors with dimensions less than the number of target words in the word vector library. In this way, the determination process of the target word vector of the target word is also relatively simple and fast.
[0101] In the embodiments of the present application, before processing the to-be-processed text information or before obtaining the to-be-processed text information, a word vector library of the target field can be constructed first. By way of example, Figure 2 FIG. is a flowchart of another information processing method provided by the embodiments of the present application. This method can be used in an information processing device and is used to construct a word vector library of the target field. As Figure 2 shown, this method may include:
[0102] Step 201, obtain multiple entity names of the target field.
[0103] In the embodiments of the present application, the target field is taken as the medical and health field as an example. The target field may also be other fields outside the medical and health field, and the target field may satisfy: there are specific language expression habits in the target field, and the words in the target field are used less frequently in human daily life.
[0104] By way of example, the information processing device may first obtain the knowledge graph of the target field, and then determine multiple entity names of the target field based on the knowledge graph of the target field. The knowledge graph may be a vertical knowledge graph of the target field. For example, the information processing device may obtain the data of the entities in the target field, perform knowledge extraction on the entity data, and then construct the knowledge graph of the target field based on the extracted knowledge. Optionally, the knowledge graph may also be a publicly constructed knowledge graph, and the information processing device may directly obtain it from the Internet.
[0105] The knowledge graph of the target domain can reflect the interconnection relationships between various entities in the target domain, and this knowledge graph includes multiple entity names in the target domain. After obtaining the knowledge graph of the target domain, the information processing device can directly obtain each entity name in this knowledge graph. By way of example, the entities in the medical and health domain can include entities of categories such as body parts, treatment objects, diseases, examinations, departments, instruments, foods, surgeries, acupoints, drugs, doctors, hospitals, symptoms, and treatment methods. There are multiple entities in each category, and each entity has a corresponding entity name. For example, the number of entity names obtained by the information processing device can be in the thousands or tens of thousands. For example, the information processing device can obtain 373,809 entity names in the medical and health domain. It should be noted that the number of entity names obtained by the information processing device is related to the specific domain and the obtained knowledge graph, and this application does not limit the number of entity names.
[0106] Step 202: Construct a word segmentation dictionary for the target domain based on the multiple entity names. Then perform Step 203.
[0107] After obtaining multiple entity names of the target domain, the information processing device can construct a word segmentation dictionary based on these multiple entity names. This word segmentation dictionary can include these multiple entity names, and this word segmentation dictionary can be used to construct a word segmenter. A word segmenter is a tool that analyzes an input text into a logical sequence of words. The word segmenter can divide the information in the text that is the same as any entity name in the word segmentation dictionary into one word.
[0108] Step 203: Obtain the training text information of the target domain.
[0109] By way of example, the information processing device can obtain text information of multiple domains from the Internet, and then screen out the text information of the target domain from this text information of multiple domains, or the information processing device can also directly obtain the text information of the target domain from the Internet. Then, the information processing device can use the obtained text information of the target domain as the training text information. It should be noted that this application embodiment takes the execution of Step 203 after Steps 201 and 202 as an example; optionally, Step 201 and Step 203 can also be executed in parallel, or Step 201 and Step 202 can also be executed after Step 203 is executed. This application embodiment does not make a limitation.
[0110] In the embodiments of the present application, the data volume of the training text information of the target field obtained by the information management device may be greater than the data volume threshold. For example, the medical text information obtained by the information management device may include millions of relevant articles on medical health. For example, 4,705,244 relevant articles on medical health are obtained, and the data volume of this text information may be approximately 21G (gigabytes). In the embodiments of the present application, the data volume of the training text information of the target field obtained is large, and this data volume is greater than the data volume of the general web text information based on which the word vectors of general words are determined in the related art. Therefore, a more stable vector distribution can be obtained based on the training text information of the target field, and the target word vectors of the determined words can accurately represent these words.
[0111] Step 204: Based on the word segmentation dictionary of the target field, perform word segmentation processing on the training text information of the target field to obtain a plurality of target words.
[0112] The information processing device may perform word segmentation processing on the obtained text information of the target field based on this word segmentation dictionary, that is, refer to this word segmentation dictionary when performing word segmentation processing on this text information. For example, the information processing device may divide the information in the text information that is the same as any entity name in the word segmentation dictionary into one word. If the text information does not contain the information of any of the plurality of entity names, the information processing device uses a general word segmentation method to perform word segmentation processing on this information. After performing word segmentation processing on the text information of the target field, the obtained plurality of target words may include at least some of the plurality of entity names and other words outside the plurality of entity names. Optionally, the information processing device may input the training text information into a word segmenter constructed based on the word segmentation dictionary of the target field, so as to perform word segmentation processing on the training text information through this word segmenter, and then obtain the plurality of target words output by the word segmenter.
[0113] In the embodiments of the present application, word segmentation is performed on the training text information of the target field based on a word segmentation dictionary including a plurality of entity names in the target field, so as to ensure that the words used in the target field are completely divided; avoid dividing the information belonging to one word into two different words, avoid the loss of some information in the word during word segmentation processing, ensure that the divided words are adapted to the language expression habits of the target field, and improve the accuracy of word division.
[0114] Optionally, the information processing device may also construct a vocabulary based on the multiple target words. Each target word is located at a position in the vocabulary, and the number of the multiple target words is the size of the vocabulary. For example, the number of the multiple target words is V. Each target word in the vocabulary may have a corresponding benchmark word vector, which is used to accurately represent the position of the word among the multiple target words obtained after segmenting the text information in the target domain. For example, the benchmark word vector can represent the position of the target word in the vocabulary, and the position where the value is 1 in the benchmark word vector is the position of the target word in the vocabulary. For example, a certain vocabulary consists of five target words, then the benchmark word vector of the third target word among them is (0, 0, 1, 0, 0).
[0115] It should be noted that the number of the multiple target words can be 10,000, 20,000, or even millions. Although the benchmark word vector can accurately represent the target word, directly using the benchmark word vector of the target word for text information processing will cause huge consumption of computing resources and storage disasters. Therefore, low-dimensional word vectors are needed to represent the target words. The information processing method provided in the embodiments of the present application for determining the target word vector is to determine a low-dimensional vector that can more accurately represent the target word in the target domain.
[0116] Step 205: Randomly configure an initial word vector for each of the multiple target words.
[0117] The information processing device may first randomly configure an initial word vector X for each of the multiple target words obtained, and then optimize and update the initial word vector X to obtain the target word vector for each target word, which is used to more accurately represent its corresponding target word. The initial word vector can be an n-dimensional vector, and the dimension of the initial word vector can be set by the staff. The dimension n of the initial word vector is much smaller than the number V of the multiple words. For example, the dimension of the initial word vector can be 100, 200, or other values. For example, the initial word vector can be an n-dimensional column vector, and the initial word vector can be represented by an n*1 matrix.
[0118] For example, the information processing device may obtain the dimension of the initial word vector input by the staff, and then, based on this dimension, randomly initialize the word vectors for each of the determined multiple target words to implement configuring an initial word vector for each target word. The information processing device may input the dimension into the function package for randomly initializing vectors and receive the initial word vectors configured for each target word output by the function package.
[0119] Step 206: Determine initial model parameters based on the initial word vectors of the multiple target words, and construct a neural network model based on the initial model parameters.
[0120] Exemplarily, the information processing device may directly determine the initial word vectors configured for the multiple target words as the initial model parameters. Optionally, the information processing device may also construct an output word matrix and use the output word matrix as the initial model parameter as well. The output word matrix is used to reflect the mapping relationship from the words around the target word to the target word. The mapping relationship from the words around each target word in the vocabulary to each target word can be reflected by the output word matrix. The output word matrix may be a matrix of V*n, that is, a matrix with V rows and n columns. The information processing device may randomly construct a matrix of V*n as the initial output word matrix, and then construct a neural network model based on the initial output word matrix.
[0121] In the embodiments of the present application, the neural network model may be a continuous bag of words (CBOW) model, and this model is actually a fully connected neural network. Figure 3 It is a schematic structural diagram of a neural network model provided by the embodiments of the present application. As Figure 3 shown, the neural network model includes an input layer, a projection layer, and an output layer. Exemplarily, the information processing device may input the initial word vectors of the target words into the input layer of the neural network model. Optionally, the initial word vectors for the multiple target words in the training text information may also be randomly configured directly through the input layer. The input layer may transmit the initial word vectors to the projection layer and may also update the initial word vectors; the projection layer may sum and average the multiple initial word vectors output by the input layer to obtain the output result of the projection layer; the output layer may multiply the output result of the projection layer by the output word matrix to obtain the output result of the output layer.
[0122] A neural network model is formed by a large number of nodes (also called "neurons" or "units") connected to each other, and each node represents a specific output function. The connection between every two nodes represents a parameter value, which is called a weight (English: weight). Different weights and activation functions will result in different outputs of the neural network. In Figure 3 the shown neural network model, the input layer has 2m*n nodes, the projection layer has n nodes, the output layer has V nodes, and the output word matrix U is the connection edge from the projection layer to the output layer, and this connection edge represents the connection from the nodes in the projection layer to the nodes in the output layer. Optionally, the connection edge from the input layer to the projection layer may be represented by S. In the embodiments of the present application, S is a constant 1, and the input layer can be directly corresponded to the projection layer.
[0123] Step 207: Obtain a training sample set based on the training text information. The training sample set includes multiple training samples and the training labels of each training sample; each training sample includes the initial word vectors of 2m surrounding words of a target word, and the training label of the training sample includes the benchmark word vector set for the target word.
[0124] Each training sample in the training sample set is used to reflect the semantics of the target word in a sample text. The sample text is a sentence in the training text information that includes at least two target words; the training label of the training sample is used to reflect the position of the target word in the multiple target words obtained by segmenting the training text information. Optionally, the sample text includes at least three target words. Each training sample can be used to reflect the semantics of some of the target words in a sample text, and the training label of the training sample is used to reflect the position of a target word in the vocabulary. The target word is different from the part of the target words, and in the sample text, the part of the target words is located around (i.e., before and after) the target word. In this application, words are characterized by the words around them. Therefore, the semantic representation of the words around the word is used as the training sample, and the actual position of the word in the vocabulary is used as the training label.
[0125] Optionally, multiple training samples can be obtained based on a sample text, and the training labels of different training samples can reflect the positions of different target words in the vocabulary in the sample text. By way of example, a certain sample text may include target words A, B, C, D, and E in sequence. The training sample 1 obtained based on this sample text can reflect the semantics of target words A and C, and the training label of this training sample can reflect the position of target word B in the vocabulary; the training sample 2 obtained based on this sample text can reflect the semantics of target words B and D, and the training label of this training sample can reflect the position of target word C in the vocabulary; the training sample 3 obtained based on this sample text can reflect the semantics of target words C and E, and the training label of this training sample can reflect the position of target word D in the vocabulary. It should be noted that this example only explains the situation of obtaining multiple training samples based on the same sample text, and does not limit the number of target words reflected by the training samples.
[0126] Optionally, for each target word in the vocabulary, multiple training samples corresponding to the target word can be determined, and each training sample in the multiple training samples can correspond to different sample texts. For a target word, corresponding training samples can be obtained based on each sample text including the target word in the training text information. Optionally, corresponding training samples can also be obtained only based on some of the sample texts including the target word in the training text information. For example, the number of training samples corresponding to each target unit can be greater than a quantity threshold, such as the number of training samples can be 100, 128 or other values. In this way, for each target word, the target word vector can be determined through multiple training samples, which can simplify the process of determining the target word vector while ensuring the accuracy of the representation of the word by the target word vector.
[0127] In the embodiments of the present application, each training sample may include the initial word vectors of 2m surrounding words of a target word, and the training label of the training sample may include a benchmark word vector set for the target word, and the benchmark word vector is used to reflect the position of the target word in the vocabulary. Wherein, m≥1, and the target word and the 2m surrounding words may belong to the same sample text. The surrounding words of a word are also the words located around (i.e., before and / or after) the word. The 2m surrounding words of each target word include: in the sample text corresponding to the training sample, the m target words closest to the target word before the target word (i.e., the upper text of the word), and the m target words closest to the target word after the target word (i.e., the lower text of the word). Relative to the 2m surrounding words of a certain target word, the target word can be called the center word or the target word. The sample text to which the center word and the 2m surrounding words of the center word belong can be a sentence or a paragraph (i.e., multiple sentences). For example, if a text message is "Xiaoming's home is in Shanghai", after word segmentation of the text message, three words "Xiaoming", "home" and "Shanghai" can be obtained. If the word "home" is used as the center word, the surrounding words of the center word include "Xiaoming" and "Shanghai". It should be noted that different words in a sentence in the text information can be surrounding words of each other. If the word "Shanghai" is used as the center word, the surrounding words of the center word can include "home"; if the word "Xiaoming" is used as the center word, the surrounding words of the center word can also include "home".
[0128] Step 208, input each training sample in the training sample set into the neural network model.
[0129] For example, the information processing device can input each training sample in the training sample set into the input layer of the constructed neural network model, and then train the neural network model based on each training sample and the corresponding training label.
[0130] Step 209: Determine the hidden vector of the target word by adding and then averaging the initial word vectors of the 2m surrounding words of the target word through the neural network model.
[0131] In this application, the 2m surrounding words of the target word are used to represent the target word, so for each target word, the initial word vectors of the 2m surrounding words of the target word are calculated. Optionally, the processing of the training samples for each target word by the neural network model can be executed in parallel.
[0132] Exemplarily, the neural network model can determine the hidden vector h of the target word based on the initial word vectors of the 2m surrounding words of the target word through the first formula. The first formula is where X i represents the initial word vector of the i-th word among the 2m surrounding words, and t represents the central position of the 2m surrounding words, that is, the position of the target word. In the embodiments of this application, the 2m surrounding words around this position are determined with the target word as the central position. The hidden vector of the target word is also the average value obtained after adding the initial word vectors of the 2m surrounding words of the target word. The initial word vector is an n-dimensional vector, and the hidden vector obtained by adding 2m n-dimensional vectors and then averaging is also an n-dimensional vector. If the hidden vector is an n-dimensional column vector, the hidden vector can be represented by an n*1 matrix. For example, the projection layer of the neural network model can use the first formula to process the initial word vectors of the 2m surrounding words of the target word output by the input layer to obtain the hidden vector of the target word.
[0133] Step 210: Determine the reference word vector of the target word by multiplying the hidden vector of the target word by the output word matrix through the neural network model.
[0134] Among them, the output word matrix can reflect the mapping relationship from the 2m surrounding words of the target word to the target word. The reference word vector of a word is used to reflect the probability that the word represented by the initial word vectors of the 2m surrounding words of the word is each target word in the vocabulary.
[0135] Exemplarily, the output layer of the neural network model can use the second formula to process the hidden vector of the target word output by the projection layer to obtain the reference word vector of the target word. In the embodiments of this application, is used to represent the reference word vector of the word. The reference word vector is a V-dimensional vector. The reference word vector can be a V-dimensional column vector, and the reference word vector can be represented by a V*1 matrix. The second formula is: Among them, U represents the output word matrix, which is used to reflect the mapping relationship from 2m surrounding words of the target word to the target word. The output word matrix can be a matrix of V * n, that is, a matrix with V rows and n columns. The hidden vector h is an n-dimensional column vector. Thus, multiplying the output word matrix of V * n by the n-dimensional hidden vector can obtain a V-dimensional reference word vector. The mapping relationship from 2m surrounding words of each target word in the vocabulary to each target word can be reflected by this output word matrix.
[0136] Exemplarily, assume n = 5, and the auxiliary word is the third word among these five words. The information processing device obtains the reference word vector of the auxiliary word as (0.8, 0.6, 0.7, 0.2, 0.4) through the second formula. This reference word vector indicates that the probability that the word represented by the 2m surrounding words of the auxiliary word is the first word among these five words is 0.8, the probability that it is the second word among these five words is 0.6, the probability that it is the third word among these five words is 0.2, the probability that it is the fourth word among these five words is 0.6, and the probability that it is the fifth word among these five words is 0.4.
[0137] Step 211: Based on the reference word vectors of each target word among the multiple target words and the training labels in the training set, determine whether there is a target word whose corresponding representation probability is less than the probability threshold among the multiple target words. When there is a target word whose corresponding representation probability is less than the probability threshold among the multiple target words, execute Step 212; when there is no target word whose corresponding representation probability is less than the probability threshold among the multiple target words, that is, when the corresponding representation probability of each target word among the multiple target words is greater than or equal to the probability threshold, execute Step 213.
[0138] Step 211 can also be to determine whether there is a representation probability corresponding to each target word in the multiple target words that is greater than or equal to a probability threshold. The representation probability corresponding to a target word is: the probability that the words represented by the initial word vectors of the 2m surrounding words of the target word are the target word. If the probability that the words represented by the initial word vectors of the 2m surrounding words of the target word are the target word is greater than or equal to the probability threshold, it indicates that the words represented by the initial word vectors of the 2m surrounding words can accurately represent the target word. However, if only the initial word vectors of the 2m surrounding words of a certain target word can accurately represent this word, it is not enough to determine the initial word vector of the target word, nor is it enough to determine that the accuracy of the position representation of the 2m surrounding words by the initial word vectors of the 2m surrounding words meets the requirements. Each target word obtained by performing word segmentation on the training text information in the target domain may be a surrounding word of each other. When the initial word vectors of the 2m surrounding words of each target word can accurately represent the corresponding target word, it can be determined that the initial word vector of each target word can accurately represent the position of the target word. The above-mentioned position of the target word refers to the position of the target word in the word list composed of multiple target words obtained by word segmentation of the training text information. Optionally, the probability threshold can be a preset fixed threshold, such as 0.98.
[0139] In an alternative implementation, whether the representation probability corresponding to each of the foregoing target words is less than the probability threshold can be reflected by a corresponding loss value, and this loss value is used to reflect the degree of difference between the reference word vector and the benchmark word vector of the target word. The information processing device can determine whether the representation probability corresponding to the target word is less than the probability threshold by calculating the loss value corresponding to each target word. For example, the information processing device can obtain a preset loss function (also called a cost function), and after determining the reference word vector for each target word, calculate the loss value L corresponding to this loss function. When the loss value corresponding to a certain target word is less than the loss value threshold, such as when the loss value converges within the target range, it is determined that the representation probability corresponding to the target word is greater than or equal to the probability threshold, and the initial word vectors of the 2m surrounding words can accurately represent the target word. When the loss value corresponding to a certain target word is greater than or equal to the loss value threshold, such as when the loss value is outside the target range, it is determined that the representation probability corresponding to the target word is less than the probability threshold, and the initial word vectors of the 2m surrounding words still cannot accurately represent the target word.
[0140] The loss value L may be a preset operation value between the reference word vector of the target word and the reference word vector of the target word. The preset operation value is the mean squared error (English: Mean squared error; abbreviation: MSE), or the mean absolute difference (that is, first obtain the absolute value of the difference of the corresponding pixel values, and then obtain the average value of all the absolute values of the differences), or the sum of absolute differences (that is, first obtain the absolute value of the difference of the corresponding pixel values, and then obtain the sum of all the absolute values of the differences), or the standard deviation, or the cross entropy (English: Cross Entropy; abbreviation: CE). Exemplarily, for any target word, the loss value may be the L2 norm between the reference word vector of the target word and the reference word vector. The information processing device may determine the loss value L based on the reference word vector of the target word through the third formula. The third formula is: Where n represents the dimension of the initial word vector, and y represents the reference word vector set for the target word.
[0141] It should be noted that in the embodiments of the present application, the information processing device can train the neural network model with at least one training sample for each target word. Each training sample of the target word includes the initial word vectors of 2m surrounding words of the target word. At least one word in the surrounding words of different training samples of the same target word can be different. If only one training sample is used for calculation for each target word, the loss value corresponding to the target word is: the loss value between the reference word vector of the target word determined based on the training sample and the reference word vector of the target word. If the information processing device uses multiple training samples of the target word for calculation for each target word, the information processing device executes the above steps 209 and 210 for each training sample; and then determines a reference word vector of the target word based on each training sample, and determines a representation probability corresponding to the target word based on the reference word vector. When the representation probabilities corresponding to the target word obtained by the calculation of each training sample are all greater than or equal to the probability threshold, step 213 is executed; when at least one of the representation probabilities corresponding to the target word obtained by the calculation of one training sample is less than the probability threshold, step 212 is executed. In this case, the loss value corresponding to the target word may be the sum of the loss values determined based on each training sample. When the sum of the loss values is within the target range, it is determined that the representation probabilities corresponding to the target word obtained by the calculation of each training sample are all greater than or equal to the probability threshold; when the sum of the loss values is outside the target range, it is determined that the representation probability corresponding to at least one training sample obtained by the calculation is less than the probability threshold.
[0142] In step 212, when updating the model parameters of the neural network model, the initial word vectors of the 2m surrounding words of the target word are updated.
[0143] When the information processing device determines that the representation probability corresponding to a certain target word is less than the probability threshold, it can determine that the initial word vectors of the 2m surrounding words of the target word still cannot meet the requirements for representing the position of the target word, and thus can update the initial word vectors of the 2m surrounding words of the target word. Exemplarily, when the information processing device determines that the loss value corresponding to the target word is outside the target range, it can update the initial word vectors of the 2m surrounding words of the target word based on the loss value. Optionally, the information processing device can also update the output word matrix based on the loss value. Optionally, after updating the initial word vectors of the 2m surrounding words of the target word in the model parameters of the neural network model, the information processing device can execute step 209 to determine again whether the updated model parameters meet the requirements.
[0144] After updating the initial word vectors of the 2m surrounding words of the target word and the output word matrix, the information processing device can verify whether the updated initial word vectors of the 2m surrounding words meet the requirements, that is, determine whether the representation probability of the initial word vectors of the 2m surrounding words for the target word is greater than or equal to the probability threshold. The information processing device can repeatedly execute the above steps 209 to 212 until the representation probability of the initial word vectors of the 2m surrounding words of the target word is greater than or equal to the probability threshold. Exemplarily, the information processing device can repeatedly execute the above steps 209 to 212 for the target word until the loss value corresponding to the target word converges to the target range. During this loop process, the function value of the loss function can continuously decrease, and finally fluctuate within a very small numerical range, and this numerical range is a relatively optimal numerical range, and this numerical range is the target range.
[0145] It should be noted that the above steps 209 to 212 are the process of training the neural network model based on each training sample and the corresponding training label in the training sample set to update the initial model parameters of the neural network model.
[0146] Step 213: Obtain the target word vectors of the multiple target words based on the model parameters of the trained neural network model to construct a word vector library for the target domain.
[0147] When it is determined in the above step 211 that there is no representation probability corresponding to the target word among the multiple target words that is less than the probability threshold, the information processing device may determine that the 2m initial word vectors of each target word can accurately represent the word, and the training of the neural network model has met the requirements, and thus the training of the neural network model can be stopped. For example, the information processing device may determine the initial word vectors of each target word among the multiple target words as the target word vectors of each target word in the model parameters of the trained neural network model. Then, the information processing device may construct a word vector library for the target domain based on the target word vectors of each target word, so as to perform subsequent text processing on the text information based on the word vector library. Optionally, the word vector library may store the correspondence between the target word and its target word vector. In the embodiments of the present application, the target word vector of a word may also be referred to as the pre-trained vector of the word.
[0148] In the embodiments of the present application, by training (also referred to as pre-training) the neural network model until the loss value corresponding to each target word converges to the target range, a trained neural network model can be obtained; furthermore, based on the trained neural network model, the target word vectors of each target word can be obtained. The model parameters of the neural network model include the initial word vectors of each target word and the output word matrix. Before training the neural network model, the model parameters of the neural network model are initialized, such as randomly setting the initial word vectors of each target word and the output word matrix. After the neural network model is trained, the target word vectors of each target word can be determined based on the model parameters of the trained neural network model. Optionally, after the neural network model is trained, the central word can also be determined based on multiple surrounding words through the neural network model.
[0149] For example, the training process of the neural network can be implemented based on the training method of the supervised learning algorithm (English: supervised learning). The supervised learning algorithm is trained through an existing training sample set (including training samples and their corresponding training labels, the training samples are known data, and the training labels can be clear identifiers or output data) to train the final parameters of the neural network. In the embodiments of the present application, the training samples in the training set include multiple intermediate words and 2m surrounding words of the intermediate words, and the training label corresponding to the training sample is the benchmark vector of each intermediate word. Multiple training samples can be used to train each intermediate word. The 2m surrounding words of the intermediate word in the multiple training samples can be different, and each training sample includes a set of initial word vectors of the surrounding words. Optionally, the training process can also be implemented by manual calibration, or an unsupervised learning algorithm, or a semi-supervised learning algorithm, etc.
[0150] It should be noted that the parameters of the neural network model in the related art include the trained word vectors of each word, the input word matrix, and the output word matrix. The input word matrix is the connection edge from the input layer to the projection layer, and the input word matrix is a matrix of V * n. When initializing the input layer of the neural network model in the related art, each word is represented by one-hot encoding to obtain the V-dimensional initial word vector corresponding to each word. Then, the initial word vectors of the 2m surrounding words of the central word are multiplied by the input word matrix to reduce the dimension of the word vector of each surrounding word, and the n-dimensional trained word vector of each surrounding word is obtained. After that, the trained word vectors of the 2m surrounding words are summed and then averaged through the projection layer, and then the output layer uses a logistic regression function (such as the softmax function) to process the output result of the projection layer to obtain the representation probability of the initial word vectors of the 2m surrounding words for each word, in order to maximize the representation probability of the central word. Then, based on the processing result of the output layer, the parameters of the neural network model are trained, and the trained word vectors of each word after the training is completed are determined as the target word vectors of the words.
[0151] In the embodiments of the present application, low-dimensional (n-dimensional) initial word vectors are directly randomly configured for each word, and then calculations can be directly performed through the projection layer and the output layer based on the initial word vectors, and the model parameters of the neural network model are trained. In the embodiments of the present application, the model parameters of the neural network model do not include the input word matrix, and there is no need to perform the step of multiplying the initial word vector by the input word matrix, and there is no need to train the input word matrix, which simplifies the information processing process and improves the information processing efficiency. The output layer in the embodiments of the present application can also use a logistic regression function (such as the softmax function) to process the output result of the projection layer.
[0152] Next, the effects of the target word vectors obtained in the embodiments of the present application in processing the text information in the target domain are introduced. For example, the text information in the target domain includes two datasets, cMedIC and cMedTC, where the task of Chinese medical intent classification needs to be performed on the dataset cMedIC, and the benchmark task of text classification needs to be performed on the dataset cMedTC.
[0153] For example, the information processing device may process the two data sets using different information processing models based on the first word vector library and the second word vector library respectively. The first word vector library includes the target word vectors of each target word obtained by training with the text information in the medical and health field in this application, that is, the word vector library of the target field created in the embodiments of this application. The second word vector library includes the target word vectors of words obtained by training with general web text information. Here, m represents the first word vector library and c represents the second word vector library. Taking the information processing device as an example of processing the two data sets using five information processing models respectively, the five information processing models include the FastText model, the TextCNN model, the TextRNN model, the TextRNN_Att model, and the Transformer model. The processing effect of the information processing device on the data set can be characterized by multiple metrics. The multiple metrics include precision, recall, F1-measure, and accuracy. The multiple metrics may also include the coverage rate of the word vector set for the words in the data set. The higher the value of each metric indicates the better the processing effect of the information processing device on the data set.
[0154] Table 1 below shows the values of each metric obtained after the information processing device processes the data set cMedIC using each information processing model. Table 2 below shows the values of each metric obtained after the information processing device processes the data set cMedTC using each information processing model. The upward arrows in Table 1 and 2 below indicate that when processing the data set based on the first word vector library compared to processing the data set based on the second word vector library, the corresponding metric values have increased.
[0155] Table 1
[0156]
[0157]
[0158] Table 2
[0159]
[0160] As can be seen from Table 1 above, on the dataset cMedIC, the coverage rate of the first word vector library for the words in this dataset has been greatly improved, and when the FastText model processes information using the first word vector library in the medical field, the values of all indicators are the largest. The amount of data in the dataset cMedIC is small, and the difference in the processing effects of some information processing models on this dataset based on different word vector libraries is small. While the amount of data in the dataset cMedTC is large. As can be seen from Table 2, the coverage rate of the first word vector library for the words in this dataset has increased by 25%, and the information processing effects obtained by most information processing models have been further improved. Moreover, the information processing model TextRNN_Att can obtain the optimal information processing effect based on the first word vector library.
[0161] Figure 4 It is a schematic diagram of the change in accuracy during the training process of an information processing model provided by an embodiment of the present application. Figure 4 It shows the change in accuracy obtained when the information processing model TextRNN_Att processes the dataset based on the first word vector library and the second word vector library respectively. Among them, the curve q1 represents the accuracy when the information processing model processes the dataset based on the first word vector library, and the curve q2 represents the accuracy when the information processing model processes the dataset based on the second word vector library. As Figure 4 shown, when the information processing model processes the dataset based on the first word vector library, the accuracy of the obtained result can converge faster, and the final effect of the model will also be better. Therefore, using the text information in the medical and health field to train the target word vector can help the information processing model converge faster and, in many cases, help the model obtain better processing effects.
[0162] In summary, in the information processing method provided by the embodiment of the present application, for the text information to be processed in the target field, the text information to be processed can be segmented based on the segmentation dictionary in the target field to obtain multiple words. In this way, the situation of dividing the information of the same word into two words when segmenting the text information in the target field can be avoided, and the accuracy of segmenting the text information in the target field is improved. And the target word vector in the word vector library based on which the target word vector is searched is determined by updating the initial model parameters of the neural network model based on the training text information in the target field. The target word vector determined in this way represents the words in the target field more accurately, and the accuracy of processing the text information to be processed based on this target word vector is relatively high, and the processing effect is good. And the initial model parameters are determined based on the initial word vectors with a dimension smaller than the number of target words in the word vector library. In this way, the determination process of the target word vector of the target word is also relatively simple and fast.
[0163] Figure 5It is a schematic structural diagram of an information processing device provided by an embodiment of the present application. As Figure 5 shown, the information processing device 50 may include:
[0164] An acquisition module 501, configured to acquire the text information to be processed in the target domain.
[0165] A first processing module 502, configured to perform word segmentation processing on the text information to be processed based on the word segmentation dictionary in the target domain, so as to obtain multiple words in the text information to be processed.
[0166] A search module 503, configured to search for the target word vectors of the multiple words in the word vector library in the target domain, where the target word vectors are used to represent the semantics of the corresponding words.
[0167] Among them, the word vector library in the target domain includes: the target word vectors of multiple target words belonging to the target domain; the target word vectors of each target word in the word vector library are determined based on the model parameters of the neural network model, and the model parameters are obtained by training the neural network model based on the training text information in the target domain to update the initial model parameters of the neural network model; the training text information includes the multiple target words, and the initial model parameters are determined based on the initial word vectors configured for each target word in the multiple target words, and the dimension of the initial word vectors is less than the number of the multiple target words.
[0168] A second processing module 504, configured to perform text processing on the text information to be processed based on the target word vectors of the target words found in the multiple words.
[0169] In summary, in the information processing device provided by the embodiment of the present application, for the text information to be processed in the target domain, word segmentation processing can be performed on the text information to be processed based on the word segmentation dictionary in the target domain to obtain multiple words. In this way, the situation that the information of the same word is divided into two words when performing word segmentation on the text information in the target domain can be avoided, and the accuracy of word segmentation of the text information in the target domain is improved. And the target word vectors in the word vector library based on which the target word vectors are searched are determined by updating the initial model parameters of the neural network model based on the training text information in the target domain. The target word vectors determined in this way represent the words in the target domain more accurately, and the accuracy of text processing of the text information to be processed based on the target word vectors is relatively high, and the processing effect is good. And the initial model parameters are determined based on the initial word vectors with a dimension less than the number of target words in the word vector library. In this way, the determination process of the target word vectors of the target words is also relatively simple and fast.
[0170] Optionally, the information processing device 50 further includes:
[0171] A second acquisition module, configured to acquire the training text information in the target domain.
[0172] A third processing module, configured to perform word segmentation on the training text information based on a word segmentation dictionary in a target domain to obtain the multiple target words.
[0173] A configuration module, configured to configure an initial word vector for each target word among the multiple target words.
[0174] A fourth acquisition module, configured to obtain a training sample set based on the training text information, where the training sample set includes multiple training samples and training labels for each training sample; each training sample is used to reflect the semantics of the target words in a sample text, and the sample text is a sentence in the training text information that includes at least two target words; the training label of the training sample is used to reflect the positions of the target words in the sample text among the multiple target words.
[0175] A training module, configured to train a neural network model based on each training sample and the corresponding training label in the training sample set to update the initial word vectors of the target words in the sample text corresponding to the training samples among the model parameters of the neural network model, so as to obtain a trained neural network model.
[0176] A first construction module, configured to obtain target word vectors for the multiple target words based on the model parameters of the trained neural network model, so as to construct a word vector library for the target domain.
[0177] Optionally, the configuration module is further configured to: randomly initialize word vectors for each target word among the multiple target words based on the dimension of the input initial word vector, so as to configure an initial word vector for each target word.
[0178] Optionally, the training module is further configured to:
[0179] input each training sample in the training sample set into the neural network model, where each training sample includes initial word vectors of 2m surrounding words of one target word among the multiple target words; where m≥1, and the 2m surrounding words include: among the sample text corresponding to the training sample, the m target words that are before the target word and closest to the one target word, and the m target words that are after each target word and closest to the one target word;
[0180] output a reference word vector for the one target word through the neural network model based on the initial word vectors of the 2m surrounding words of the one target word; the reference word vector of the target word is used to reflect: the probabilities that the words represented by the initial word vectors of the 2m surrounding words of the target word are each target word among the multiple target words;
[0181] When, based on the reference word vectors of each target word among the multiple target words and the training labels in the training set, it is determined that the representation probability corresponding to each target word among the multiple target words is greater than or equal to the probability threshold, stop training the neural network model; the representation probability corresponding to a target word is: the probability that the word represented by the initial word vectors of the 2m surrounding words of the target word is the target word;
[0182] When, based on the reference word vector of the one target word and the training label corresponding to the one target word, it is determined that the representation probability corresponding to the one target word is less than the probability threshold, update the initial word vectors of the 2m surrounding words of the one target word in the model parameters of the neural network model; return to execute the step of outputting the reference word vector of the one target word through the neural network model based on the initial word vectors of the 2m surrounding words of the one target word.
[0183] Optionally, the training module is further configured to:
[0184] Through the neural network model, determine the hidden vector of the one target word as the vector obtained by adding and then averaging the initial word vectors of the 2m surrounding words of the one target word;
[0185] Through the neural network model, determine the reference word vector of the one target word as the vector obtained by multiplying the hidden vector of the one target word by the output word matrix, and the output word matrix is used to reflect the mapping relationship from the 2m surrounding words to the one target word.
[0186] Optionally, the training label of the training sample includes a benchmark word vector set for the one target word, and the benchmark word vector is used to reflect the position of the one target word among the multiple target words;
[0187] The training module is further configured to:
[0188] Based on the reference word vector of the one target word and the benchmark word vector set for the one target word, determine the loss value corresponding to each target word, and the loss value is used to reflect the degree of difference between the reference word vector and the benchmark word vector; when the loss value corresponding to the one target word is less than the loss value threshold, determine that the representation probability corresponding to the one target word is less than the probability threshold;
[0189] When it is determined that the representation probability corresponding to the one target word is less than the probability threshold, update the initial word vectors of the 2m surrounding words of the one target word in the model parameters of the neural network model, including: based on the loss value corresponding to the one target word, update the initial word vectors of the 2m surrounding words of the one target word and the output word matrix.
[0190] Optionally, the information processing device further includes:
[0191] A fifth acquisition module, configured to acquire multiple entity names of a target domain before performing word segmentation on the text information to be processed based on a word segmentation dictionary of the target domain;
[0192] A second construction module, configured to construct a word segmentation dictionary of the target domain based on the multiple entity names.
[0193] In summary, in the information processing device provided in the embodiment of the present application, for the text information to be processed in the target domain, word segmentation processing can be performed on the text information to be processed based on the word segmentation dictionary of the target domain to obtain multiple words. In this way, the situation where the information of the same word is divided into two words when performing word segmentation on the text information in the target domain can be avoided, and the accuracy of word segmentation on the text information in the target domain is improved. And the target word vector in the word vector library based on which the target word vector is searched is determined by updating the initial model parameters of the neural network model based on the training text information of the target domain. The target word vector determined in this way represents the words in the target domain more accurately, and the accuracy of text processing on the text information to be processed based on the target word vector is relatively high, and the processing effect is good. And the initial model parameters are determined based on the initial word vectors with a dimension smaller than the number of target words in the word vector library. In this way, the determination process of the target word vector of the target word is also relatively simple and fast.
[0194] In an exemplary embodiment, an electronic device is further provided. The electronic device may include a processor and a memory, and at least one instruction is stored in the memory. The at least one instruction is configured to be executed by one or more processors to implement any one of the above information processing methods.
[0195] Figure 6 It is a schematic structural diagram of a server provided in the embodiment of the present application. The server may be the electronic device described in the above embodiment. As Figure 6 shown, the server 80 includes a central processing unit (CPU) 801, a system memory 804 including a random access memory (RAM) 802 and a read-only memory (ROM) 803, and a system bus 805 connecting the system memory 804 and the central processing unit 801. The server 800 also includes a basic input / output system (I / O system) 806 for transmitting information between various devices in the computer, and a mass storage device 807 for storing an operating system 813, application programs 814, and other program modules 815.
[0196] The basic input / output system 806 includes a display 808 for displaying information and input devices 809 such as a mouse, keyboard, etc. for user input of information. Both the display 808 and the input devices 809 are connected to the central processing unit 801 through an input / output controller 810 connected to the system bus 805. The basic input / output system 806 may also include an input / output controller 810 for receiving and processing inputs from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 810 also provides outputs to a display screen, printer, or other types of output devices.
[0197] The mass storage device 807 is connected to the central processing unit 801 through a mass storage controller (not shown) connected to the system bus 805. The mass storage device 807 and its associated computer-readable medium provide non-volatile storage for the server 800. That is, the mass storage device 807 may include a computer-readable medium (not shown) such as a hard disk or a CD-ROM drive. In the embodiments of the present application, the training data for training the model and the data of the model can be stored in this storage device.
[0198] Without loss of generality, computer-readable media may include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, EPROM, EEPROM, flash memory or other solid-state storage devices, CD-ROM, DVD or other optical storage, magnetic tape cartridges, tapes, magnetic disk storage or other magnetic storage devices. Of course, those skilled in the art will know that computer storage media is not limited to the above several types. The above-mentioned system memory 804 and mass storage device 807 may be collectively referred to as memory. The above memory also includes one or more programs, and one or more programs are stored in the memory and are configured to be executed by the CPU.
[0199] According to various embodiments of the present application, the server 80 may also run by connecting to a remote computer on the network through a network such as the Internet. That is, the server 80 may be connected to the network 812 through a network interface unit 811 connected to the system bus 805, or in other words, the network interface unit 811 may also be used to connect to other types of networks or remote computer systems (not shown).
[0200] Optionally, the training data for training the model and the data of the model in the embodiments of the present application can also be stored based on the blockchain. It should be noted that the blockchain is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain, in essence, is a decentralized database, a string of data blocks generated by using cryptographic methods, and each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. The blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer.
[0201] The blockchain underlying platform can include processing modules such as user management, basic services, smart contracts, and operation monitoring. Among them, the user management module is responsible for the identity information management of all blockchain participants, including maintaining the generation of public and private keys (account management), key management, and the maintenance of the corresponding relationship between the user's real identity and the blockchain address (permission management), and under authorized circumstances, supervising and auditing the transaction situations of certain real identities, providing the rule configuration for risk control (risk control audit); the basic service module is deployed on all blockchain node devices, used to verify the validity of business requests, and record them on the storage after consensus on the valid requests. For a new business request, the basic service first performs interface adaptation parsing and authentication processing (interface adaptation), then encrypts the business information through the consensus algorithm (consensus management), transmits it to the shared ledger intact and consistently after encryption (network communication), and records and stores it; the smart contract module is responsible for the registration and issuance of contracts, contract triggering, and contract execution. Developers can define contract logic through a certain programming language, publish it to the blockchain (contract registration), trigger the execution according to the logic of the contract terms, call keys or other events to complete the contract logic, and also provide functions for contract upgrade and cancellation; the operation monitoring module is mainly responsible for the deployment, configuration modification, contract setting, cloud adaptation during the product release process, and the visual output of the real-time status during product operation, such as: alarming, monitoring network conditions, monitoring the health status of node devices, etc.
[0202] The platform product service layer provides the basic capabilities and implementation frameworks of typical applications. Developers can build on these basic capabilities and overlay the characteristics of the business to complete the blockchain implementation of the business logic. The application service layer provides application services based on the blockchain solution for business participants to use.
[0203] In an exemplary embodiment, a computer-readable storage medium is also provided. At least one instruction, at least one program, a code set, or an instruction set is stored in the storage medium, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to perform any of the above information processing methods.
[0204] It should be noted that when the information processing device provided in the above embodiment obtains the target word vector, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. The method embodiments provided in the embodiments of the present application can be referred to each other with the corresponding device embodiments, and the embodiments of the present application do not make any limitations in this regard. The sequence of steps in the method embodiments provided in the embodiments of the present application can be appropriately adjusted, and the steps can also be increased or decreased accordingly. Any person skilled in the art can easily think of a changed method within the technical scope disclosed in the present application, and all such methods should be covered within the protection scope of the present application, so details will not be described herein.
[0205] In the embodiments of the present application, terms such as "first" and "second" are used to distinguish between identical or similar items with basically the same functions. It should be understood that there is no logical or temporal dependence between "first", "second", and "nth", nor are the quantity and execution order limited. In the embodiments of the present application, the meaning of the term "at least one" refers to one or more, and the meaning of the term "a plurality" in the present application refers to two or more. The term "and / or" used in the embodiments of the present application refers to and encompasses any and all possible combinations of one or more of the associated listed items. The term "and / or" is a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present application generally represents an "or" relationship between the preceding and following associated objects.
[0206] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the protection scope of the present application.
Claims
1. An information processing method, characterized in that, the method includes: obtaining text information to be processed in a target field; performing word segmentation processing on the text information to be processed based on a word segmentation dictionary in the target field to obtain multiple words in the text information to be processed; searching for target word vectors of the multiple words in a word vector library in the target field, where the target word vectors are used to represent the semantics of the corresponding words; performing text processing on the text information to be processed based on the target word vectors of the target words found in the multiple words; wherein, the word vector library in the target field includes: target word vectors of multiple target words belonging to the target field; the target word vector of each target word in the word vector library is determined based on model parameters of a neural network model, and the model parameters are obtained by training the neural network model with a training sample set obtained based on training text information in the target field to update initial model parameters of the neural network model; the training text information includes the multiple target words, the initial model parameters are determined based on initial word vectors configured for each of the multiple target words, and the dimension of the initial word vectors is less than the number of the multiple target words; the training sample set includes multiple training samples and training labels for each of the training samples; each training sample is used to reflect the semantics of the target words in a sample text, and the sample text is a sentence including at least two target words in the training text information; the training label of the training sample is used to reflect the positions of the target words in the sample text among the multiple target words.
2. The method according to claim 1, characterized in that, the method further includes: obtaining training text information in the target field; performing word segmentation processing on the training text information based on the word segmentation dictionary in the target field to obtain the multiple target words; configuring the initial word vectors for each of the multiple target words; obtaining the training sample set based on the training text information; training the neural network model based on each training sample in the training sample set and the corresponding training label to update the initial word vectors of the target words in the sample text corresponding to each training sample in the model parameters of the neural network model, so as to obtain a trained neural network model; obtaining the target word vectors of the multiple target words based on the model parameters of the trained neural network model to construct the word vector library in the target field.
3. The method according to claim 2, characterized in that, the configuring the initial word vectors for each of the multiple target words includes: performing random initialization of word vectors for each of the multiple target words based on the dimension of the input initial word vectors to configure the initial word vectors for each of the target words.
4. The method according to claim 2 or 3, characterized in that, Training the neural network model based on each training sample and the corresponding training label in the training sample set to update the initial word vector of the target word in the sample text corresponding to each training sample in the model parameters of the neural network model includes: Inputting each training sample in the training sample set into the neural network model, where each training sample includes the initial word vectors of 2m surrounding words of one target word among the multiple target words; where m≥1, and the 2m surrounding words include: among the sample text corresponding to the training sample, the m target words closest to and before the target word, and the m target words closest to and after the target word; Based on the initial word vectors of the 2m surrounding words of the one target word, outputting a reference word vector of the one target word through the neural network model; the reference word vector of the target word is used to reflect: the probability that the words represented by the initial word vectors of the 2m surrounding words of the target word are each target word among the multiple target words; When it is determined, based on the reference word vectors of each target word among the multiple target words and the training labels in the training set, that the representation probability corresponding to each target word among the multiple target words is greater than or equal to the probability threshold, stop training the neural network model; the representation probability corresponding to the target word is: the probability that the words represented by the initial word vectors of the 2m surrounding words of the target word are the target word; When it is determined, based on the reference word vector of the one target word and the training label corresponding to the one target word, that the representation probability corresponding to the one target word is less than the probability threshold, update the initial word vectors of the 2m surrounding words of the one target word in the model parameters of the neural network model; return to execute the step of outputting a reference word vector of the one target word through the neural network model based on the initial word vectors of the 2m surrounding words of the one target word.
5. The method according to claim 4, wherein, Outputting a reference word vector of the one target word through the neural network model based on the initial word vectors of the 2m surrounding words of the one target word includes: Through the neural network model, determining the vector obtained by adding and then averaging the initial word vectors of the 2m surrounding words of the one target word as the hidden vector of the one target word; Through the neural network model, determining the vector obtained by multiplying the hidden vector of the one target word by the output word matrix as the reference word vector of the one target word, where the output word matrix is used to reflect the mapping relationship from the 2m surrounding words to the one target word.
6. The method according to claim 5, wherein, The training label of the training sample includes a benchmark word vector set for the one target word, and the benchmark word vector is used to reflect the position of the one target word among the multiple target words; Training the neural network model based on each training sample and the corresponding training label in the training sample set to update the model parameters of the neural network model, and for the initial word vector of the target word in the sample text corresponding to the training sample, further comprising: Determining a loss value corresponding to each target word based on the reference word vector of the one target word and the benchmark word vector set for the one target word, where the loss value is used to reflect the degree of difference between the reference word vector and the benchmark word vector; When the loss value corresponding to the one target word is less than the loss value threshold, determining that the representation probability corresponding to the one target word is less than the probability threshold; When determining that the representation probability corresponding to the one target word is less than the probability threshold, updating the initial word vectors of 2m surrounding words of the one target word in the model parameters of the neural network model, including: Updating the initial word vectors of 2m surrounding words of the one target word and the output word matrix based on the loss value corresponding to the one target word.
7. The method according to any one of claims 1 to 3, characterized in that, before performing word segmentation processing on the to-be-processed text information based on the word segmentation dictionary of the target domain, the method further comprises: Obtaining a plurality of entity names of the target domain; Constructing the word segmentation dictionary based on the plurality of entity names.
8. An information processing apparatus, characterized in that, the information processing apparatus comprises: A first acquisition module, configured to acquire to-be-processed text information of a target domain; A first processing module, configured to perform word segmentation processing on the to-be-processed text information based on the word segmentation dictionary of the target domain to obtain a plurality of words in the to-be-processed text information; A lookup module, configured to look up target word vectors of the plurality of words in a word vector library of the target domain, where the target word vectors are used to represent the semantics of the corresponding words; A second processing module, configured to perform text processing on the to-be-processed text information based on the target word vectors of the target words found in the plurality of words; Among them, the word vector library of the target field includes: target word vectors of multiple target words belonging to the target field; the target word vectors of each target word in the word vector library are determined based on the model parameters of a neural network model, and the model parameters are obtained by training the neural network model with a training sample set acquired based on the training text information of the target field to update the initial model parameters of the neural network model; the training text information includes the multiple target words, the initial model parameters are determined based on the initial word vectors configured for each of the multiple target words, and the dimension of the initial word vectors is less than the number of the multiple target words; the training sample set includes multiple training samples and the training labels of each training sample; each training sample is used to reflect the semantics of the target words in a sample text, and the sample text is a sentence including at least two target words in the training text information; the training labels of the training samples are used to reflect the positions of the target words in the sample text among the multiple target words.
9. An electronic device, characterized in that the electronic device includes: a processor and a memory, and at least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the information processing method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that at least one instruction, at least one program, a code set or an instruction set is stored in the storage medium, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the information processing method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Supervised word vector training method and device
CN108776655A
Keyword generation method, apparatus and device, and storage medium
CN112364136A