Semantic recognition method and device, electronic equipment, storage medium and product
By inputting the user's personalized voice data into the target large language model, and using the user's personalized voice feature knowledge that passed the verification for semantic recognition, the problem that the existing technology cannot accurately recognize strong personalized voice is solved, and the accurate recognition of user intentions is achieved.
Patent Information
- Application Number
- CN202510164924.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art cannot accurately recognize semantics of strong personalized speech, especially local dialect pronunciation and grammatical errors.
By obtaining the user's personalized voice data and inputting it into the target large language model for semantic recognition, the target large language model is recognized based on the user's personalized voice feature knowledge that passes the verification. The feature knowledge includes the user's personalized voice data, semantic recognition results, and speech context data.
Accurate semantic recognition of strong personalized speech is achieved, and user intentions can be accurately identified, solving the problem that cannot be recognized by the prior art.
Smart Images

Figure CN120106077A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of artificial intelligence technology, and in particular, to a semantic recognition method, device, electronic device, storage medium and product. Background Art
[0002] A very important link in the voice dialogue system is natural language processing. Natural language processing is an important direction in the fields of computer science and artificial intelligence. It mainly studies various theories and methods that can enable effective communication between people and computers using natural language.
[0003] Natural language processing mainly includes semantic recognition, which can accurately understand the user's intention. Existing semantic recognition methods mainly perform semantic recognition on standard voice data. If the user's voice data is highly personalized, such as local dialect voice, grammatically incorrect voice; for such user's voice data, existing voice recognition methods cannot accurately recognize the user's intention. Summary of the invention
[0004] The present invention provides a semantic recognition method, device, electronic device, storage medium and product to solve the problem that the prior art cannot accurately perform semantic recognition on highly personalized speech.
[0005] According to one aspect of the present invention, there is provided a semantic recognition method, comprising:
[0006] Obtain the user's personalized voice data;
[0007] Inputting the personalized voice data into a target large language model to obtain a final semantic recognition result of the user;
[0008] The target large language model performs semantic recognition based on verified user personalized speech feature knowledge, wherein the user personalized speech feature knowledge includes the user's personalized speech data, semantic recognition results, and speech context data.
[0009] According to another aspect of the present invention, there is provided a semantic recognition device, comprising:
[0010] An acquisition module is used to acquire the user's personalized voice data;
[0011] A recognition module, used for inputting the personalized voice data into a target large language model to obtain a final semantic recognition result of the user;
[0012] The target large language model performs semantic recognition based on verified user personalized speech feature knowledge, wherein the user personalized speech feature knowledge includes the user's personalized speech data, semantic recognition results, and speech context data.
[0013] According to another aspect of the present invention, there is provided an electronic device, the electronic device comprising: at least one processor;
[0014] and a memory communicatively coupled to the at least one processor;
[0015] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the semantic recognition method described in any embodiment of the present invention.
[0016] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the semantic recognition method described in any embodiment of the present invention when executed.
[0017] According to another aspect of the present invention, a computer program product is provided. The computer program product includes a computer program. When the computer program is executed by a processor, the semantic recognition method described in any embodiment of the present invention is implemented.
[0018] The technical solution of the embodiment of the present invention solves the problem that the prior art is unable to perform accurate semantic recognition of strong personalized voices, by inputting the personalized voice data into the target large language model to perform semantic recognition based on verified user personalized voice feature knowledge, and achieves the beneficial effect of being able to accurately recognize user intentions.
[0019] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0021] Figure 1 A flowchart of a semantic recognition method provided in Embodiment 1 of the present invention;
[0022] Figure 2 A flowchart of a semantic recognition method provided in Embodiment 2 of the present invention;
[0023] Figure 3A schematic diagram of the structure of a semantic recognition device provided in Embodiment 3 of the present invention;
[0024] Figure 4 The present invention is a schematic diagram of the structure of an electronic device for a semantic recognition method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0025] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiment of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiment of the present invention. Obviously, the described embodiment is only an embodiment of a part of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work should belong to the scope of protection of the present invention. It should be understood that the various steps recorded in the method implementation of the present invention can be performed in different orders and / or in parallel. In addition, the method implementation may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0026] The term "including" and its variations used herein are open inclusions, i.e., "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0028] It should be noted that the modifications of "one" and "plurality" mentioned in the present invention are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0029] The names of the messages or information exchanged between multiple devices in the embodiments of the present invention are only used for illustrative purposes, and are not used to limit the scope of these messages or information.
[0030] Embodiment 1
[0031] Figure 1 A flow chart of a semantic recognition method provided in Embodiment 1 of the present invention is applicable to human-computer dialogue scenarios, and the situation of semantic recognition of the user's non-standardized speech. The method can be executed by a semantic recognition device, wherein the device can be implemented by software and / or hardware and is generally integrated on an electronic device. In this embodiment, the electronic device includes but is not limited to: a computer device.
[0032] like Figure 1 As shown, a semantic recognition method provided by Embodiment 1 of the present invention includes the following steps:
[0033] S110: Obtain personalized voice data of the user.
[0034] Among them, personalized voice data can be understood as non-standardized voice data with user speaking characteristics. Different users have different speaking characteristics. For example, some users have confusing speech logic, some users speak without a subject, some users speak in a long-winded manner, some users speak too briefly, and some users speak in local dialects.
[0035] Optionally, the user's personalized voice data includes one or more of the following:
[0036] Speech data without subject; speech data in local dialect; complicated speech data; speech data without logic; short speech data.
[0037] In this embodiment, the user's personalized voice data can be obtained in a variety of ways. One feasible way is to collect the user's personalized voice data through the built-in microphone of the electronic device in a human-computer interaction scenario; another feasible way is to collect the user's personalized voice data through the voice collection function of the operating system; and yet another feasible way is to collect the user's personalized voice data through third-party voice collection software and the voice collection function of cloud services.
[0038] S120: Input the personalized voice data into a target large language model to obtain a final semantic recognition result of the user.
[0039] The target large language model performs semantic recognition based on verified user personalized speech feature knowledge, wherein the user personalized speech feature knowledge includes the user's personalized speech data, semantic recognition results, and speech context data.
[0040] In this embodiment, the target large language model is a large language model obtained after model training and model verification. Preferably, the large language model may be a large language model with small parameters.
[0041] The target large language model is trained and verified based on the personalized speech feature knowledge of multiple users. The semantic recognition results and speech context data included in the personalized speech feature knowledge of the users are obtained by inputting the personalized speech data of the users into the ultra-large parameter large language model for semantic analysis and then output.
[0042] The user personalized speech feature knowledge of multiple users is verified and marked with a verification success label. The verification process of the user personalized speech feature knowledge includes verifying the correctness, language rationality and logical correctness of the user's personalized speech data, semantic recognition results and speech context data.
[0043] In this embodiment, the target large language model can be deployed in the natural language processing branch of the voice dialogue system. After the target large language model is deployed, on the one hand, the target large language model can be used to perform semantic recognition on the user's voice data acquired in real time, and on the other hand, the user's voice data can be used for self-learning of the model.
[0044] A semantic recognition method provided in the first embodiment of the present invention first obtains the personalized voice data of the user; then the personalized voice data is input into the target large language model to obtain the final semantic recognition result of the user; wherein the target large language model performs semantic recognition based on the verified personalized voice feature knowledge of the user, and the personalized voice feature knowledge of the user includes the personalized voice data of the user, the semantic recognition result and the voice context data. The above method performs semantic recognition based on the personalized voice feature knowledge of the user, and can accurately identify the user's intention based on the personalized voice data of the user.
[0045] Based on the above embodiment, a variant embodiment of the above embodiment is proposed. It should be noted that in order to make the description concise, only the differences from the above embodiment are described in the variant embodiment.
[0046] In one embodiment, the process of generating user personalized speech feature knowledge includes:
[0047] If the personalized voice data is valid personalized voice data, determining whether there is historical voice data corresponding to the valid personalized voice data in the user voice knowledge base;
[0048] If yes, input the prompt information template, the effective personalized voice data and the historical voice data into a first large language model, analyze the personalized effective voice data and the historical voice data based on the prompt information template through the first large language model, and output a first semantic recognition result and a first voice context data;
[0049] The effective personalized voice data, the first semantic recognition result and the first voice context data are established in a corresponding relationship and stored in the user voice knowledge base.
[0050] It can be understood that the above-mentioned process of generating user personalized speech feature knowledge is the process of generating user personalized speech feature knowledge of one user, and the process of generating personalized language feature knowledge of each user is the same, which will not be repeated here.
[0051] In this embodiment, the obtained personalized voice data of the user needs to be judged whether it is valid voice data to ensure that the personalized voice data of the user is voice data with normal semantics. If the personalized voice data of the user does not have normal semantics, it cannot be used for model self-learning. For example, if the personalized voice data of the user is an interjection, such as "ah", "um", or "oh", it is invalid voice data.
[0052] Among them, the first large language model is preferably a large language model with ultra-large parameters.
[0053] The historical voice data is the historical voice data of the user to whom the effective personalized voice belongs, and the historical voice data is obtained from the user voice knowledge base.
[0054] In this embodiment, the first large language model analyzes the effective personalized voice data and the historical voice data according to the prompt information template and outputs a semantic recognition result as the first semantic recognition result, as well as voice context data.
[0055] Among them, the prompt information template is a template of a prompt language with a specific format, structure and expression method, which is equivalent to a pre-set standard framework. Users can fill it with specific content, such as intentions and real needs, so that the first language model can execute and output according to the prompt content formed by the prompt information template.
[0056] The speech context data is the speech data related to the effective personalized speech data. The speech context data refers to the related speech data that appears before and after the effective personalized speech data. The related speech data helps to understand the meaning, intention, emotion of the effective personalized speech data and its role in the entire speech communication scenario. The related context data is obtained by analyzing the historical speech data.
[0057] Exemplarily, if the effective personalized voice data does not contain pronouns, it is possible to confirm from the historical voice data whether it includes the subject or object of the pronoun. For example, if the effective personalized voice data is "turn up a bit", the subject of "turn up a bit" can be analyzed based on the historical voice data. If the historical voice data includes voice-related content, it is analyzed that the subject is the music volume, and the action is to turn it up. For example, if the effective personalized voice data is "the screen is a bit dark", it is analyzed that the subject is the screen brightness, and the action is to turn it up.
[0058] In this embodiment, the effective personalized voice data, the first semantic recognition result and the voice context data may be stored in the user voice knowledge base after establishing a correspondence relationship, so as to be used for self-learning of the target large language model.
[0059] Further, if not, the effective personalized voice data is input into the first language model for semantic analysis to output a second semantic recognition result; a corresponding relationship is established between the effective personalized voice data and the second semantic recognition result and then stored in the user voice knowledge base.
[0060] If the effective personalized voice data has no corresponding historical voice data, the first language model can directly perform ordinary analysis on the effective personalized voice data and output a semantic recognition result as the second semantic recognition result.
[0061] In this embodiment, the effective personalized voice data and the second semantic recognition result may be stored in the user voice knowledge base after establishing a correspondence relationship, so as to be used for self-learning of the target large language model.
[0062] In one embodiment, the verification process of the user's personalized voice feature knowledge includes:
[0063] Determining whether the amount of user personalized voice feature knowledge in the user voice knowledge base reaches a preset value;
[0064] If so, all the user personalized voice feature knowledge in the user voice knowledge base is input into the second largest language model, and the correctness, logic and completeness of all the user personalized voice feature knowledge are analyzed by the second largest language model, and the verification result is output.
[0065] The second largest language model is a large language model with huge parameters, and the number of parameters of the second largest language model is greater than the number of parameters of the first largest language model.
[0066] In this embodiment, a preset value is set in advance. When the amount of user personalized voice feature knowledge in the user voice knowledge base reaches the preset value, all user personalized voice feature knowledge in the user voice knowledge base can be input into the second largest language model for verification, and the verification result is output.
[0067] Among them, the second largest language model can analyze whether the user's personalized voice feature knowledge is correct and whether there are any unreasonable aspects in terms of language, logic, etc. If it is correct, the user's personalized voice feature knowledge will be marked as passed. The user's personalized voice feature knowledge that is not marked needs to wait for manual re-verification.
[0068] In one embodiment, the target large language model is obtained by self-learning the third large language model, and the number of parameters of the third large language model is less than the number of parameters of the first large language model;
[0069] The self-learning process of the third language model includes:
[0070] The verified user personalized voice feature knowledge is format-converted and divided into training set data and verification set data;
[0071] Inputting the training set data into the third language model for fine-tuning training to obtain a fine-tuned third language model;
[0072] Inputting the validation set data into the fine-tuned third language model for model validation;
[0073] If the verification result meets the preset standard, the fine-tuned third language model is used as the target large language model;
[0074] If the verification result does not meet the preset standard, new training set data is obtained to perform fine-tuning training on the fine-tuned third language model until the verification result meets the preset standard.
[0075] Preferably, the third largest language model is a large language model with small parameters.
[0076] In this embodiment, a portion of the verified user personalized voice feature knowledge is converted according to the requirements of the training set format and used as the training set data; another portion of the verified user personalized voice feature knowledge is converted according to the requirements of the verification set format and used as the verification set. For example, the voice data in the user personalized voice feature knowledge is divided into words or subwords one by one, and the voice data is marked as a series of words or subwords using a marker, each word or subword corresponds to a unique number, and the third language model processes the voice data through these numbers.
[0077] In this embodiment, the process of inputting the training set data into the third language model for fine-tuning training includes:
[0078] Data input: According to the set batch size, the training set data is input batch by batch into the third language model;
[0079] Forward propagation: The training set data is calculated in the third language model according to the network structure to obtain the model output; the training set data is calculated through the multi-layer Transformer architecture and finally outputs the semantic recognition result;
[0080] Calculate loss: Based on the real label data output by the third language model, use the selected loss function to calculate the loss value; this loss value represents the gap between the semantic recognition result output by the model and the real result;
[0081] Back propagation: Based on the calculated loss value, the model parameters are updated through the back propagation algorithm; starting from the output layer, the loss value is back propagated layer by layer, the gradient of each parameter is calculated according to the chain rule, and then the optimization algorithm is used to update the parameters.
[0082] In this embodiment, the validation set data is input into the fine-tuned third largest language model for verification. If the accuracy of the fine-tuned third largest language model on the validation set reaches a stable value, it means that the model has been trained and the target large language model is obtained. If the accuracy of the fine-tuned third largest language model on the validation set does not reach the preset value, it is necessary to obtain new training set data, and use the new training set data to repeat the above-mentioned forward propagation, loss calculation, and back propagation processes for the fine-tuned third largest language model until the verification result meets the preset standard.
[0083] Embodiment 2
[0084] The embodiment of the present invention provides a specific implementation method based on the technical solutions of the above embodiments.
[0085] As a specific implementation method of this embodiment, Figure 2 A flowchart of a semantic recognition method provided in the second embodiment of the present invention is shown as follows: Figure 2 As shown, the following process is included:
[0086] The user's personalized voice data is distributed to the analysis processing module and the self-learning module at the same time through the cloud-based semantic analysis service of the voice dialogue system; the analysis processing module uses the large language model to analyze the user's personalized voice data and returns the semantic recognition result for the client to execute; the self-learning module uses the user's personalized voice data to complete model self-learning to obtain the target large language model; the target large language model is deployed to the analysis processing module to use the target large language model to perform semantic recognition on the personalized voice data input by the user.
[0087] Among them, the process of the self-learning module using the user's personalized voice data to complete the model self-learning to obtain the target large language model includes:
[0088] Generate user personalized voice feature knowledge: determine whether the user's personalized voice data is valid personalized voice data; if so, input the user's personalized voice data into the ultra-large parameter large language model to output semantic recognition results and voice context data; establish a one-to-one correspondence between the user's personalized voice data, semantic recognition results and voice context data to obtain the user's personalized voice feature knowledge, and store it in the user voice knowledge base;
[0089] Verify the user's personalized speech feature knowledge: When the user's personalized speech feature knowledge in the user speech knowledge base reaches a certain amount, all the user's personalized speech feature knowledge in the user speech knowledge base is input into a large language model with huge parameters to verify the correctness of the user's personalized speech feature knowledge;
[0090] Model self-learning is performed based on user personalized voice feature knowledge: the verified user personalized voice feature knowledge is divided into a training set and a validation set; the training set is used to fine-tune the large language model with small parameters, and the validation set is used to verify the fine-tuned model. After verification, the target large language model is obtained.
[0091] Embodiment 3
[0092] Figure 3 This is a structural schematic diagram of a semantic recognition device provided in Example 3 of the present invention. The device can be applicable to human-computer dialogue scenarios and performs semantic recognition on the user's non-standardized voice. The device can be implemented by software and / or hardware and is generally integrated on an electronic device.
[0093] like Figure 3 As shown, the device includes: an acquisition module 110 and an identification module 120.
[0094] An acquisition module 110 is used to acquire personalized voice data of a user;
[0095] Recognition module 120, used for inputting the personalized voice data into a target large language model to obtain a final semantic recognition result of the user;
[0096] The target large language model performs semantic recognition based on verified user personalized speech feature knowledge, wherein the user personalized speech feature knowledge includes the user's personalized speech data, semantic recognition results, and speech context data.
[0097] In this embodiment, the device first obtains the user's personalized voice data through the acquisition module 110; then inputs the personalized voice data into the target large language model through the recognition module 120 to obtain the final semantic recognition result of the user; wherein the target large language model performs semantic recognition based on the verified user's personalized voice feature knowledge, and the user's personalized voice feature knowledge includes the user's personalized voice data, semantic recognition results and voice context data.
[0098] This embodiment provides a semantic recognition device that can accurately recognize the user's intention from the user's personalized voice data.
[0099] Furthermore, the process of generating the user personalized speech feature knowledge includes:
[0100] If the personalized voice data is valid personalized voice data, determining whether there is historical voice data corresponding to the valid personalized voice data in the user voice knowledge base;
[0101] If yes, input the prompt information template, the effective personalized voice data and the historical voice data into a first large language model, analyze the personalized effective voice data and the historical voice data based on the prompt information template through the first large language model, and output a first semantic recognition result and voice context data;
[0102] The effective personalized voice data, the first semantic recognition result and the voice context data are established in a corresponding relationship and stored in the user voice knowledge base.
[0103] On the basis of the above optimization, it also includes: if not, inputting the effective personalized voice data into the first large language model for semantic analysis to output a second semantic recognition result; establishing a corresponding relationship between the effective personalized voice data and the second semantic recognition result and storing them in the user voice knowledge base.
[0104] Based on the above technical solution, the verification process of the user's personalized voice feature knowledge includes:
[0105] Determining whether the amount of user personalized voice feature knowledge in the user voice knowledge base reaches a preset value;
[0106] If yes, all the user personalized voice feature knowledge in the user voice knowledge base is input into the second largest language model, and the correctness, logic and completeness of all the user personalized voice feature knowledge are analyzed by the second largest language model, and the verification result is output;
[0107] The number of parameters of the second largest language model is greater than the number of parameters of the first largest language model.
[0108] Furthermore, the target large language model is obtained by self-learning the third large language model, and the number of parameters of the third large language model is less than the number of parameters of the first large language model;
[0109] The self-learning process of the third language model includes:
[0110] The verified user personalized voice feature knowledge is format-converted and divided into training set data and verification set data;
[0111] Inputting the training set data into the third language model for fine-tuning training to obtain a fine-tuned third language model;
[0112] Inputting the validation set data into the fine-tuned third language model for model validation;
[0113] If the verification result meets the preset standard, the fine-tuned third language model is used as the target large language model;
[0114] If the verification result does not meet the preset standard, new training set data is obtained to perform fine-tuning training on the fine-tuned third language model until the verification result meets the preset standard.
[0115] Furthermore, the user's personalized voice data includes one or more of the following:
[0116] Speech data without subject; speech data in local dialect; complicated speech data; speech data without logic; short speech data.
[0117] The above-mentioned semantic recognition device can execute the semantic recognition method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0118] Embodiment 4
[0119] Figure 4 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.
[0120] like Figure 4As shown, the electronic device 10 includes at least one processor 11, and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0121] A number of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0122] The processor 11 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as a semantic recognition method.
[0123] In some embodiments, a semantic recognition method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps in the semantic recognition method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to perform a semantic recognition method in any other appropriate manner (e.g., by means of firmware).
[0124] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0125] In some embodiments, the semantic recognition method can be implemented as a computer program, which is invisibly included in a computer program product. The computer program implements the semantic recognition method of the present invention when executed by a processor. The computer program product can be understood as a software product that mainly implements its solution through a computer program. The computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the computer program, when executed by the processor, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The computer program can be executed entirely on the machine, partially on the machine, as an independent software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0126] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in combination with an instruction execution system, device or equipment. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0127] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).
[0128] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0129] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.
[0130] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.
[0131] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A semantic recognition method, characterized in that: The method comprises: Obtain the user's personalized voice data; Inputting the personalized voice data into a target large language model to obtain a final semantic recognition result of the user; The target large language model performs semantic recognition based on verified user personalized speech feature knowledge, wherein the user personalized speech feature knowledge includes the user's personalized speech data, semantic recognition results, and speech context data.
2. The method according to claim 1, characterized in that The process of generating the user personalized speech feature knowledge includes: If the personalized voice data is valid personalized voice data, determining whether there is historical voice data corresponding to the valid personalized voice data in the user voice knowledge base; If yes, input the prompt information template, the effective personalized voice data and the historical voice data into a first large language model, analyze the personalized effective voice data and the historical voice data based on the prompt information template through the first large language model, and output a first semantic recognition result and voice context data; The effective personalized voice data, the first semantic recognition result and the voice context data are established in a corresponding relationship and stored in the user voice knowledge base.
3. The method according to claim 2, characterized in that Also includes: If not, inputting the valid personalized voice data into the first language model for semantic analysis to output a second semantic recognition result; The effective personalized voice data and the second semantic recognition result are established in a corresponding relationship and stored in the user voice knowledge base.
4. The method according to claim 2, characterized in that: The verification process of the user's personalized voice feature knowledge includes: Determining whether the amount of user personalized voice feature knowledge in the user voice knowledge base reaches a preset value; If yes, all the user personalized voice feature knowledge in the user voice knowledge base is input into the second largest language model, and the correctness, logic and completeness of all the user personalized voice feature knowledge are analyzed by the second largest language model, and the verification result is output; The number of parameters of the second largest language model is greater than the number of parameters of the first largest language model.
5. The method according to claim 1, characterized in that The target large language model is obtained by self-learning the third large language model, and the number of parameters of the third large language model is less than the number of parameters of the first large language model; The self-learning process of the third language model includes: The verified user personalized voice feature knowledge is format-converted and divided into training set data and verification set data; Inputting the training set data into the third language model for fine-tuning training to obtain a fine-tuned third language model; Inputting the validation set data into the fine-tuned third language model for model validation; If the verification result meets the preset standard, the fine-tuned third language model is used as the target large language model; If the verification result does not meet the preset standard, new training set data is obtained to perform fine-tuning training on the fine-tuned third language model until the verification result meets the preset standard.
6. The method according to claim 1, characterized in that The user's personalized voice data includes one or more of the following: Speech data without subject; speech data in local dialect; complicated speech data; speech data without logic; short speech data.
7. A semantic recognition device, characterized in that: The device comprises: An acquisition module is used to acquire the user's personalized voice data; A recognition module, used for inputting the personalized voice data into a target large language model to obtain a final semantic recognition result of the user; The target large language model performs semantic recognition based on verified user personalized speech feature knowledge, wherein the user personalized speech feature knowledge includes the user's personalized speech data, semantic recognition results, and speech context data.
8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the semantic recognition method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the semantic recognition method according to any one of claims 1 to 6 when executed.
10. A computer program product, characterized in that The computer program product comprises a computer program, which, when executed by a processor, implements the semantic recognition method according to any one of claims 1 to 6.