Voice instruction response method and device based on large model fine tuning, equipment and medium

By fine-tuning the pre-trained large language model and combining the instruction response data set, a voice command response method based on the large model is realized, solving the problem of insufficient accuracy and real-time sound triggering of voice commands in the prior art, and improving user experience and response efficiency.

CN119943046APending Publication Date: 2025-05-06GUANGZHOU BAOLUN ELECTRONICS CO LTD

Patent Information

Application Number
CN202510101466.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing intelligent voice interaction methods have shortcomings in the accuracy and real-time performance of voice command triggering. Users need to remember a lot of voice commands, and their response efficiency is low and the user experience is poor.

Method used

The speech instruction response method based on large model fine-tuning is adopted to fine-tune the pre-trained large language model through the instruction response data set to realize local instruction response capabilities, and the text understanding and analysis capabilities of the large model are used for instruction fuzzy matching and joint instruction response.

Benefits of technology

It improves the accuracy and efficiency of voice commands, reduces the difficulty of user memory trigger conditions, enhances user experience, and realizes the effects of fuzzy matching and joint response of instructions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943046A_ABST
    Figure CN119943046A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses a voice instruction response method and device based on large model fine tuning, equipment and a medium, and the method comprises the steps: collecting a first voice signal, and carrying out the analysis of the first voice signal, and obtaining voice information; judging whether the voice information is matched with a preset wake-up word or not; if the voice information is matched with the preset wake-up word, waking up a voice instruction response module, collecting a second voice signal, and converting the second voice signal into text information through a voice recognition model; analyzing and processing the text information by utilizing an instruction response large model, and outputting instruction response information corresponding to a plurality of target instructions; wherein the instruction response large model is obtained by finely adjusting a pre-trained large language model according to an instruction response data set; and triggering a voice instruction response device based on large model fine tuning through the application interface to execute corresponding response operation according to the instruction response information corresponding to each target instruction. By applying the technical scheme of the invention, the accuracy and efficiency of processing the voice instruction can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of artificial intelligence technology, and specifically to a voice command response method, device, equipment and computer-readable storage medium based on large model fine-tuning. Background Art

[0002] The existing intelligent voice interaction method pre-sets all commands as voice trigger conditions, and has high requirements for voice command trigger conditions. The corresponding command can only be matched when the input audio is exactly the same as the preset audio, and the response rate is low. When there are many voice commands, users need to accurately remember these cumbersome and large number of voice commands, and users need to highlight the differences between voice commands to avoid false triggering problems. The numerous trigger commands also increase the difficulty of use for users. Therefore, the command triggering accuracy of the existing voice command interaction method is low and the real-time performance is poor, resulting in poor user experience.

[0003] In addition, existing technologies also usually use natural language processing technology to perform a series of operations such as intent analysis, demand matching and information retrieval on user text to obtain response results. The overall process is relatively cumbersome and heavily dependent on response information at each stage. The instruction response efficiency is low and the accuracy is not high. Summary of the invention

[0004] In view of the above problems, the embodiments of the present invention provide a voice command response method, apparatus, device and computer-readable storage medium based on large model fine-tuning, which are used to improve the accuracy and efficiency of processing voice commands.

[0005] According to one aspect of an embodiment of the present invention, a voice command response method based on large model fine-tuning is provided, the method comprising:

[0006] Collecting a first voice signal, analyzing the first voice signal, and obtaining voice information;

[0007] Determine whether the voice information matches a preset wake-up word; the preset wake-up word is pre-set by the user;

[0008] If the voice information matches the preset wake-up word, then wake up the voice command response module;

[0009] When the voice command response module is awakened, a second voice signal is collected, and the second voice signal is converted into text information through a voice recognition model;

[0010] The text information is analyzed and processed by using the command response big model, and command response information corresponding to a number of target commands is output; the command response information corresponding to the target command includes: the command ID marked by the target command and the command reply information of the target command; wherein the command response big model is obtained by fine-tuning the pre-trained big language model according to the command response data set; the command ID is pre-bound to the application interface of the voice command response device fine-tuned based on the big model;

[0011] The voice command response device is triggered through the application interface to perform a corresponding response operation according to the command response information corresponding to each target command.

[0012] The voice command response method based on large model fine-tuning provided by the present invention uses the command response data set to fine-tune the LLM to obtain a fine-tuned large language model, so that it has local command response capabilities, and at the same time, with the help of the text understanding and analysis capabilities of the large model, the effect of command fuzzy matching and joint command response is achieved. At the same time, for the numerous commands of the voice command response device based on large model fine-tuning, the user can only perform voice interaction based on the names of tools and applications in the interface display of the voice command response device based on large model fine-tuning, and can effectively trigger the function command without accurately remembering the trigger conditions, thereby reducing the difficulty of user use.

[0013] In an optional manner, determining whether the voice information matches a preset wake-up word further includes:

[0014] If the judgment result matches the preset wake-up word, a voice response is triggered;

[0015] If the judgment result is that it does not match the preset wake-up word, wait silently.

[0016] In an optional manner, after waking up the voice command response module, the method further includes:

[0017] When no third voice signal is received within a first preset time after the second voice signal is collected, causing the voice command response module to sleep;

[0018] The voice command response module cannot be awakened again during the second preset sleep time.

[0019] In an optional manner, the text information is analyzed and processed using a large instruction response model to output instruction response information corresponding to a number of target instructions, wherein the large instruction response model is obtained by fine-tuning a pre-trained large language model according to an instruction response data set, and includes:

[0020] The command response data set includes a plurality of preset command response data pairs;

[0021] The plurality of preset instruction response data pairs include: text instruction information of a plurality of preset instructions and corresponding preset response information; the preset response information includes: the instruction ID marked by the preset instruction and the instruction reply information of the preset instruction;

[0022] The command response data set is a collection of text command information of several preset commands and their corresponding preset response information, which is summarized and sorted according to the functions of the voice command response device;

[0023] In the process of fine-tuning the pre-trained large language model, the instruction response large model learns the relationship between the text instruction information of each of the preset instructions and its corresponding preset response information;

[0024] The command response large model is used to understand and analyze the text information to match the corresponding target command response data pair, and the target command response data pair is analyzed based on the learned knowledge to obtain the target command and the corresponding command response information, and output a number of command response information corresponding to the target command; the target command response data pair is any one of a number of preset command response data pairs.

[0025] In an optional manner, fine-tuning the pre-trained large language model according to the command response dataset includes:

[0026] The text instruction information of the preset instruction in the instruction response data set is used as the prompt word of the pre-trained large language model, and the preset response information corresponding to the text instruction information of the preset instruction is used as the expected response parameter of the pre-trained large language model to fine-tune the pre-trained large language model.

[0027] In an optional manner, the instruction response macromodel outputs instruction response information corresponding to a plurality of target instructions in a preset format;

[0028] The command response information in the preset format is a json object.

[0029] In an optional manner, triggering the voice command response device through the application interface to perform a corresponding response operation according to the command response information corresponding to each of the target commands also includes:

[0030] Execute the corresponding instruction operation according to the instruction ID marked by the target instruction;

[0031] The command reply information of the target command is synchronously played in the form of voice through the voice synthesis module.

[0032] According to another aspect of an embodiment of the present invention, there is provided a voice command response device based on large model fine-tuning, comprising: a voice acquisition module, a wake-up module and a response module;

[0033] The voice collection module is used to collect the first voice signal, analyze the first voice signal, and obtain voice information;

[0034] The wake-up module is used to determine whether the voice information matches a preset wake-up word; the preset wake-up word is pre-set by the user; if the voice information matches the preset wake-up word, the voice command response module is awakened;

[0035] The response module is used to collect the second voice signal when waking up the voice command response module, and convert the second voice signal into text information through the voice recognition model;

[0036] The text information is analyzed and processed by using the command response big model, and command response information corresponding to a number of target commands is output; the command response information corresponding to the target command includes: the command ID marked by the target command and the command reply information of the target command; wherein the command response big model is obtained by fine-tuning the pre-trained big language model according to the command response data set; the command ID is pre-bound to the application interface of the voice command response device fine-tuned based on the big model;

[0037] The voice command response device is triggered through the application interface to perform a corresponding response operation according to the command response information corresponding to each target command.

[0038] According to another aspect of an embodiment of the present invention, there is provided a voice command response device based on large model fine-tuning, comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; the memory is used to store at least one executable instruction, wherein the executable instruction enables the processor to perform the operation of the voice command response method based on large model fine-tuning as described in any one of the above items.

[0039] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is provided, wherein the storage medium stores at least one executable instruction, wherein the executable instruction enables a voice command response device / apparatus based on large model fine-tuning to perform the operation of the voice command response method based on large model fine-tuning as described in any one of the above.

[0040] The above description is only an overview of the technical solution of the embodiment of the present invention. In order to more clearly understand the technical means of the embodiment of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the embodiment of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The accompanying drawings are only used to illustrate the embodiments and are not to be considered as limiting the present invention. In addition, the same reference symbols are used to represent the same components throughout the accompanying drawings. In the accompanying drawings:

[0042] Figure 1 A flow chart of an embodiment of a voice command response method based on large model fine-tuning provided by the present invention is shown;

[0043] Figure 2 A structural schematic diagram of an embodiment of a voice command response device based on large model fine-tuning provided by the present invention is shown;

[0044] Figure 3 A structural schematic diagram of an embodiment of a voice command response device based on large model fine-tuning provided by the present invention is shown. DETAILED DESCRIPTION

[0045] Exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited to the embodiments set forth herein.

[0046] Embodiment 1, Figure 1 The flowchart of an embodiment of the voice command response method based on large model fine-tuning provided by the present invention is shown, and the method is executed by a voice command response device based on large model fine-tuning. In this embodiment, the voice command response device based on large model fine-tuning is an all-in-one machine, which can be used for the voice command response method and system based on large model fine-tuning in scenes such as conferences and education. Figure 1 As shown, the method comprises the following steps:

[0047] Step 110: Collect a first voice signal, analyze the first voice signal, and obtain voice information.

[0048] In some embodiments of the present invention, a microphone configured in the integrated device is used to collect the first voice signal in real time, and the voice information in the first voice signal is synchronously monitored through the voice wake-up module.

[0049] Preferably, the voice wake-up module is used to monitor the voice information in the voice signal collected by the microphone in real time.

[0050] Step 120: Determine whether the voice information matches a preset wake-up word; the preset wake-up word is pre-set by the user; if the voice information matches the preset wake-up word, wake up the voice command response module.

[0051] Specifically, the preset wake-up words can be customized according to the personalized attributes of the application. Usually 4-6 Chinese characters are sufficient. The wake-up words should be of different syllables as much as possible and have large pronunciation differences, which can effectively improve the wake-up rate.

[0052] In some embodiments of the present invention, determining whether the voice information matches a preset wake-up word further includes:

[0053] If the judgment result matches the preset wake-up word, a voice response is triggered;

[0054] If the judgment result is that it does not match the preset wake-up word, wait silently.

[0055] Specifically, the announcement in the form of voice includes using the speaker of the all-in-one machine to announce a preset voice signal. For example, if the voice information matches the preset wake-up word, a voice signal of "I am here" is issued by the speaker, indicating that the voice command response system has been awakened. If the voice information does not match the preset wake-up word, then wait silently.

[0056] In some embodiments of the present invention, after waking up the voice command response module, the method further includes:

[0057] When no third voice signal is received within a first preset time after the second voice signal is collected, causing the voice command response module to sleep;

[0058] The voice command response module cannot be awakened again during the second preset sleep time.

[0059] Preferably, the preset duration of voice recognition is 3s. The all-in-one machine uses the configured microphone to collect the second voice signal, each time for 3s. Subsequent voice commands need to re-awaken the system, and a time difference is required between consecutive wake-ups of the system to avoid malfunctions caused by frequent system awakenings.

[0060] Step 130: When the voice command response module is awakened, a second voice signal is collected, and the second voice signal is converted into text information through a voice recognition model; the text information is analyzed and processed using a large command response model, and command response information corresponding to a number of target commands is output; the command response information corresponding to the target command includes: a command ID marked by the target command and command reply information of the target command; wherein the large command response model is obtained by fine-tuning a pre-trained large language model according to a command response data set; the command ID is pre-bound to an application interface of a voice command response device fine-tuned based on the large model.

[0061] Specifically, large language models (LLMs), such as Llama, CHATGLM, and Tongyi Qianwen, are good at semantic understanding and analysis and are commonly used in intelligent dialogue application systems, but they cannot handle other information beyond the knowledge reserve of the model itself. The present invention fine-tunes the general LLM through a self-made command response data set, and with its semantic understanding and analysis capabilities, it can also achieve the effect of command fuzzy matching and joint command response. Whether it is "open whiteboard", "start whiteboard" or "help me open whiteboard", it can effectively trigger the "open whiteboard" command to improve user experience.

[0062] The present invention changes audio instructions into text instructions, and with the help of the semantic understanding and analysis capabilities of the instruction LLM, it is possible to achieve an instruction fuzzy matching effect without strictly relying on instruction trigger conditions.

[0063] Specifically, if the command response model cannot effectively identify the command content corresponding to the text information, a voice prompt is triggered. For example, a voice signal of "this command does not exist, please re-enter" is sent out through the speaker to instruct the user to re-enter. At the same time, the command response model determines that the voice signal corresponding to the text information is invalid information such as noise or an unfamiliar command, and continues to monitor the microphone voice input. On the other hand, the frequency of occurrence of unfamiliar commands can be counted to facilitate the subsequent addition and optimization of commands.

[0064] In some embodiments of the present invention, the text information is analyzed and processed using a large instruction response model to output instruction response information corresponding to a number of target instructions, wherein the large instruction response model is obtained by fine-tuning a pre-trained large language model according to an instruction response data set, and includes:

[0065] The command response data set includes a plurality of preset command response data pairs;

[0066] The plurality of preset instruction response data pairs include: text instruction information of a plurality of preset instructions and corresponding preset response information; the preset response information includes: the instruction ID marked by the preset instruction and the instruction reply information of the preset instruction;

[0067] The command response data set is a collection of text command information of several preset commands and their corresponding preset response information, which is summarized and sorted according to the functions of the voice command response device;

[0068] In the process of fine-tuning the pre-trained large language model, the instruction response large model learns the relationship between the text instruction information of each of the preset instructions and its corresponding preset response information;

[0069] The command response large model is used to understand and analyze the text information to match the corresponding target command response data pair, and the target command response data pair is analyzed based on the learned knowledge to obtain the target command and the corresponding command response information, and output a number of command response information corresponding to the target command; the target command response data pair is any one of a number of preset command response data pairs.

[0070] Specifically, the command response macromodel is used to understand and analyze the text information, and the matched target command response data pair is the one most similar to the text information among several preset command response data pairs.

[0071] In some embodiments of the present invention, fine-tuning the pre-trained large language model according to the instruction response dataset includes:

[0072] The text instruction information of the preset instruction in the instruction response data set is used as the prompt word of the pre-trained large language model, and the preset response information corresponding to the text instruction information of the preset instruction is used as the expected response parameter of the pre-trained large language model to fine-tune the pre-trained large language model.

[0073] Specifically, the present invention fine-tunes the general LLM through the command response data set, and the obtained command response large model (command LLM) has the ability to respond to all-in-one machine commands, fine-tunes the general model into a dedicated model, and returns output information in the expected specific format. After the wake-up command is triggered, the all-in-one machine uniformly converts the collected second voice signal into text information, and then uses the command LLM to analyze the text command information and command response information of the command (basic command, application command and central control command), wherein the command response information includes the command ID marked by the command and the command reply information of the command. The application interface of the linked all-in-one machine executes the corresponding command.

[0074] The categories of commands include basic commands for all-in-one machines, application commands, and central control commands. Basic commands refer to the operation of the tools that come with the all-in-one machine, such as hotspots and calendars; application commands refer to the operation of applications installed on the all-in-one machine, such as video conferencing and application market; central control commands refer to the operation of conference room equipment bound to the all-in-one machine, such as air conditioners and curtains.

[0075] The instructions and instruction IDs are a list of all-in-one machine functions that are manually collected, summarized and organized, and can be customized according to different all-in-one machine devices. There is a linkage relationship between the instructions in the all-in-one machine function list, such as turning on and off the all-in-one machine application function, raising and lowering the temperature and brightness, etc.

[0076] Preferably, the all-in-one machine uses SFT technology to fine-tune the general dialogue LLM to obtain the instruction LLM, including: the all-in-one machine collects and summarizes the function list of the all-in-one machine, including basic functions such as hotspots, calendars and file managers, application functions such as whiteboards and video conferencing, and central control functions such as temperature and curtains. The instruction LLM organizes it into instruction form, such as turning on / off the hotspot, and raising / lowering the temperature, etc., presets a corresponding ID for each instruction, and organizes it according to the format of the LLM fine-tuning data set to obtain a labeled instruction response data set, including input prompts and expected responses, such as: {"messages": [{"role":"user","content":"Open whiteboard."},{"role":"assistant","content":{"id":1001,"nb":"OK, the whiteboard has been opened for you."}}]}. Among them, the text instruction information of the instruction is "Open whiteboard", the ID of the instruction is 1001, and the instruction reply information of the instruction is "OK, the whiteboard has been opened for you". The all-in-one machine uses the labeled command response data set to fine-tune the general LLM to obtain the command LLM, thereby having local command response capabilities.

[0077] Specifically, the present invention fine-tunes the general LLM by presetting the command response data set, so that it learns the correspondence between text commands and command IDs. When the text content is sent to the fine-tuned command LLM, the command LLM will first match the corresponding command, and then output the command ID and command reply information according to the learned mapping relationship.

[0078] In one of the examples, the single command response process is as follows: after the voice command response module is successfully awakened, when the microphone collects the voice "open the whiteboard", the voice recognition module parses the text command information of the command as "open the whiteboard", and the all-in-one machine uses the command LLM analysis to obtain the target command A. The command ID corresponding to the target command A is "1001" and the command reply information is "OK, the whiteboard has been opened for you", thereby triggering the device to execute the corresponding whiteboard function.

[0079] In another example, the joint command response process is as follows: after the voice command response module is successfully awakened, when the microphone collects the voice "open the whiteboard and settings", the voice recognition module parses the text command information of the command as "open the whiteboard and settings", and uses the command LLM analysis to obtain the target command B and target command C. The command ID corresponding to the target command B is "1001" and the command reply information is "OK, the whiteboard has been opened for you"; the command ID corresponding to the target command C is "1002" and the command reply information is "OK, the settings have been opened for you", thereby triggering the device to execute the corresponding whiteboard functions and setting functions in sequence.

[0080] In some embodiments of the present invention, the instruction response macromodel outputs instruction response information corresponding to a plurality of target instructions in a preset format;

[0081] The command response information in the preset format is a json object.

[0082] Exemplarily, the command information in the preset format is a json object, including id and nb information, id corresponds to the command name, nb is the reply information generated by the large model for different commands, and the output reply information is, for example, "OK, the whiteboard has been opened for you" and "OK, the settings have been opened for you", etc., which are used to broadcast the command response information in voice form later. The command LLM can effectively deal with command and non-command text, not just a simple correspondence between text and command.

[0083] Step 140: triggering the voice command response device through the application interface to perform a corresponding response operation according to the command response information corresponding to each of the target commands.

[0084] In an optional manner, triggering the voice command response device through the application interface to perform a corresponding response operation according to the command response information corresponding to each of the target commands also includes:

[0085] Execute the corresponding instruction operation according to the instruction ID marked by the target instruction;

[0086] The command reply information of the target command is synchronously played in the form of voice through the voice synthesis module. Specifically, the application interface (Application Programming Interface, referred to as API) of the all-in-one machine is a set of definitions, programs and protocols, through which the interaction between different software components can be realized. The all-in-one machine application interface is the entrance to the execution instruction. Applications or other modules outside the all-in-one machine need to trigger the internal execution instruction through the application interface.

[0087] This application uses a variety of artificial intelligence algorithms to implement functions such as voice wake-up, voice recognition and voice synthesis, and combines these functions to build an all-in-one voice command response system to realize voice control of all-in-one devices and improve user experience.

[0088] The voice command response method based on large model fine-tuning provided by the present invention uses the command response data set to fine-tune the LLM to obtain a fine-tuned large language model, so that it has local command response capabilities, and at the same time, with the help of the text understanding and analysis capabilities of the large model, the effect of command fuzzy matching and joint command response is achieved. At the same time, for the numerous commands of the voice command response device based on large model fine-tuning, the user can only perform voice interaction based on the names of tools and applications in the interface display of the voice command response device based on large model fine-tuning, and can effectively trigger the function command without accurately remembering the trigger conditions, thereby reducing the difficulty of user use.

[0089] Embodiment 2, Figure 2 FIG. 1 shows a schematic diagram of a structure of an embodiment of a voice command response device based on large model fine-tuning provided by the present invention; Figure 2 As shown, the device 200 includes: a voice collection module 210, a wake-up module 220 and a response module 230.

[0090] The voice collection module 210 is used to collect a first voice signal, analyze the first voice signal, and obtain voice information;

[0091] The wake-up module 220 is used to determine whether the voice information matches a preset wake-up word; the preset wake-up word is pre-set by the user; if the voice information matches the preset wake-up word, the voice command response module is awakened;

[0092] The response module 230 is used to collect the second voice signal when waking up the voice command response module, and convert the second voice signal into text information through the voice recognition model;

[0093] The text information is analyzed and processed by using the command response big model, and command response information corresponding to a number of target commands is output; the command response information corresponding to the target command includes: the command ID marked by the target command and the command reply information of the target command; wherein the command response big model is obtained by fine-tuning the pre-trained big language model according to the command response data set; the command ID is pre-bound to the application interface of the voice command response device fine-tuned based on the big model;

[0094] The voice command response device is triggered through the application interface to perform a corresponding response operation according to the command response information corresponding to each target command.

[0095] In an optional manner, in the wake-up module 220, determining whether the voice information matches a preset wake-up word further includes:

[0096] If the judgment result matches the preset wake-up word, a voice response is triggered;

[0097] If the judgment result is that it does not match the preset wake-up word, wait silently.

[0098] In an optional manner, the wake-up module 220 is further configured to put the voice command response module into sleep mode when no third voice signal is received within a first preset time after the second voice signal is collected;

[0099] The voice command response module cannot be awakened again during the second preset sleep time.

[0100] In an optional manner, in the response module 230, the text information is analyzed and processed using a large instruction response model, and instruction response information corresponding to a plurality of target instructions is output, wherein the large instruction response model is obtained by fine-tuning a pre-trained large language model according to an instruction response data set, and includes:

[0101] The command response data set includes a plurality of preset command response data pairs;

[0102] The plurality of preset instruction response data pairs include: text instruction information of a plurality of preset instructions and corresponding preset response information; the preset response information includes: the instruction ID marked by the preset instruction and the instruction reply information of the preset instruction;

[0103] The command response data set is a collection of text command information of several preset commands and their corresponding preset response information, which is summarized and sorted according to the functions of the voice command response device;

[0104] In the process of fine-tuning the pre-trained large language model, the instruction response large model learns the relationship between the text instruction information of each of the preset instructions and its corresponding preset response information;

[0105] The command response large model is used to understand and analyze the text information to match the corresponding target command response data pair, and the target command response data pair is analyzed based on the learned knowledge to obtain the target command and the corresponding command response information, and output a number of command response information corresponding to the target command; the target command response data pair is any one of a number of preset command response data pairs.

[0106] In an optional manner, in the response module 230, fine-tuning the pre-trained large language model according to the instruction response data set includes:

[0107] The text instruction information of the preset instruction in the instruction response data set is used as the prompt word of the pre-trained large language model, and the preset response information corresponding to the text instruction information of the preset instruction is used as the expected response parameter of the pre-trained large language model to fine-tune the pre-trained large language model.

[0108] In an optional manner, the response module 230, the instruction response large model outputs instruction response information corresponding to a plurality of target instructions in a preset format;

[0109] The command response information in the preset format is a json object.

[0110] In an optional manner, the response module 230 triggers the voice command response device through the application interface to perform a corresponding response operation according to the command response information corresponding to each target command, and further includes:

[0111] Execute the corresponding instruction operation according to the instruction ID marked by the target instruction;

[0112] The command reply information of the target command is synchronously played in the form of voice through the voice synthesis module.

[0113] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0114] Embodiment 3, Figure 3 A structural schematic diagram of an embodiment of a voice command response device based on large model fine-tuning provided by the present invention is shown. The specific embodiment of the present invention does not limit the specific implementation of the voice command response device based on large model fine-tuning.

[0115] like Figure 3 As shown, the voice command response device based on large model fine-tuning may include: a processor (processor) 302, a communication interface (Communications Interface) 304, a memory (memory) 306, and a communication bus 308.

[0116] The processor 302, the communication interface 304, and the memory 306 communicate with each other via the communication bus 308. The communication interface 304 is used to communicate with other devices such as a client or other server network elements. The processor 302 is used to execute the program 310, which can specifically execute the relevant steps in the above-mentioned embodiment of the voice command response method for fine-tuning based on a large model.

[0117] Specifically, the program 310 may include program code including computer executable instructions.

[0118] The processor 302 may be a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. The one or more processors included in the voice command response device based on large model fine-tuning may be processors of the same type, such as one or more CPUs; or may be processors of different types, such as one or more CPUs and one or more ASICs.

[0119] The memory 306 is used to store the program 310. The memory 306 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0120] The algorithms or displays provided herein are not inherently related to any particular computer, virtual system or other device. In addition, the embodiments of the present invention are not directed to any particular programming language.

[0121] In the description provided herein, a large number of specific details are described. However, it is understood that embodiments of the present invention can be practiced without these specific details. Similarly, in order to simplify the present invention and help understand one or more of the various inventive aspects, in the above description of exemplary embodiments of the present invention, the various features of the embodiments of the present invention are sometimes grouped together into a single embodiment, figure, or description thereof. Wherein, the claims that follow the specific embodiment are hereby expressly incorporated into the specific embodiment, wherein each claim itself is a separate embodiment of the present invention.

[0122] Those skilled in the art will appreciate that the modules in the devices in the embodiments may be adaptively changed and arranged in one or more devices different from the embodiments. The modules or units or components in the embodiments may be combined into one module or unit or component, and further may be divided into a plurality of submodules or subunits or subcomponents. Except that at least some of such features and / or processes or units are mutually exclusive.

[0123] It should be noted that the above embodiments illustrate the present invention rather than limit it, and that those skilled in the art may devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference symbol between brackets shall not be construed as a limitation on the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "one" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention may be implemented by means of hardware comprising a number of different elements and by means of a suitably programmed computer. In a unit claim enumerating a number of devices, several of these devices may be embodied by the same hardware item. The use of the words first, second, and third, etc. does not indicate any order. These words may be interpreted as names. The steps in the above embodiments, unless otherwise specified, should not be understood as limitations on the order of execution.

Claims

1. A voice command response method based on large model fine-tuning, characterized in that: The method comprises: Collecting a first voice signal, analyzing the first voice signal, and obtaining voice information; Determine whether the voice information matches a preset wake-up word; the preset wake-up word is pre-set by the user; If the voice information matches the preset wake-up word, then wake up the voice command response module; When the voice command response module is awakened, a second voice signal is collected, and the second voice signal is converted into text information through a voice recognition model; The text information is analyzed and processed by using the command response big model, and command response information corresponding to a number of target commands is output; the command response information corresponding to the target command includes: the command ID marked by the target command and the command reply information of the target command; wherein the command response big model is obtained by fine-tuning the pre-trained big language model according to the command response data set; the command ID is pre-bound to the application interface of the voice command response device fine-tuned based on the big model; The voice command response device is triggered through the application interface to perform a corresponding response operation according to the command response information corresponding to each target command.

2. The voice command response method based on large model fine-tuning according to claim 1 is characterized in that: Determining whether the voice information matches a preset wake-up word also includes: If the judgment result matches the preset wake-up word, a voice response is triggered; If the judgment result is that it does not match the preset wake-up word, wait silently.

3. The voice command response method based on large model fine-tuning according to claim 1 is characterized in that: After waking up the voice command response module, it also includes: When no third voice signal is received within a first preset time after the second voice signal is collected, causing the voice command response module to sleep; The voice command response module cannot be awakened again during the second preset sleep time.

4. The voice command response method based on large model fine-tuning according to claim 1, characterized in that: The text information is analyzed and processed using a large instruction response model, and instruction response information corresponding to a number of target instructions is output, wherein the large instruction response model is obtained by fine-tuning a pre-trained large language model according to an instruction response data set, and includes: The command response data set includes a plurality of preset command response data pairs; The plurality of preset instruction response data pairs include: text instruction information of a plurality of preset instructions and corresponding preset response information; the preset response information includes: the instruction ID marked by the preset instruction and the instruction reply information of the preset instruction; The command response data set is a collection of text command information of several preset commands and their corresponding preset response information, which is summarized and sorted according to the functions of the voice command response device; In the process of fine-tuning the pre-trained large language model, the instruction response large model learns the relationship between the text instruction information of each of the preset instructions and its corresponding preset response information; The command response large model is used to understand and analyze the text information to match the corresponding target command response data pair, and the target command response data pair is analyzed based on the learned knowledge to obtain the target command and the corresponding command response information, and output a number of command response information corresponding to the target command; the target command response data pair is any one of a number of preset command response data pairs.

5. The voice command response method based on large model fine-tuning according to claim 4 is characterized in that: Fine-tune the pre-trained large language model based on the command response dataset, including: The text instruction information of the preset instruction in the instruction response data set is used as the prompt word of the pre-trained large language model, and the preset response information corresponding to the text instruction information of the preset instruction is used as the expected response parameter of the pre-trained large language model to fine-tune the pre-trained large language model.

6. The voice command response method based on large model fine-tuning according to any one of claims 1 to 5, characterized in that: The instruction response large model outputs instruction response information corresponding to target instructions in a plurality of preset formats; The command response information in the preset format is a json object.

7. The voice command response method based on large model fine-tuning according to claim 6 is characterized in that: The voice command response device is triggered by the application interface to perform a corresponding response operation according to the command response information corresponding to each target command, and further includes: Execute the corresponding instruction operation according to the instruction ID marked by the target instruction; The command reply information of the target command is synchronously played in the form of voice through the voice synthesis module.

8. A voice command response device based on large model fine-tuning, characterized in that: The device is used to perform the operation of the voice command response method based on large model fine-tuning as described in any one of claims 1 to 7, and the device includes: a voice acquisition module, a wake-up module and a response module; The voice collection module is used to collect the first voice signal, analyze the first voice signal, and obtain voice information; The wake-up module is used to determine whether the voice information matches a preset wake-up word; the preset wake-up word is pre-set by the user; if the voice information matches the preset wake-up word, the voice command response module is awakened; The response module is used to collect the second voice signal when waking up the voice command response module, and convert the second voice signal into text information through the voice recognition model; The text information is analyzed and processed by using the command response big model, and command response information corresponding to a number of target commands is output; the command response information corresponding to the target command includes: the command ID marked by the target command and the command reply information of the target command; wherein the command response big model is obtained by fine-tuning the pre-trained big language model according to the command response data set; the command ID is pre-bound to the application interface of the voice command response device fine-tuned based on the big model; The voice command response device is triggered through the application interface to perform a corresponding response operation according to the command response information corresponding to each target command.

9. A voice command response device based on large model fine-tuning, characterized in that: include: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform the operation of the voice command response method based on large model fine-tuning as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that: The storage medium stores at least one executable instruction. When the executable instruction is executed on a voice command response device / apparatus based on large model fine-tuning, the voice command response device / apparatus based on large model fine-tuning performs the operation of the voice command response method based on large model fine-tuning as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Acquisition method of false triggering voice information and device , equipment and storage medium

    CN112712799A

  • Information interaction method and device, electronic equipment and medium

    CN116894078A

  • Control method and device, electronic equipment and storage medium

    CN117174089A

  • Intelligent voice recognition method and system for AR helmet

    CN118379994A

  • Interaction method and device based on large language model, electronic equipment and medium

    CN119067078A

Cited By

  • False wakeup corpus acquisition method, related method, device, equipment and storage medium

    CN121306106A