Multi-model fusion decision-making method and device based on AI earphone, earphone and medium
By adopting a multi-model fusion decision-making method in smart headphones, a single recognition model is solved, and a more accurate voice processing and personalized services are achieved, improving user experience.
Patent Information
- Application Number
- CN202510343070.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-24
AI Technical Summary
In the voice control system of smart headphones, a single recognition model is difficult to cope with different voice requests from different types of users, resulting in poor user experience.
The multi-model fusion decision-making method based on AI headsets is adopted. By receiving the user's wake-up instructions, the voice information is obtained, the control type to which the voice information belongs, and the corresponding model is called according to the type to identify and make decisions, and the headset is finally controlled to perform related operations.
It realizes more accurate processing of user voice information, improves user experience, adapts to the needs of different users for different scenarios, and provides users with personalized and intelligent services.
Smart Images

Figure CN120201343A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of intelligent earphones, and in particular, to a multi-model fusion decision-making method, device, intelligent earphone and storage medium based on an AI earphone. Background Art
[0002] Intelligent earphones not only have the functions of traditional listening devices, but also have rich functions such as speech recognition, intelligent assistants, and multimedia playback, becoming an important tool for users to obtain information and manage personal affairs. Users can interact with intelligent earphones through voice commands to perform various operations such as playing music and answering calls.
[0003] Currently, in the voice control system of intelligent earphones, the processing of user voice commands usually relies on a single recognition model for decision-making. However, in actual applications, a single recognition model is difficult to handle different voice requests of different types of users, resulting in a poor experience for users during the use of intelligent earphones.
[0004] Therefore, how to make more accurate decisions on user voice information to improve the user experience has become a technical problem that needs to be solved urgently by those skilled in the art. It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0005] In view of the above, this application provides a multi-model fusion decision-making method, device, intelligent earphone and storage medium based on an AI earphone, aiming to solve the above technical problems.
[0006] In a first aspect, this application provides a multi-model fusion decision-making method based on an AI earphone. The method is applied to an intelligent earphone and includes:
[0007] Receiving a wake-up instruction initiated by a user and starting a voice control mode;
[0008] Obtaining the voice information sent by the user and determining the control type to which the voice information belongs;
[0009] According to the control type to which the voice information belongs, calling a corresponding model to recognize the voice information to obtain a decision result;
[0010] Controlling the intelligent earphone to perform relevant operations according to the decision result.
[0011] In a second aspect, this application provides a multi-model fusion decision-making device based on an AI earphone. The device includes:
[0012] A start-up module: for receiving a wake-up instruction initiated by a user and starting a voice control mode;
[0013] Determination module: configured to obtain the voice information sent by the user and determine the control type to which the voice information belongs;
[0014] Decision-making module: configured to call a corresponding model to identify the voice information according to the control type to which the voice information belongs and obtain a decision result;
[0015] Control module: configured to control the intelligent earphone to perform related operations according to the decision result.
[0016] In a third aspect, the present application provides an intelligent earphone, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0017] The memory is used to store a computer program;
[0018] The processor is configured to implement the multi-model fusion decision-making method based on an AI earphone according to any one of the embodiments in the first aspect when executing the program stored on the memory.
[0019] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it implements the multi-model fusion decision-making method based on an AI earphone according to any one of the embodiments in the first aspect.
[0020] The above technical solutions provided by the embodiments of the present application have the following advantages compared with the prior art:
[0021] The present application receives a wake-up instruction initiated by a user, starts a voice control mode, obtains the voice information sent by the user, determines the control type to which the voice information belongs, calls a corresponding model to identify the voice information according to the control type to which the voice information belongs and obtains a decision result, and controls the intelligent earphone to perform related operations according to the decision result. It can flexibly call a corresponding model to identify the voice information according to the control type to which the voice information belongs, and process the user's voice information through models with different computing resources, so that the earphone can quickly make a feedback when processing the user's voice information instruction, adapt to the needs of different users for different scenarios, provide personalized and intelligent services for users, and improve the user experience. Description of the Drawings
[0022] The drawings here are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0023] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0024] Figure 1 It is a schematic flowchart of an embodiment of the multi-model fusion decision-making method based on an AI headset of the present application;
[0025] Figure 2 It is a schematic module diagram of an embodiment of the multi-model fusion decision-making device based on an AI headset of the present application;
[0026] Figure 3 It is a schematic diagram of an embodiment of the intelligent headset of the present application;
[0027] The realization of the purpose of the present application, functional features and advantages will be further described in conjunction with the embodiments with reference to the drawings. Detailed implementation manners
[0028] To make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.
[0029] The following disclosure provides many different embodiments or examples for implementing different structures of the present application. To simplify the disclosure of the present application, the components and settings of specific examples are described below. Of course, they are only examples and are not intended to limit the present application. In addition, the present application may repeat reference numerals and / or letters in different examples. This repetition is for the purpose of simplification and clarity, and does not itself indicate the relationship between the various embodiments and / or settings discussed.
[0030] The present application provides a multi-model fusion decision-making method based on an AI headset. Referring to Figure 1 As shown, it is a schematic flowchart of the method of an embodiment of the multi-model fusion decision-making method based on an AI headset of the present application. This method can be executed by an intelligent headset, and the intelligent headset can be implemented by software and / or hardware. The multi-model fusion decision-making method based on an AI headset includes:
[0031] Step S10: Receive a wake-up instruction initiated by the user and start the voice control mode;
[0032] Step S20: Obtain the voice information sent by the user and determine the control type to which the voice information belongs;
[0033] Step S30: According to the control type to which the voice information belongs, call the corresponding model to recognize the voice information to obtain a decision result;
[0034] Step S40: Control the intelligent earphone to perform related operations according to the decision result.
[0035] The intelligent earphone of this application can be connected to the AI large model. Users can control the related operations of the earphone through voice. The earphone cabin of the intelligent earphone has a 4G card insertion function and can be connected to the exclusive audio content APP. Therefore, when users use the intelligent earphone, they can obtain the audio content they want to listen to anytime and anywhere without relying on devices such as mobile phones.
[0036] In the scenario of voice interaction, to facilitate the interaction between the user and the intelligent earphone, the user can wake up the voice control mode of the intelligent earphone. The user can send a voice wake-up command to the earphone to activate the voice control function of the intelligent earphone. For example, the voice wake-up command "Kitty Kitty" is preset. The wake-up command module of the earphone recognizes this command and starts the voice control mode, ready to receive subsequent voice inputs.
[0037] After the voice control mode of the intelligent earphone is started, the earphone continuously monitors and obtains the voice information command of the user to ensure that it can respond to the user's needs in a timely manner. For example, the user then says "Turn up the volume a little bit". At this time, the earphone records the user's voice information through the voice acquisition module and can convert it into a digital signal.
[0038] The control types to which the voice information belongs can be divided into basic control voices and non-basic control voices. Basic control voices refer to basic control types and belong to direct request operations. For example, if the earphone recognizes that the voice information is "Play the next song", it can be determined that the voice information is a basic control voice. Non-basic control voices refer to request operations that require in-depth understanding.
[0039] After determining the voice information type, the system calls the corresponding recognition model according to the type to process the voice information. Specifically, according to the control type to which the voice information belongs, calling the corresponding model to recognize the voice information to obtain a decision result includes:
[0040] If the voice information is a basic control voice, recognize the voice information through the local recognition model to obtain a first decision result;
[0041] If the voice information is a non-basic control voice, call the model in the cloud to recognize the non-basic control voice to obtain a second decision result.
[0042] For basic control voice, a lightweight model deployed locally on the earphone can be used. For non-basic control voice, a cloud model can be used. For example, assuming the voice information sent by the user is "increase the volume" or "play the next song", the local lightweight voice recognition model is called for quick processing to obtain a decision result (denoted as the first decision result), that is, "volume increased" or "next song".
[0043] Assume the voice information sent by the user is "recommend a song suitable for running", and the system sends this request to the large model in the cloud for analysis, so as to obtain an accurate decision result (denoted as the second decision result).
[0044] Control the intelligent earphone to perform relevant operations according to the decision result to respond to the instruction corresponding to the user's voice information. For example, for the "volume increased" instruction returned by the model, the earphone immediately performs the operation of adjusting the volume. For the complex request "recommend a song suitable for running" returned by the cloud model, the intelligent earphone can reply to the user: "The song recommended for you is 'Running'. Do you want to play it immediately?" and then execute after waiting for the user to confirm.
[0045] This application flexibly calls the corresponding model according to the control type to which the voice information belongs to identify the voice information and obtain the decision result. By using models with different computing resources to process the user's voice information, the earphone can quickly make a feedback when processing the instructions of the user's voice information, adapt to the needs of different users in different scenarios, and provide personalized and intelligent services for users.
[0046] In one embodiment, determining the control type to which the voice information belongs includes:
[0047] Convert the voice information into text information;
[0048] Match the text information with a preset text set to determine whether the target intention of the voice information is matched in the preset text set;
[0049] If so, determine that the voice information is basic control voice;
[0050] If not, determine that the voice information is non-basic control voice.
[0051] The intelligent earphone utilizes a built-in voice conversion module to convert the user's voice information into text information. After obtaining the converted text information, the earphone matches it with a preset text set, which contains texts corresponding to common basic control voices, such as "increase volume", "stop playing", "next song", etc. This text set can be stored in the memory of the earphone and can be called at any time for matching. If the target intention of the voice information is matched in the preset text set, the voice information will be recognized as a basic control, that is, the instruction issued by the user belongs to a simple control operation. For example, if the text information is successfully matched to "play", the system determines that the voice information is a basic control voice. On the contrary, if no corresponding target intention is matched, the voice information will be determined as a non-basic control voice. For example, the user's voice information is "recommend a song suitable for running for me", and the text information corresponding to this voice information does not match the target intention in the preset text set, then this voice information is determined as a non-basic control voice. Thus, it is possible to quickly distinguish between basic control and non-basic control voices. For the voice commands of basic control, they can be processed immediately, reducing the user's waiting time, while non-basic control is parsed by a more complex model to ensure the accuracy of the commands.
[0052] In one embodiment, the controlling the intelligent earphone to execute relevant operations according to the decision result includes:
[0053] Performing an audio control operation on the intelligent earphone through the first decision result;
[0054] Feeding back conversation information to the user through the second decision result.
[0055] Since the user hopes that the earphone can immediately execute audio-related operations through voice commands, such as playing, pausing, adjusting the volume, etc. In addition, the user may also need the earphone to feed back relevant conversation information. To effectively handle these requirements, the intelligent earphone needs to take corresponding operations according to different types of decision results. Specifically, after the voice information issued by the user is accurately recognized and converted into the first decision result, the earphone will perform corresponding audio control operations according to this result, including but not limited to operations such as playing, pausing, and adjusting the volume. For example, if the user's voice information is "increase the volume", the first decision result is "increase the volume", and the earphone will increase the volume.
[0056] In addition to audio control operations, the smart headphones can also respond to the voice information requested by the user. That is, based on the second decision result, the headphones will generate corresponding conversation information and feedback it to the user. The feedback content includes confirmation information, suggestions, or providing additional information, etc. For example, if the user's voice information is "What's the weather like today", the second decision result will obtain the latest weather information through the cloud model, and the headphones will give a voice feedback like "Today's weather is cloudy, and the temperature is between 15 and 20 degrees". The conversation between the user and the headphones can enhance the smart experience of the headphones. Through context-based information feedback, users can more conveniently obtain the information they need, reducing the time for searching and confirmation.
[0057] Among them, the audio control operations include:
[0058] Performing a control operation on the playing state of the currently playing audio, performing a control operation on the volume size of the currently playing audio, or performing a switching operation on the currently playing audio.
[0059] The headphones can immediately change the playing state of the audio according to the user's voice information. The playing states of the audio mainly include playing, pausing, stopping, etc. It can also adjust the volume size according to the user's voice information, including increasing, decreasing, or setting to a specific volume. For example, if the user requests "Lower the volume to 50%", the headphones will adjust the volume to 50%. It can also perform a switching operation on the currently playing audio according to the user's voice information. The switching operation includes switching to the next or previous song in the current playlist, or selecting a specific song to play.
[0060] In one embodiment, the calling of the cloud model to identify the non-basic control voice to obtain the second decision result includes:
[0061] Calling the emotion feature extraction model in the cloud to extract the emotion features of the non-basic control voice;
[0062] Converting the non-basic control voice into text information, and calling the text feature extraction model in the cloud to extract the text features of the text information;
[0063] Based on the emotion features and the text features, calling the semantic recognition model in the cloud to identify the non-basic control voice to obtain the second decision result.
[0064] When the user issues a non-basic control voice command (e.g., expressing emotions or making requests) in the smart earphone, the earphone converts the captured voice segment into a predetermined encoding format (e.g., PCM or WAV format) and transmits it to the cloud server. The cloud inputs the received audio data into an emotion feature extraction model, which can analyze the acoustic wave features using a convolutional neural network, extract data reflecting the user's emotions, and output an emotion feature vector covering emotion states such as "happy", "sad", "angry", "neutral", etc.
[0065] Convert the non-basic control voice into text information, determine the keywords in the text, such as important information like "need help", "recommend music", and "today's weather", use natural language processing (NLP) technology to analyze the structure and meaning of the sentence, so as to extract the text features corresponding to the text information. The semantic recognition model in the cloud combines the received emotion feature vector with the text features for comprehensive analysis to identify the user's specific intention. For example, if the user's emotion feature is "tired" and the text features show that the user requests "play light music", the second decision result recognized by the model may be "play light music to relax the mood". If the user expresses anxiety and requests "share some encouraging words", the model will accordingly generate "provide comfort and positive information". By combining emotions with text features, the semantic recognition model can understand the user's intention more meticulously to improve the accuracy of the decision result.
[0066] Furthermore, based on the emotion features and the text features, calling the semantic recognition model in the cloud to identify the non-basic control voice to obtain a second decision result includes:
[0067] Concatenate the emotion features and the text features to obtain a concatenated feature;
[0068] Input the concatenated feature into the semantic recognition model to obtain the semantic recognition result of the non-basic control voice;
[0069] Determine the second decision result based on the semantic recognition result.
[0070] Fuse and splice the vector of the emotional feature with the vector of the corresponding text feature to form a new spliced feature vector. For example, the emotional feature is a 1D array [0.8, 0.2] (indicating happiness and surprise), and the text feature is a 1D array [1, 0, 1] (indicating keywords and grammar structures). The spliced feature will form a multi-dimensional array [0.8, 0.2, 1, 0, 1]. Input the spliced feature into the semantic recognition model to obtain the semantic recognition result of the non-basic control speech. The semantic recognition model can be trained based on the long short-term memory network. The semantic recognition model is trained with a large-scale dataset and has the ability to recognize emotional tones and text semantics, and can understand the complex information contained in the spliced feature. The result output by the model reflects the specific intention expressed by the non-basic control speech, such as the intention classification result, sentiment analysis result, and related tags. Suppose the spliced feature implies "interested in relaxing activities", the recognition result that the model may return is "request for relaxing music" or "need suggestions for relaxation". For example, if the semantic recognition result points to "request for relaxing music", the intelligent earphone will formulate relevant feedback instructions, including playing a piece of light music. The intelligent earphone can also return personalized suggestions according to the recognition result, such as "listening to some relaxing podcasts may help you relax", so as to enhance the user's interaction experience and operation efficiency.
[0071] Refer to Figure 2 As shown, it is a schematic diagram of the functional modules of the multi-model fusion decision-making device 100 based on an AI earphone according to the present application.
[0072] The multi-model fusion decision-making device 100 based on an AI earphone described in the present application can be installed in an intelligent earphone. According to the implemented functions, the multi-model fusion decision-making device 100 based on an AI earphone can include a startup module 110, a determination module 120, a decision module 130, and a control module 140. The modules described in the present application can also be referred to as units, which refer to a series of computer program segments that can be executed by the intelligent earphone processor and can complete fixed functions, and are stored in the memory of the intelligent earphone.
[0073] In this embodiment, the functions of each module / unit are as follows:
[0074] Startup module 110: used to receive the wake-up instruction initiated by the user and start the voice control mode;
[0075] Determination module 120: used to obtain the voice information sent by the user and determine the control type to which the voice information belongs;
[0076] Decision module 130: used to call the corresponding model according to the control type to which the voice information belongs to recognize the voice information and obtain a decision result;
[0077] Control module 140: configured to control the intelligent earphone to perform relevant operations according to the decision result.
[0078] In one embodiment, determining the control type to which the voice information belongs includes:
[0079] Converting the voice information into text information;
[0080] Matching the text information with a preset text set to determine whether the target intention of the voice information is matched in the preset text set;
[0081] If so, determining that the voice information is basic control voice;
[0082] If not, determining that the voice information is non-basic control voice.
[0083] In one embodiment, according to the control type to which the voice information belongs, calling a corresponding model to identify the voice information to obtain a decision result includes:
[0084] If the voice information is basic control voice, identifying the voice information through a local recognition model to obtain a first decision result;
[0085] If the voice information is non-basic control voice, calling a cloud model to identify the non-basic control voice to obtain a second decision result.
[0086] In one embodiment, according to the decision result, controlling the intelligent earphone to perform relevant operations includes:
[0087] Performing an audio control operation on the intelligent earphone through the first decision result;
[0088] Feeding back conversation information to the user through the second decision result.
[0089] In one embodiment, the audio control operation includes:
[0090] Performing a control operation on the playing state of the currently playing audio, performing a control operation on the volume of the currently playing audio, or performing a switching operation on the currently playing audio.
[0091] In one embodiment, calling a cloud model to identify the non-basic control voice to obtain a second decision result includes:
[0092] Calling a cloud emotion feature extraction model to extract the emotion feature of the non-basic control voice;
[0093] Converting the non-basic control voice into text information, and calling a cloud text feature extraction model to extract the text feature of the text information;
[0094] Based on the emotional feature and the text feature, call the semantic recognition model in the cloud to recognize the non-basic control speech to obtain a second decision result.
[0095] In one embodiment, the step of based on the emotional feature and the text feature, calling the semantic recognition model in the cloud to recognize the non-basic control speech to obtain a second decision result includes:
[0096] Concatenate the emotional feature and the text feature to obtain a concatenated feature;
[0097] Input the concatenated feature into the semantic recognition model to obtain the semantic recognition result of the non-basic control speech;
[0098] Determine the second decision result based on the semantic recognition result.
[0099] Refer to Figure 3 shown, which is a schematic diagram of a preferred embodiment of the intelligent earphone of the present application.
[0100] The intelligent earphone includes a processor 111, a communication interface 112, a memory 113, and a communication bus 114. Among them, the processor 111, the communication interface 112, and the memory 113 complete mutual communication through the communication bus 114;
[0101] The memory 113 is used to store computer programs, for example, a multi-model fusion decision program based on an AI earphone;
[0102] Among them, the processor 111 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips in some embodiments. The processor 111 is generally used to control the overall operation of the intelligent earphone, such as performing control and processing related to data interaction or communication. In this embodiment, the processor 111 is used to run the program code stored in the memory 113 or process data, such as running the program code of the multi-model fusion decision program based on the AI earphone.
[0103] The communication interface 112 may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface), and the communication interface 112 can also be used to establish a communication connection between the intelligent earphone and other intelligent earphones.
[0104] The memory 113 includes at least one type of readable storage medium, which includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 113 may be an internal storage unit of the smart headset, such as the hard disk or memory of the smart headset. In other embodiments, the memory 113 may also be an external storage device of the smart headset, such as a plug-in hard disk equipped with the smart headset, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Of course, the memory 113 may also include both the internal storage unit of the smart headset and its external storage device. In this embodiment, the memory 11 is generally used to store the operating system installed in the smart headset and various computer programs, such as the program code of the multi-model fusion decision-making program based on the AI headset. In addition, the memory 113 may also be used to temporarily store various types of data that have been output or will be output.
[0105] Figure 3 Only the smart headset with components 111-114 is shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.
[0106] In an embodiment of the present application, when the processor 111 is used to execute the program stored on the memory 113, it implements the multi-model fusion decision-making method based on the AI headset provided by any one of the foregoing method embodiments, including:
[0107] Receiving a wake-up instruction initiated by the user and starting the voice control mode;
[0108] Obtaining the voice information sent by the user and determining the control type to which the voice information belongs;
[0109] According to the control type to which the voice information belongs, calling the corresponding model to identify the voice information to obtain a decision result;
[0110] Controlling the smart headset to perform related operations according to the decision result.
[0111] For a detailed introduction to the above steps, please refer to the above Figure 1 Explanation of the flowchart of the embodiment of the multi-model fusion decision-making method based on the AI headset.
[0112] In addition, an embodiment of the present application further provides a computer-readable storage medium, which may be non-volatile or volatile. The computer-readable storage medium includes a storage data area and a storage program area. The storage program area stores a multi-model fusion decision program based on an AI headset. When the multi-model fusion decision program based on the AI headset is executed by a processor, the following operations are implemented:
[0113] Receive a wake-up instruction initiated by the user and start the voice control mode;
[0114] Obtain the voice information sent by the user and determine the control type to which the voice information belongs;
[0115] According to the control type to which the voice information belongs, call the corresponding model to identify the voice information to obtain a decision result;
[0116] Control the intelligent headset to perform related operations according to the decision result.
[0117] The specific implementation manner of the computer-readable storage medium of the present application is substantially the same as the specific implementation manner of the above multi-model fusion decision method based on an AI headset, and will not be described in detail here.
[0118] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0119] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the related technology, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0120] It should be noted that the descriptions involving "first", "second", etc. in this application are only for descriptive purposes and should not be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. Additionally, the technical solutions between various embodiments may be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by this application.
[0121] It should be understood that the terms used herein are for the purpose of describing particular example embodiments only and are not intended to be limiting. Unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" as used herein may also include the plural forms. The terms "comprising", "including", "containing", and "having" are inclusive and thus specify the presence of the stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring them to be performed in the particular order described or illustrated, unless the order of performance is explicitly stated. It should also be understood that additional or alternative steps may be used.
[0122] The above are only specific embodiments of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application will not be limited to these embodiments shown herein, but rather will conform to the broadest scope consistent with the principles and novel features claimed herein.
Claims
1. A multi-model fusion decision method based on AI headset, characterized in that: The method is applied to a smart headset, and the method comprises: Receive a wake-up command initiated by the user and start the voice control mode; Acquire the voice information sent by the user, and determine the control type to which the voice information belongs; According to the control type to which the voice information belongs, calling the corresponding model to identify the voice information to obtain a decision result; The smart headset is controlled to perform relevant operations according to the decision result.
2. The multi-model fusion decision-making method based on AI headset according to claim 1, characterized in that: The determining the control type to which the voice information belongs includes: Converting the voice information into text information; Matching the text information with a preset text set to determine whether the target intent of the voice information is matched in the preset text set; If so, determining that the voice information is the basic control voice; If not, it is determined that the voice information is non-basic control voice.
3. The multi-model fusion decision method based on AI headset according to claim 1, characterized in that: The step of calling a corresponding model to identify the voice information and obtaining a decision result according to the control type to which the voice information belongs includes: If the voice information is the basic control voice, the voice information is recognized by a local recognition model to obtain a first decision result; If the voice information is non-basic control voice, a cloud-based model is called to identify the non-basic control voice to obtain a second decision result.
4. The multi-model fusion decision method based on AI headset according to claim 1, characterized in that: The controlling the smart headset to perform related operations according to the decision result includes: Performing an audio control operation on the smart headset according to the first decision result; The dialogue information is fed back to the user through the second decision result.
5. The multi-model fusion decision method based on AI headset according to claim 4, characterized in that: The audio control operation includes: Control the playback status of the currently playing audio, control the volume of the currently playing audio, or switch the currently playing audio.
6. The multi-model fusion decision method based on AI headset according to claim 3, characterized in that: The calling of the cloud model to identify the non-basic control voice to obtain a second decision result includes: Calling the emotion feature extraction model in the cloud to extract the emotion feature of the non-basic control speech; Convert the non-basic control voice into text information, and call a text feature extraction model in the cloud to extract text features of the text information; Based on the emotion feature and the text feature, a cloud-based semantic recognition model is called to recognize the non-basic control voice to obtain a second decision result.
7. The multi-model fusion decision method based on AI headset according to claim 6, characterized in that: The calling of a cloud-based semantic recognition model to identify the non-basic control voice based on the emotion feature and the text feature to obtain a second decision result includes: Concatenating the emotion feature and the text feature to obtain a concatenated feature; Inputting the concatenated features into a semantic recognition model to obtain a semantic recognition result of the non-basic control speech; The second decision result is determined based on the semantic recognition result.
8. A multi-model fusion decision-making device based on AI headphones, characterized in that: The device comprises: Startup module: used to receive the wake-up command initiated by the user and start the voice control mode; Determination module: used for acquiring the voice information sent by the user and determining the control type to which the voice information belongs; Decision-making module: used to call the corresponding model to identify the voice information and obtain a decision result according to the control type to which the voice information belongs; Control module: used to control the smart headset to perform relevant operations according to the decision result.
9. A smart headset, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, used to implement the multi-model fusion decision method based on AI headset according to any one of claims 1 to 7 when executing the program stored in the memory.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the multi-model fusion decision-making method based on AI headphones as described in any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Bluetooth earphone multi-device switching and scene perception control system based on voice recognition
CN122201293A