Smart loudspeaker box processing method and apparatus, electronic device, and storage medium
Through the target dialogue model, the voice content of the smart speaker is analyzed and the expected execution action information is generated, which solves the accuracy problem of the smart speaker when recognizing user voice commands, and achieves smarter and more efficient functional execution.
Patent Information
- Application Number
- PCT/CN2024/127887
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-10-28
- Publication Date
- 2025-07-03
AI Technical Summary
When the smart speaker recognizes the user's voice commands, it may not be able to accurately recognize it, resulting in the wrong selection of the function to execute.
The voice content is analyzed through the target dialogue model, the expected execution action information is generated, and the corresponding execution operation function is determined based on the information, and the smart speaker is controlled.
It improves the accuracy of the recognition of voice commands by smart speakers, enhances functional scalability and interaction efficiency, and provides more intelligent and personalized services.
Smart Images

Figure CN2024127887_03072025_PF_FP_ABST
Abstract
Description
Smart speaker processing method, device, electronic device and storage medium
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on December 29, 2023, with application number 202311865315.8 and invention name “Smart speaker processing method, device, electronic device and storage medium”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The embodiments of the present disclosure relate to the field of smart speaker technology, and in particular to a smart speaker processing method, device, electronic device, and storage medium. Background Art
[0003] With the continuous development of intelligent technology, smart speakers are becoming increasingly popular. Most smart speakers have voice interaction functions. After receiving voice, smart speakers respond based on the content of the voice to provide voice answers or smart device control functions. However, when implementing different functions such as voice answers or smart device control, smart speakers may not accurately recognize user voice commands, resulting in the selection of the wrong function.
[0004] Summary of the Invention
[0005] The present disclosure provides a smart speaker processing method, device, electronic device and storage medium to achieve the operation function of accurately recognizing voice instructions only for controlling the smart speaker.
[0006] In a first aspect, an embodiment of the present disclosure provides a smart speaker processing method, the method comprising:
[0007] Determine the current text content corresponding to the current voice content input to the smart speaker;
[0008] [Corrected 14.05.2025 according to Rule 91] Parsing the current text content using a target dialogue model to obtain information about the desired action corresponding to the current speech content, wherein the target dialogue model is pre-trained based on a natural language model;
[0009] Determining a current execution operation function used by the expected execution action information corresponding to the current voice content;
[0010] The smart speaker is controlled according to the currently executed operation function.
[0011] In a second aspect, an embodiment of the present disclosure further provides a smart speaker processing device, the device comprising:
[0012] A determination module, configured to determine the current text content corresponding to the current voice content input to the smart speaker;
[0013] [Corrected 14.05.2025 according to Rule 91] A parsing module, configured to parse the current text content using a target dialogue model to obtain information about the desired action corresponding to the current speech content, wherein the target dialogue model is pre-trained based on a natural language model;
[0014] A matching module, configured to determine a current execution operation function used by the expected execution action information corresponding to the current voice content;
[0015] A control module is used to control the smart speaker according to the currently executed operation function.
[0016] In a third aspect, an embodiment of the present disclosure further provides an electronic device, the electronic device comprising:
[0017] at least one processor; and
[0018] a memory communicatively connected to the at least one processor; wherein,
[0019] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the smart speaker processing method described in any one of the above embodiments.
[0020] In a fourth aspect, a computer-readable medium is further provided in an embodiment of the present disclosure, wherein the computer-readable medium stores computer instructions, and the computer instructions are used to enable a processor to implement the smart speaker processing method described in any one of the above embodiments when executed.
[0021] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.
[0023] FIG1 is a flow chart of a smart speaker processing method provided by an embodiment of the present disclosure;
[0024] FIG2 is a schematic diagram of a framework of a smart speaker processing provided by an embodiment of the present disclosure;
[0025] FIG3 is a schematic structural diagram of a smart speaker processing device provided by an embodiment of the present disclosure;
[0026] FIG4 is a schematic structural diagram of an electronic device for implementing a smart speaker processing method provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0027] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0028] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0029] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.
[0030] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0031] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0032] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0033] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0034] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0035] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0036] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0037] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0038] Figure 1 is a flow chart of a smart speaker processing method provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to situations where a smart speaker can accurately identify the need to perform an action and thereby adapt to control the corresponding operation function. The method can be executed by a smart speaker processing device, which can be implemented in the form of software and / or hardware and is generally integrated into any electronic device with network communication function, which can be a mobile terminal, PC or server, etc.
[0039] As shown in FIG1 , the smart speaker processing method according to an embodiment of the present disclosure may include the following steps:
[0040] S110. Determine the current text content corresponding to the current voice content input to the smart speaker.
[0041] As shown in Figure 2, the smart speaker receives the current voice content from the external input in real time and converts the current voice content into the corresponding current text content through the built-in voice recognition function of the smart speaker. For example, the current voice content can be pre-processed, including but not limited to noise reduction, reverberation, and enhancement, to improve the quality and recognizability of the current voice content. Then, the current voice content is processed by voice-to-text to obtain the corresponding current text content.
[0042] As an optional but non-limiting implementation, determining the current text content corresponding to the current voice content input to the smart speaker includes steps A1-A2:
[0043] Step A1: Obtain the current voice content collected by the microphone configured on the smart speaker.
[0044] Step A2: Perform speech-to-text recognition on the current speech content to obtain the current text content corresponding to the current speech content.
[0045] Referring to FIG2 , a microphone (e.g., a microphone) is built into the smart speaker, and the microphone or other microphones can capture the voice signal emitted from the outside of the smart speaker. Optionally, considering that some errors may occur in the process of voice-to-text conversion, for example, some words in the voice may be mistakenly recognized as other words, a language model can be used to predict the possible text content, and context information can be used to determine the correct text content, and the result of voice-to-text conversion can be corrected to ensure its accuracy. For the current text content generated by voice-to-text conversion, the recognized current text content is pre-processed by removing punctuation, converting to lowercase, and removing stop words to improve the quality and comprehensibility of the text content.
[0046] [Corrected 14.05.2025 according to Rule 91] S120. Analyze the current text content through the target dialogue model to obtain the expected execution action information corresponding to the current voice content. The target dialogue model is pre-trained based on the natural language model.
[0047] The target dialogue model is a natural language model based on deep learning. For example, a natural language model based on the Transformer structure is trained using a large-scale corpus to enable the target dialogue model to learn the statistical laws and semantic representation of the language. The target dialogue model can be used for various natural language processing tasks, such as text generation, question answering, and machine translation.
[0048] Referring to Figure 2, the current text content corresponding to the current voice content is input into the target dialogue model. The target dialogue model can automatically parse the expected execution action requirements for the smart speaker included in the current text content, and then generate expected execution action information that matches the current voice content. The expected execution action information corresponding to the current voice content includes the specific action that the smart speaker is expected to perform. The expected execution action information may include device control instructions, emotional dialogue instructions, etc. If the expected execution action information is a device control instruction, the corresponding device control operation can be performed through the voice control function of the smart speaker; if the expected execution action information is an emotional dialogue instruction, the corresponding emotional dialogue can be carried out through the voice interaction function of the smart speaker.
[0049] As an optional but non-limiting implementation, the target dialogue model is used to parse the current text content to obtain the expected action information corresponding to the current voice content, including steps B1-B2:
[0050] Step B1: When it is detected that a target dialogue model needs to be used for voice dialogue interaction, the reference dialogue guidance information corresponding to the target dialogue model is determined through the target dialogue model. The reference dialogue guidance information includes the previous text content input to the target dialogue model and the expected execution action information corresponding to the previous voice content obtained by recognizing the previous text content through the target dialogue model. The previous text content is the text content corresponding to the previous voice content input to the smart speaker.
[0051] Step B2: parse the reference dialogue guidance information and the current text content through the target dialogue model to obtain the first expected execution action information corresponding to the current voice content. The first expected execution action information is used to instruct the smart speaker to perform a first task operation. The first task operation is to feedback matching voice content for the current voice content.
[0052] As shown in Figure 2, device control of smart speakers includes voice interaction, home control, and voice broadcast (including but not limited to weather forecasts, music playback, etc.) control functions. During voice interaction, the dialogue must be understood and analyzed between both parties before the dialogue results can be given. Without contextual information, it may not be able to accurately understand the needs of the voice interaction and thus provide an accurate response. For example, if you say "I want to travel to place A" to the smart speaker, the target dialogue model may respond with some travel information about place A. However, without more contextual information, it may not accurately understand the dialogue needs and may not know whether you want to learn about tourist attractions in place A or about the food and accommodations in place A.
[0053] Refer to Figure 2. To this end, when the target dialogue model is used to parse the current text content in the voice dialogue interaction scenario, not only the current text content corresponding to the current voice content is required, but also the dialogue requirements of the voice dialogue interaction need to be understood through continuous context, the previous voice content input to the smart speaker is obtained, and speech-to-text is converted to the corresponding previous text content. The target dialogue model is used to parse and identify the previous text content to determine the expected execution action information corresponding to the previous voice content; then, the previous text content and the expected execution action information corresponding to the previous voice content are combined into reference dialogue guidance information, and the reference dialogue guidance information is provided as input to the target dialogue model.
[0054] As shown in Figure 2, based on the reference dialogue guidance information and the current text content, the target dialogue model parses the current text content to obtain the first desired action information corresponding to the current speech content. The target dialogue model then uses the reference dialogue guidance information and the context of the current text content to generate the first desired action information corresponding to the current speech content. This demonstrates that understanding the conversational needs of voice dialogue interaction through continuous context creates a more natural and fluid conversational experience. This better understanding of conversational needs leads to more accurate and useful responses.
[0055] The above process determines the reference dialogue guidance information based on the target dialogue model. The context dimension is reflected by including the previous text content and the desired action information corresponding to the previous speech content. This allows the target dialogue model to refer to the previous dialogue content and the desired action when generating responses, improving the coherence and accuracy of the dialogue.
[0056] As another optional but non-limiting implementation, parsing the current text content through the target dialogue model to obtain the expected execution action information corresponding to the current voice content includes the following steps:
[0057] When it is detected that a target dialogue model is needed to control the device, the target dialogue model is used to obtain second expected execution action information corresponding to the current voice content. The second expected execution action information is used to instruct the smart speaker to perform a second task operation. The second task operation is to generate a matching device control instruction for the current voice content.
[0058] When controlling a device, it is usually not necessary for the two parties to conduct a dialogue, understand, and analyze each other before a more appropriate result can be given. Therefore, the target dialogue model can be used to directly obtain the second expected execution action information corresponding to the current voice content for the current text content. Among them, device control includes home control, voice broadcast (including but not limited to weather forecast, music playback, etc.) and other control functions. Among them, the weather forecast supports obtaining 7-day weather forecast information for cities across the country, supports locating the current IP address to the city, and supports extended questions, such as: Is place A suitable for hiking this Saturday? Tomorrow's clothing matching suggestions, etc.; music playback supports music playback, specifying songs, and recommending songs based on mood, time period, and atmosphere.
[0059] As an optional but non-limiting implementation, before parsing the current text content through the target dialogue model to obtain the expected execution action information corresponding to the current voice content, the following steps C1-C3 are also included:
[0060] Step C1: Detect whether the current text content contains reference text description content through the target dialogue model. The reference text description content is used to indicate whether a smart speaker is used for device control.
[0061] Step C2: If the reference text description content is not detected, the target dialogue model is used for voice dialogue interaction.
[0062] Step C3: If the reference text description content is detected, the target dialogue model is used to control the device.
[0063] Referring to Figure 2, when the target dialogue model is used to identify the needs of the current voice content, it is determined whether the detection is for device control or emotional dialogue. The target dialogue model can be used to detect whether the current text content contains the reference text description content, which can be achieved by calculating the similarity between the current text content and the reference text description content. If there is text content in the current text content whose similarity with the reference text description content is higher than the preset similarity threshold, then it can be considered that the current text content contains the reference text description content, and in this case it is determined to use the smart speaker for device control. If there is no text content in the current text content whose similarity with the reference text description content is higher than the preset similarity threshold, it can be considered that the current text content does not contain the reference text description content, and in this case it is determined to use the target dialogue model for voice dialogue interaction.
[0064] The reference text description content may be a first keyword and a second keyword, wherein the first keyword refers to the vocabulary related to device control or emotional dialogue contained in the current text content input. For example, if the input text content contains words such as "play music" and "adjust volume", then the first keyword can be analyzed to determine whether the demand of the current voice content is for device control; the second keyword refers to the vocabulary related to emotion contained in the current text content input. For example, if the input text content contains words such as "happy" and "sad", then the emotional words can be analyzed to determine whether the demand of the current voice content is for voice dialogue interaction.
[0065] As another optional but non-limiting implementation method, the context or voice features corresponding to the current text content are parsed by the target dialogue model, and based on the parsed context or voice features, it is determined whether to use the target dialogue model for voice dialogue interaction or to use the target dialogue model for device control. Context refers to the environment and background of the input text content. For example, if the input text content is about device control, the context is analyzed to determine whether the current voice content requires device control. Voice features refer to the features of the input voice signal that are related to device control or emotional dialogue. For example, if the input voice signal contains specific features such as intonation and speech speed, the voice features are analyzed to determine whether the current voice content requires device control.
[0066] S130: Determine a current execution operation function adopted by the expected execution action information corresponding to the current voice content.
[0067] S140: Control the smart speaker according to the currently executed operation function.
[0068] Referring to Figure 2, after receiving the desired action information corresponding to the current speech content output by the target dialogue model, the action information is parsed and understood. Based on the pre-defined association relationship with the target dialogue model output, the specific meaning and intent of the action information corresponding to the current speech content is determined. Furthermore, based on the specific meaning and intent of the action information corresponding to the current speech content, the current action function skill used by the action information corresponding to the current speech content is matched. The specific task operation is executed by calling the corresponding current action function.
[0069] Among them, a skill is an execution function, execution service, execution module, or other executable code block that performs corresponding processing according to the requirements of an action. Skills are usually associated with actions. Actions represent the specific action or intention to be performed, while skills are the specific capabilities to implement that action. If the action is "play music," then the corresponding skill may be a music playback function that can play music files or tracks specified. In actual applications, skills can be defined and implemented according to different fields and needs. They can be various types of tasks such as simple logical judgment, data processing, image recognition, natural language processing, and device control. By separating and defining actions and skills, it can be more flexible and scalable. Different actions can correspond to different skills, thereby realizing diversified functions and task execution.
[0070] As an optional but non-limiting implementation, determining the current execution operation function adopted by the expected execution action information corresponding to the current voice content includes steps D1-D2:
[0071] Step D1: parse the expected action information corresponding to the current voice content to obtain the action type and action content.
[0072] Step D2: According to the parsed execution action type and execution action content, a current execution operation function that matches the expected execution action information corresponding to the current voice content is selected from the preset execution operation function set.
[0073] Referring to Figure 2, the expected action information "action" corresponding to the current voice content is parsed and understood, and the specific meaning and intention of the expected action information "action" corresponding to the current voice content is determined based on the association relationship pre-defined with the target dialogue model output. Furthermore, the execution action type and execution action content corresponding to the expected action information can be determined, and the current execution operation function skill that matches the expected action information corresponding to the current voice content can be found based on the execution action type or execution action content of the action. Once the skill to be called is determined, the skill operation can be executed by calling the corresponding function, method, or service. When calling a skill, it may be necessary to pass some necessary parameters or context information so that the skill can be executed correctly. After the skill is executed, some output results are returned. These outputs can be processed as needed, such as displayed to the user, stored in a database, or continued to perform other operations.
[0074] As an optional but non-limiting implementation, controlling the smart speaker according to the currently executed operation function includes steps E1-E2:
[0075] Step E1: Call the executable function corresponding to the current execution operation function, and execute the task operation according to the execution action content corresponding to the expected execution action information corresponding to the current voice content.
[0076] Step E2: Control the smart speaker according to the output result after the executable function is executed.
[0077] Skills pre-define corresponding executable functions for different execution functions. For example, for a smart lighting device, you can define a skill to control the on / off status of the smart lighting device. In the smart home system, HomeAssistant supports the control of multiple smart home devices and realizes the interconnection of multiple smart devices.
[0078] Optionally, common Skills include but are not limited to the following: Lighting control: control the on / off, brightness, color, etc. of lights in the home. Curtain control: control the on / off, percentage of curtains, etc. Temperature control: control the temperature and mode of air conditioners, heaters and other equipment. Home appliance control: control the on / off and functions of home appliances such as TVs, stereos, washing machines, ovens, etc. Security monitoring: monitor doors, windows, cameras, etc. in the home, and provide intrusion alarm functions. Smart door locks: control the on / off of door locks, and provide remote unlocking, password unlocking and other functions. Health monitoring: monitor environmental factors such as air quality and humidity in the home, and provide health advice. Energy management: monitor energy usage in the home and provide energy-saving advice. These are just some common Skills. Smart home systems can provide more types of Skills based on needs and device functions.
[0079] The technical solution of the embodiment of the present disclosure determines in real time the current text content corresponding to the current voice content input to the smart speaker, and parses the current text content through the target dialogue model to obtain the expected execution action information corresponding to the current voice content, and determines the current execution operation function adopted by the expected execution action information corresponding to the current voice content. By accurately identifying the input voice content, the appropriate execution operation function can be matched according to the action required to be executed by the voice content, so that the smart speaker can better understand the usage needs, thereby providing more intelligent and personalized services. Moreover, since different operation functions are executed according to needs, such as playing music, checking the weather, controlling smart homes, etc., the functional scalability of the smart speaker is enhanced, the efficiency and convenience of interaction with the smart speaker are greatly improved, and the smart speaker can be used more easily and conveniently.
[0080] Figure 3 is a structural diagram of a smart speaker processing device provided by an embodiment of the present disclosure. The embodiment of the present disclosure is suitable for enabling a smart speaker to accurately identify the need to perform an action so as to adapt to and control the corresponding operation function. The smart speaker processing device can be implemented in the form of software and / or hardware and is generally integrated into any electronic device with network communication function, which can be a mobile terminal, PC or server, etc.
[0081] As shown in FIG3 , the smart speaker processing device according to an embodiment of the present disclosure may include:
[0082] A determination module 310 is configured to determine the current text content corresponding to the current voice content input to the smart speaker;
[0083] [Corrected 14.05.2025 according to Rule 91] Parsing module 320, for parsing the current text content using a target dialogue model to obtain information about the desired action corresponding to the current speech content, wherein the target dialogue model is pre-trained based on a natural language model;
[0084] A matching module 330 is used to determine the current execution operation function used by the expected execution action information corresponding to the current voice content;
[0085] The control module 340 is used to control the smart speaker according to the currently executed operation function.
[0086] Based on the optional solutions of the above embodiment, optionally, determining the current text content corresponding to the current voice content input to the smart speaker includes:
[0087] Get the current voice content collected by the microphone configured on the smart speaker;
[0088] By performing speech-to-text recognition on the current speech content, current text content corresponding to the current speech content is obtained.
[0089] Based on the optional solution of the above embodiment, optionally, the current text content is parsed by the target dialogue model to obtain the expected execution action information corresponding to the current voice content, including:
[0090] When it is detected that a target dialogue model needs to be used for voice dialogue interaction, reference dialogue guidance information corresponding to the target dialogue model is determined through the target dialogue model, where the reference dialogue guidance information includes the previous text content input to the target dialogue model and expected execution action information corresponding to the previous voice content obtained by recognizing the previous text content through the target dialogue model, where the previous text content is the text content corresponding to the previous voice content input to the smart speaker;
[0091] The reference dialogue guidance information and the current text content are parsed through the target dialogue model to obtain the first expected execution action information corresponding to the current voice content. The first expected execution action information is used to instruct the smart speaker to perform a first task operation, and the first task operation is to feedback matching voice content for the current voice content.
[0092] Based on the optional solution of the above embodiment, optionally, the current text content is parsed by the target dialogue model to obtain the expected execution action information corresponding to the current voice content, including:
[0093] When it is detected that the target dialogue model needs to be used to control the device, the target dialogue model is used to obtain the second expected execution action information corresponding to the current voice content. The second expected execution action information is used to instruct the smart speaker to perform a second task operation. The second task operation is to generate a matching device control instruction for the current voice content.
[0094] Based on the optional solution of the above embodiment, optionally, before parsing the current text content through the target dialogue model to obtain the expected execution action information corresponding to the current voice content, the method further includes:
[0095] Detecting, through the target dialogue model, whether the current text content contains reference text description content, wherein the reference text description content is used to indicate whether the smart speaker is used for device control;
[0096] If no reference text description content is detected, the target dialogue model is used for voice dialogue interaction;
[0097] If the reference text description content is detected, the target dialogue model is used to control the device.
[0098] Based on the optional solutions of the above embodiment, optionally, determining the current execution operation function adopted by the expected execution action information corresponding to the current voice content includes:
[0099] Parsing the expected action type and action content from the expected action information corresponding to the current voice content;
[0100] According to the parsed execution action type and execution action content, a current execution operation function that matches the expected execution action information corresponding to the current voice content is selected from the preset execution operation function set.
[0101] Based on the optional solution of the above embodiment, optionally, controlling the smart speaker according to the currently executed operation function includes:
[0102] Calling the executable function corresponding to the currently executed operation function, and executing the task operation according to the execution action content corresponding to the expected execution action information corresponding to the current voice content;
[0103] The smart speaker is controlled according to the output result after the executable function is executed.
[0104] The technical solution of the embodiment of the present disclosure determines in real time the current text content corresponding to the current voice content input to the smart speaker, and parses the current text content through the target dialogue model to obtain the expected execution action information corresponding to the current voice content, and determines the current execution operation function adopted by the expected execution action information corresponding to the current voice content. By accurately identifying the input voice content, the appropriate execution operation function can be matched according to the action required to be executed by the voice content, so that the smart speaker can better understand the usage needs, thereby providing more intelligent and personalized services. Moreover, since different operation functions are executed according to needs, such as playing music, checking the weather, controlling smart homes, etc., the functional scalability of the smart speaker is enhanced, the efficiency and convenience of interaction with the smart speaker are greatly improved, and the smart speaker can be used more easily and conveniently.
[0105] The smart speaker processing device provided in the embodiments of the present disclosure can execute the smart speaker processing method provided in any embodiment of the present disclosure, and has functional modules and beneficial effects corresponding to the execution method.
[0106] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the protection scope of the embodiments of the present disclosure.
[0107] FIG4 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Referring to FIG4 , a schematic diagram of the structure of an electronic device (such as a terminal device or server in FIG4 ) 400 suitable for implementing an embodiment of the present disclosure is shown below. The terminal device in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. The electronic device shown in FIG4 is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present disclosure.
[0108] As shown in FIG4 , the electronic device 400 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 into a random access memory (RAM) 403. Various programs and data required for the operation of the electronic device 400 are also stored in the RAM 403. The processing device 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An edit / output (I / O) interface 405 is also connected to the bus 404.
[0109] Typically, the following devices may be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 408 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 409. The communication device 409 may allow the electronic device 400 to communicate with other devices wirelessly or by wire to exchange data. Although FIG4 shows the electronic device 400 with various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may alternatively be implemented or present.
[0110] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 409, or installed from the storage device 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0111] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0112] The electronic device provided by the embodiment of the present disclosure and the smart speaker processing method provided by the above embodiment belong to the same inventive concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0113] An embodiment of the present disclosure provides a computer storage medium on which a computer program is stored. When the program is executed by a processor, the smart speaker processing method provided by the above embodiment is implemented.
[0114] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0115] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.
[0116] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0117] [Corrected on 14.05.2025 according to Rule 91] The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device: determines the current text content corresponding to the current voice content input into the smart speaker; parses the current text content through a target dialogue model to obtain the expected execution action information corresponding to the current voice content, and the target dialogue model is pre-trained based on a natural language model; determines the current execution operation function adopted by the expected execution action information corresponding to the current voice content; and controls the smart speaker according to the current execution operation function.
[0118] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0119] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0120] The units involved in the embodiments described in this disclosure may be implemented in software or hardware. In some cases, the name of a unit does not limit the unit itself. For example, the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses."
[0121] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0122] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0123] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0124] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0125] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. An intelligent speaker processing method, the method comprising: Determine the current text content corresponding to the current voice content input to the intelligent speaker; Parse the current text content through a target dialogue model to obtain the expected execution action information corresponding to the current voice content, and the target question-and-answer model is pre-trained based on a natural language model; Determine the current execution operation function adopted by the expected execution action information corresponding to the current voice content; Control the intelligent speaker according to the current execution operation function.
2. The method according to claim 1, wherein, Determine the current text content corresponding to the current voice content input to the intelligent speaker, including: Obtain the current voice content collected by a microphone configured on the intelligent speaker; Perform speech-to-text recognition on the current voice content to obtain the current text content corresponding to the current voice content.
3. The method according to claim 1, wherein, Parse the current text content through a target dialogue model to obtain the expected execution action information corresponding to the current voice content, including: When it is detected that the target dialogue model needs to be used for voice dialogue interaction, determine the reference dialogue guidance information corresponding to the target dialogue model through the target dialogue model, where the reference dialogue guidance information includes the previous text content input to the target dialogue model and the expected execution action information corresponding to the previous voice content obtained by the target dialogue model through recognizing the previous text content, and the previous text content is the text content corresponding to the previous voice content input to the intelligent speaker; Parse the reference dialogue guidance information and the current text content through the target dialogue model to obtain the first expected execution action information corresponding to the current voice content, and the first expected execution action information is used to instruct the intelligent speaker to execute a first task operation, and the first task operation is to feedback a matching voice content for the current voice content.
4. The method according to claim 1, wherein Parse the current text content through a target dialogue model to obtain the expected execution action information corresponding to the current voice content, including: When it is detected that the target dialogue model needs to be used for device control, parse the current text content through the target dialogue model to obtain the second expected execution action information corresponding to the current voice content, and the second expected execution action information is used to instruct the intelligent speaker to execute a second task operation, and the second task operation is to generate a matching device control instruction for the current voice content.
5. The method according to claim 4 or 5, wherein Before parsing the current text content through the target dialogue model to obtain the expected execution action information corresponding to the current voice content, it further includes: Detect whether there is a reference text description content in the current text content through the target dialogue model, and the reference text description content is used to indicate whether to use the intelligent speaker for device control; If the reference text description content is not detected, use the target dialogue model for voice dialogue interaction; If the reference text description content is detected, use the target dialogue model for device control.
6. The method according to claim 1, wherein, Determine the current execution operation function adopted by the expected execution action information corresponding to the current voice content, including: Parse the execution action type and execution action content from the expected execution action information corresponding to the current voice content; Select the current execution operation function that matches the expected execution action information corresponding to the current voice content from the preset execution operation function set according to the parsed execution action type and execution action content.
7. The method according to claim 1, wherein, Control the smart speaker according to the current execution operation function, including: Call the executable function corresponding to the current execution operation function, and execute the task operation according to the execution action content corresponding to the expected execution action information corresponding to the current voice content; Control the smart speaker according to the output result after the execution of the executable function.
8. A smart speaker processing device, the device includes: A determination module, configured to determine the current text content corresponding to the current voice content input to the smart speaker; An analysis module, configured to analyze the current text content through a target dialogue model to obtain the expected execution action information corresponding to the current voice content, and the target question-and-answer model is pre-trained based on a natural language model; A matching module, configured to determine the current execution operation function adopted by the expected execution action information corresponding to the current voice content; A control module, configured to control the smart speaker according to the current execution operation function.
9. An electronic device, the electronic device includes: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the smart speaker processing method according to any one of claims 1-7.
10. A storage medium containing computer-executable instructions, the computer-executable instructions are used to execute the smart speaker processing method according to any one of claims 1-7 when executed by a computer processor.
Citation Information
Patent Citations
Conversation state tracking method and device and calculation equipment
CN111090728A
Conversation understanding method and device, readable medium and electronic equipment
CN112307774A
Voice control method and device, electronic equipment and readable storage medium
CN116705018A
Using augmentation to create natural language models
US20220084511A1
Smart dialogue system and method of integrating enriched semantics from personal and contextual learning
WO2019214799A1
Cited By
Intelligent data analysis method and device based on model, medium and electronic equipment
CN122369037A