Intelligent sound box processing method and device, electronic equipment and storage medium
By converting text and analyzing the voice content input by the smart speaker, obtaining the expected execution action information, the problem of inaccurate recognition of smart speakers in voice response and device control is solved, and more efficient and convenient user interaction is achieved.
Patent Information
- Application Number
- CN202311865315.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-01
AI Technical Summary
When smart speakers realize voice response or smart device control, it is difficult to accurately recognize user voice commands, resulting in the wrong selection of functions.
By determining the text content corresponding to the current voice content and parsing it with a pre-trained natural language model, the expected execution action information is obtained, thereby matching and executing the corresponding operation function.
It realizes accurate recognition and understanding of voice content, enhances the functional scalability and interactive efficiency of smart speakers, and allows users to use smart speakers more easily and conveniently.
Smart Images

Figure CN120236583A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the technical field of smart speakers, and in particular, to a smart speaker processing method, device, electronic device, and storage medium. Background Art
[0002] With the continuous development of intelligent technologies, smart speakers are becoming increasingly popular. Most smart speakers have a voice interaction function, and the smart speaker responds according to the voice content after receiving the voice to implement voice response or perform intelligent device control functions. However, when implementing different functions such as voice response or intelligent device control, the smart speaker may not accurately recognize the user's voice command, resulting in the execution of the wrong function. Summary of the Invention
[0003] The present disclosure provides a smart speaker processing method, device, electronic device, and storage medium to accurately recognize that the operation function indicated by the voice is only for controlling the smart speaker.
[0004] In a first aspect, an embodiment of the present disclosure provides a smart speaker processing method, the method including:
[0005] Determine the current text content corresponding to the current voice content input to the smart speaker;
[0006] Parse the current text content through a target dialogue model to obtain the expected execution action information corresponding to the current voice content, where the target question-and-answer model is pre-trained based on a natural language model;
[0007] Determine the current execution operation function adopted by the expected execution action information corresponding to the current voice content;
[0008] Control the smart speaker according to the current execution operation function.
[0009] In a second aspect, an embodiment of the present disclosure further provides a smart speaker processing device, the device including:
[0010] A determination module, configured to determine the current text content corresponding to the current voice content input to the smart speaker;
[0011] An analysis module, configured to parse the current text content through a target dialogue model to obtain the expected execution action information corresponding to the current voice content, where the target question-and-answer model is pre-trained based on a natural language model;
[0012] A matching module, configured to determine the current execution operation function adopted by the expected execution action information corresponding to the current voice content;
[0013] A control module for controlling the smart speaker according to the current execution operation function.
[0014] In a third aspect, an electronic device is further provided in the embodiments of the present disclosure. The electronic device includes:
[0015] At least one processor; and
[0016] A memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores a computer program executable by the at least one processor. The computer program is executed by the at least one processor so that the at least one processor can execute the smart speaker processing method according to any one of the above embodiments.
[0018] In a fourth aspect, a computer-readable medium is further provided in the embodiments of the present disclosure. The computer-readable medium stores computer instructions for causing a processor to implement the smart speaker processing method according to any one of the above embodiments when executed.
[0019] The technical solution of the embodiments of the present disclosure determines in real time the current text content corresponding to the current voice content input to the smart speaker, and parses the current text content through a target dialogue model to obtain the expected execution action information corresponding to the current voice content, and determines the current execution operation function adopted by the expected execution action information corresponding to the current voice content. By accurately recognizing the input voice content, an appropriate execution operation function can be matched according to the action required by the voice content, so that the smart speaker can better understand the usage requirements, thereby providing more intelligent and personalized services. Moreover, since different operation functions are executed according to the requirements, such as playing music, querying the weather, controlling smart home appliances, etc., the function expansion of the smart speaker is enhanced, and the interaction efficiency and convenience with the smart speaker are greatly improved, making it possible to use the smart speaker more easily and conveniently.
[0020] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In combination with the drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and the original elements and elements are not necessarily drawn to scale.
[0022] Figure 1It is a schematic flowchart of a method for processing an intelligent speaker provided by an embodiment of the present disclosure;
[0023] Figure 2 It is a schematic framework diagram of a processing of an intelligent speaker provided by an embodiment of the present disclosure;
[0024] Figure 3 It is a schematic structural diagram of a device for processing an intelligent speaker provided by an embodiment of the present disclosure;
[0025] Figure 4 It is a schematic structural diagram of an electronic device for implementing a method for processing an intelligent speaker provided by an embodiment of the present disclosure. Detailed implementation manners
[0026] The embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0027] It should be understood that the various steps recorded in the method embodiments of the present disclosure can be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0028] The term "including" and its variants used herein are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0029] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence relationship of the functions performed by these devices, modules or units.
[0030] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".
[0031] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0032] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained through appropriate means in accordance with relevant laws and regulations.
[0033] For example, when responding to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.
[0034] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0035] It is understandable that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present disclosure, and other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0036] It is understandable that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of corresponding laws, regulations and related provisions.
[0037] Figure 1 FIG. is a schematic flowchart of a method for processing an intelligent speaker provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to the situation where the intelligent speaker can accurately identify the execution action requirement to adapt to the corresponding execution operation function for control. This method can be executed by an intelligent speaker processing device, and the intelligent speaker processing device can be implemented in the form of software and / or hardware and is generally integrated on any electronic device with network communication function, and the electronic device can be a mobile terminal, a PC terminal or a server, etc.
[0038] As Figure 1 shown, the method for processing an intelligent speaker in the embodiment of the present disclosure may include the following processes:
[0039] S110. Determine the current text content corresponding to the current voice content input to the intelligent speaker.
[0040] See Figure 2, The smart speaker receives the current voice content input from the outside in real time, and converts the current voice content into the corresponding current text content through the built-in voice recognition function in the smart speaker. For example, the current voice content can be preprocessed, including but not limited to noise reduction, reverberation removal, enhancement, etc., to improve the quality and recognizability of the current voice content, and then perform speech-to-text processing on the current voice content to obtain the corresponding current text content.
[0041] As an optional but non-limiting implementation manner, determining the current text content corresponding to the current voice content input to the smart speaker includes steps A1 - A2:
[0042] Step A1, obtain the current voice content collected by the microphone configured on the smart speaker.
[0043] Step A2, through speech-to-text recognition of the current voice content, obtain the current text content corresponding to the current voice content.
[0044] See Figure 2 , A microphone (such as a microphone) is built into the smart speaker, and the voice signal emitted can be captured from the outside of the smart speaker through the microphone and other microphones. Optionally, considering that some errors may occur during the speech-to-text process, for example, some words in the speech may be misrecognized as other words, so a language model can be used to predict the possible text content, and context information can be used to determine the correct text content, and the result of speech-to-text is corrected to ensure its accuracy. For the current text content generated by speech-to-text, preprocessing such as removing punctuation marks, converting to lowercase, and removing stop words is performed on the recognized current text content to improve the quality and comprehensibility of the text content.
[0045] S120, Parse the current text content through the target dialogue model to obtain the expected execution action information corresponding to the current voice content, and the target question-and-answer model is pre-trained based on the natural language model.
[0046] The target dialogue model is a natural language model based on deep learning. For example, a natural language model based on the Transformer structure is trained using a large-scale corpus to enable the target dialogue model to learn the statistical laws and semantic representations of the language. The target dialogue model can be used for various natural language processing tasks, such as text generation, question and answer, machine translation, etc.
[0047] See Figure 2, the current text content corresponding to the current voice content is input into the target dialogue model, which can automatically parse out the required execution action requirements for the smart speaker included in the current text content, and then generate the expected execution action information that matches the current voice content. The expected execution action information corresponding to the current voice content includes the specific actions expected to be performed by the smart speaker, and the expected execution action information can include device control instructions, emotional dialogue instructions, etc. If the expected execution action information is a device control instruction, the corresponding device control operation can be performed through the voice control function of the smart speaker; if the expected execution action information is an emotional dialogue instruction, the corresponding emotional dialogue can be carried out through the voice interaction function of the smart speaker.
[0048] As an optional but non-limiting implementation, by parsing the current text content through the target dialogue model, the expected execution action information corresponding to the current voice content is obtained, including steps B1 - B2:
[0049] Step B1, when it is detected that voice dialogue interaction needs to be performed using the target dialogue model, determine the reference dialogue guidance information corresponding to the target dialogue model through the target dialogue model. The reference dialogue guidance information includes the previous text content input to the target dialogue model and the expected execution action information corresponding to the previous voice content obtained by the target dialogue model through the recognition of the previous text content. The previous text content is the text content corresponding to the previous voice content input to the smart speaker.
[0050] Step B2, parse the reference dialogue guidance information and the current text content through the target dialogue model to obtain the first expected execution action information corresponding to the current voice content. The first expected execution action information is used to instruct the smart speaker to perform the first task operation, and the first task operation is to feedback the matching voice content for the current voice content.
[0051] See Figure 2 , the device control of the smart speaker includes voice dialogue interaction, home control, voice broadcast (including but not limited to weather forecast, music playback, etc.) control functions. During voice dialogue interaction, dialogue understanding and analysis need to be carried out between both parties to give the dialogue result. Without context information, it may not be possible to accurately understand the requirements of voice dialogue interaction, and thus an accurate answer cannot be provided. For example, if you say to the smart speaker "I want to travel to place A", the target dialogue model may answer some travel information about place A, but without providing more context information, it may not be possible to accurately understand the dialogue requirements and it is not clear whether you want to know the tourist attractions in place A, or the food and accommodation in place A, etc.
[0052] See Figure 2, therefore, when using the target dialogue model to parse the current text content in the voice dialogue interaction scenario, not only the current text content corresponding to the current voice content is required, but also the dialogue requirements of the voice dialogue interaction need to be understood through continuous context. The previous voice content input to the smart speaker is obtained, and the corresponding previous text content is obtained by voice-to-text conversion. The target dialogue model is used to parse and identify the previous text content to determine the expected execution action information corresponding to the previous voice content. Furthermore, the previous text content and the expected execution action information corresponding to the previous voice content are combined into reference dialogue guidance information, and the reference dialogue guidance information is provided as input to the target dialogue model.
[0053] See Figure 2 , based on the reference dialogue guidance information and the current text content, the target dialogue model parses the current text content to obtain the first expected execution action information corresponding to the current voice content. The target dialogue model uses the reference dialogue guidance information and the context of the current text content to generate the first expected execution action information corresponding to the current voice content. It can be seen that by continuously understanding the dialogue requirements of the voice dialogue interaction through context, a more natural and fluent dialogue experience is brought, and a more accurate and useful answer is given because the dialogue requirements can be better understood.
[0054] Through the above process, the process of determining the reference dialogue guidance information according to the target dialogue model is realized, and the context dimension is reflected by including the previous text content corresponding to the previous voice content and the expected execution action information. In this way, the target dialogue model can refer to the previous dialogue content and the expected execution action when generating a response, improving the coherence and accuracy of the dialogue.
[0055] As another optional but non-limiting implementation manner, parsing the current text content by the target dialogue model to obtain the expected execution action information corresponding to the current voice content includes the following steps:
[0056] When it is detected that the target dialogue model needs to be used for device control, the current text content is parsed by the target dialogue model to obtain the second expected execution action information corresponding to the current voice content. The second expected execution action information is used to instruct the smart speaker to execute a second task operation, and the second task operation is to generate a matching device control instruction for the current voice content.
[0057] When performing device control, it is usually not necessary to conduct dialogue understanding and analysis between both parties to obtain a relatively appropriate result. Therefore, the second expected execution action information corresponding to the current voice content can be directly obtained through the target dialogue model for the current text content. Among them, device control includes control functions such as home control and voice broadcast (including but not limited to weather forecast, music playback, etc.). Among them, the weather forecast supports obtaining 7-day weather forecast information for cities across the country, supporting current IP address location to the city, and supporting extended questions, such as: Is it suitable for hiking in Area A this Saturday? What is the clothing matching suggestion for tomorrow?; Music playback supports music playback, specifying songs, and supporting song recommendations based on mood, time period, and atmosphere.
[0058] As an optional but non-limiting implementation method, before parsing the current text content through the target dialogue model to obtain the expected execution action information corresponding to the current voice content, the following steps C1 - C3 are also included:
[0059] Step C1: Detect whether there is reference text description content in the current text content through the target dialogue model, and the reference text description content is used to indicate whether to use a smart speaker for device control.
[0060] Step C2: If no reference text description content is detected, perform voice dialogue interaction using the target dialogue model.
[0061] Step C3: If reference text description content is detected, perform device control using the target dialogue model.
[0062] See Figure 2 , when identifying the requirements for the current voice content through the target dialogue model, determine whether to perform device control or emotional dialogue. The target dialogue model can be used to detect whether the current text content contains reference text description content, which can be specifically implemented by calculating the similarity between the current text content and the reference text description content. If there is text content in the current text content whose similarity to the reference text description content is higher than the preset similarity threshold, then it can be considered that the current text content contains reference text description content, and at this time, it is determined to use a smart speaker for device control. If there is no text content in the current text content whose similarity to the reference text description content is higher than the preset similarity threshold, it can be considered that the current text content does not contain reference text description content, and at this time, it is determined to use the target dialogue model for voice dialogue interaction.
[0063] Among them, the reference text description content can be the first keyword and the second keyword. The first keyword refers to the words related to device control or emotional dialogue contained in the current text content input. For example, if the input text content contains words such as "play music" and "adjust volume", then the needs of the current voice content can be determined by analyzing the first keyword to perform device control; the second keyword refers to the words related to emotion contained in the current text content input. For example, if the input text content contains words such as "happy" and "sad", then the needs of the current voice content can be determined by analyzing the emotion words to perform voice dialogue interaction.
[0064] As another optional but non-limiting implementation, the context or voice features corresponding to the current text content are parsed by the target dialogue model, and based on the parsed context or voice features, it is determined to perform voice dialogue interaction or device control using the target dialogue model. Among them, the context refers to the environment and background in which the input text content is located. For example, if the input text content is about device control, the needs of the current voice content are determined by analyzing the context to perform device control; the voice features refer to the features related to device control or emotional dialogue contained in the input voice signal. For example, if the input voice signal contains specific intonation, speech rate and other features, then the needs of the current voice content are determined by analyzing the voice features to perform device control.
[0065] S130. Determine the current execution operation function used for the expected execution action information corresponding to the current voice content.
[0066] S140. Control the smart speaker according to the current execution operation function.
[0067] See Figure 2 , after receiving the expected execution action information action corresponding to the current voice content output by the target dialogue model, the expected execution action information action corresponding to the current voice content is parsed and understood, and according to the association relationship defined in advance before the output of the target dialogue model, the specific meaning and intention of the expected execution action information action corresponding to the current voice content are determined. Furthermore, based on the specific meaning and intention of the expected execution action information action corresponding to the current voice content, the current execution operation function skill used for the expected execution action information corresponding to the current voice content is matched. The specific task operation is executed by calling the corresponding current execution operation function.
[0068] Among them, skill is an execution function, execution service, execution module, or other executable code block that performs corresponding processing according to the requirements of action. Skill is usually associated with action. Action represents the specific action or intention to be executed, while skill is the specific ability to implement that action. If the action is "play music", the corresponding skill may be a music playback function that can play according to the specified music file or track. In actual applications, skills can be defined and implemented according to different fields and requirements, and can be various types of tasks such as simple logical judgment, data processing, image recognition, natural language processing, device control, etc. By separating and defining action and skill, it can be more flexible and extensible. Different actions can correspond to different skills, thus realizing diverse functions and task executions.
[0069] As an optional but non-limiting implementation manner, the current execution operation function used to determine the expected execution action information corresponding to the current speech content includes steps D1 - D2:
[0070] Step D1: Parse the execution action type and execution action content from the expected execution action information corresponding to the current speech content.
[0071] Step D2: According to the parsed execution action type and execution action content, select the current execution operation function that matches the expected execution action information corresponding to the current speech content from the preset execution operation function set.
[0072] See Figure 2 , parse and understand the expected execution action information action corresponding to the current speech content. According to the pre-defined association relationship before the output of the target dialogue model, determine the specific meaning and intention of the expected execution action information action corresponding to the current speech content. Furthermore, the execution action type and execution action content corresponding to the expected execution action information can be determined, and the current execution operation function skill that matches the expected execution action information corresponding to the current speech content can be found according to the execution action type or execution action content of action. Once the skill to be called is determined, the operation of the skill can be executed by calling the corresponding function, method, or service. When calling the skill, some necessary parameters or context information may need to be passed so that the skill can execute correctly. After the skill is executed, some output results are returned. These outputs can be processed as needed, such as displayed to the user, stored in the database, or continue to execute other operations.
[0073] As an optional but non-limiting implementation, the smart speaker is controlled according to the currently executed operation function, including steps E1 - E2:
[0074] Step E1: Invoke the executable function corresponding to the currently executed operation function, and execute the task operation according to the execution action content corresponding to the expected execution action information corresponding to the current voice content.
[0075] Step E2: Control the smart speaker according to the output result after the executable function is executed.
[0076] For different execution operation functions (Skills), their respective corresponding executable functions are predefined. For example, for a smart lighting device, a Skill can be defined to control the on / off state of this smart lighting device. In a smart home system, it is supported to control multiple smart homes through HomeAssistant to achieve interconnection and interoperability of multiple smart devices.
[0077] Optionally, common Skills include but are not limited to the following: Lighting control: Control the on / off, brightness, color, etc. of the lights at home. Curtain control: Control the on / off, percentage of the curtain, etc. Temperature control: Control the temperature, mode, etc. of devices such as air conditioners and heaters. Home appliance control: Control the on / off, functions, etc. of home appliances such as TVs, stereos, washing machines, and ovens. Security monitoring: Monitor doors, windows, cameras, etc. at home and provide intrusion alarm functions. Smart lock: Control the on / off of the lock and provide functions such as remote unlocking and password unlocking. Health monitoring: Monitor environmental factors such as air quality and humidity at home and provide health suggestions. Energy management: Monitor the energy usage situation at home and provide energy-saving suggestions. These are just some common Skills, and the smart home system can provide more types of Skills according to requirements and the functions of devices.
[0078] The technical solution of the embodiments of the present disclosure determines in real time the current text content corresponding to the currently input voice content to the smart speaker, and parses the current text content through the target dialogue model to obtain the expected execution action information corresponding to the current voice content, determines the currently executed operation function adopted by the expected execution action information corresponding to the current voice content. By accurately identifying the input voice content, the appropriate execution operation function can be matched according to the action required by the voice content, enabling the smart speaker to better understand the usage requirements, thereby providing more intelligent and personalized services. Moreover, since different operation functions are executed according to requirements, such as playing music, querying the weather, controlling smart homes, etc., the function extensibility of the smart speaker is enhanced, greatly improving the interaction efficiency and convenience with the smart speaker, making it possible to use the smart speaker more easily and conveniently.
[0079] Figure 3The following is a schematic structural diagram of an intelligent speaker processing device provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to the situation where an intelligent speaker can accurately identify an execution action requirement and thus adapt to a corresponding execution operation function for control. The intelligent speaker processing device can be implemented in the form of software and / or hardware and is generally integrated on any electronic device with network communication functions. The electronic device can be a mobile terminal, a PC terminal, a server, or the like.
[0080] As Figure 3 shown, the intelligent speaker processing device of the embodiment of the present disclosure may include:
[0081] A determination module 310, configured to determine the current text content corresponding to the current voice content input to the intelligent speaker;
[0082] An analysis module 320, configured to analyze the current text content through a target dialogue model to obtain the expected execution action information corresponding to the current voice content. The target question-and-answer model is pre-trained based on a natural language model;
[0083] A matching module 330, configured to determine the current execution operation function adopted by the expected execution action information corresponding to the current voice content;
[0084] A control module 340, configured to control the intelligent speaker according to the current execution operation function.
[0085] Based on the optional solution of the above embodiment, optionally, determining the current text content corresponding to the current voice content input to the intelligent speaker includes:
[0086] Obtaining the current voice content collected by a microphone configured on the intelligent speaker;
[0087] Performing speech-to-text recognition on the current voice content to obtain the current text content corresponding to the current voice content.
[0088] Based on the optional solution of the above embodiment, optionally, analyzing the current text content through a target dialogue model to obtain the expected execution action information corresponding to the current voice content includes:
[0089] When it is detected that a target dialogue model needs to be used for voice dialogue interaction, determining, through the target dialogue model, reference dialogue guidance information corresponding to the target dialogue model. The reference dialogue guidance information includes the previous text content input to the target dialogue model and the expected execution action information corresponding to the previous voice content obtained by recognizing the previous text content through the target dialogue model. The previous text content is the text content corresponding to the previous voice content input to the intelligent speaker;
[0090] The reference dialogue guiding information and the current text content are parsed by the target dialogue model to obtain first expected execution action information corresponding to the current voice content. The first expected execution action information is used to instruct the smart speaker to perform a first task operation, and the first task operation is to feedback matching voice content for the current voice content.
[0091] Based on the optional solutions of the above embodiments, optionally, parsing the current text content by the target dialogue model to obtain expected execution action information corresponding to the current voice content includes:
[0092] When it is detected that the target dialogue model needs to be used for device control, the current text content is parsed by the target dialogue model to obtain second expected execution action information corresponding to the current voice content. The second expected execution action information is used to instruct the smart speaker to perform a second task operation, and the second task operation is to generate a matching device control instruction for the current voice content.
[0093] Based on the optional solutions of the above embodiments, optionally, before parsing the current text content by the target dialogue model to obtain expected execution action information corresponding to the current voice content, it further includes:
[0094] Detect whether there is reference text description content in the current text content through the target dialogue model. The reference text description content is used to indicate whether to use the smart speaker for device control;
[0095] If the reference text description content is not detected, the target dialogue model is used for voice dialogue interaction;
[0096] If the reference text description content is detected, the target dialogue model is used for device control.
[0097] Based on the optional solutions of the above embodiments, optionally, the current execution operation function used to determine the expected execution action information corresponding to the current voice content includes:
[0098] Parse the execution action type and execution action content from the expected execution action information corresponding to the current voice content;
[0099] According to the parsed execution action type and execution action content, select the current execution operation function that matches the expected execution action information corresponding to the current voice content from the preset execution operation function set.
[0100] Based on the optional solutions of the above embodiments, optionally, controlling the smart speaker according to the current execution operation function includes:
[0101] Call the executable function corresponding to the currently executed operation function, and perform a task operation according to the execution action content corresponding to the expected execution action information corresponding to the current voice content;
[0102] Control the smart speaker according to the output result after the execution of the executable function.
[0103] The technical solution of the embodiments of the present disclosure determines in real time the current text content corresponding to the current voice content input to the smart speaker, and parses the current text content through a target dialogue model to obtain the expected execution action information corresponding to the current voice content, and determines the current execution operation function adopted by the expected execution action information corresponding to the current voice content. By accurately identifying the input voice content, it is possible to match a suitable execution operation function according to the action required by the voice content, so that the smart speaker can better understand the usage requirements, thereby providing more intelligent and personalized services. Moreover, since different operation functions are executed according to the requirements, such as playing music, querying the weather, controlling smart home appliances, etc., the function expansion of the smart speaker is enhanced, and the interaction efficiency and convenience with the smart speaker are greatly improved, making it possible to use the smart speaker more easily and conveniently.
[0104] The smart speaker processing device provided by the embodiments of the present disclosure can execute the smart speaker processing method provided by any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects for executing the method.
[0105] It should be noted that the various units and modules included in the above device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the embodiments of the present disclosure.
[0106] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. The following refers to Figure 4 , which shows a schematic structural diagram of an electronic device 400 suitable for implementing the embodiments of the present disclosure (such as Figure 4 the terminal device or server in). The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 4 The electronic device shown is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present disclosure.
[0107] As Figure 4As shown, the electronic device 400 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 401, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the electronic device 400 are also stored. The processing device 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. An editing / output (I / O) interface 405 is also connected to the bus 404.
[0108] Generally, the following devices may be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 408 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 409. The communication device 409 may allow the electronic device 400 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 4 an electronic device 400 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.
[0109] Specifically, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for executing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device 409, or installed from the storage device 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above functions defined in the method of the embodiment of the present disclosure are executed.
[0110] The names of the messages or information interacted between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0111] The electronic device provided by the embodiment of the present disclosure and the intelligent speaker processing method provided by the above embodiment belong to the same inventive concept. Technical details not described in detail in this embodiment may be referred to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0112] The embodiment of the present disclosure provides a computer storage medium, on which a computer program is stored, and when the program is executed by a processor, the intelligent speaker processing method provided by the above embodiment is implemented.
[0113] It should be noted that the above-mentioned computer-readable medium in the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0114] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks (“LAN”), wide area networks (“WAN”), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed network.
[0115] The above-mentioned computer-readable medium may be included in the above-mentioned electronic device; or it may exist separately and not be assembled into the electronic device.
[0116] The above computer-readable medium carries one or more programs which, when executed by the electronic device, cause the electronic device to: determine the current text content corresponding to the current voice content input to the smart speaker; parse the current text content through a target dialogue model to obtain the expected execution action information corresponding to the current voice content, where the target question-and-answer model is pre-trained based on a natural language model; determine the current execution operation function for the expected execution action information corresponding to the current voice content; and control the smart speaker according to the current execution operation function.
[0117] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0118] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that, in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0119] The units involved in the embodiments of the present disclosure can be implemented in software or in hardware. Among them, the name of a unit does not constitute a limitation on the unit itself in some cases. For example, the first acquisition unit can also be described as "the unit for acquiring at least two Internet protocol addresses".
[0120] The functions described above herein can be performed, at least in part, by one or more hardware logic components. By way of example and not limitation, the types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0121] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0122] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the present disclosure.
[0123] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment may also be implemented separately or in any suitable subcombination in multiple embodiments.
[0124] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. An intelligent speaker processing method, characterized in that, The method includes: Determine the current text content corresponding to the current voice content input to the smart speaker; Parse the current text content through a target dialogue model to obtain the expected execution action information corresponding to the current voice content, and the target question-and-answer model is pre-trained based on a natural language model; Determine the current execution operation function for the expected execution action information corresponding to the current voice content; Control the smart speaker according to the current execution operation function.
2. The method according to claim 1, wherein Determine the current text content corresponding to the current voice content input to the smart speaker, including: Obtain the current voice content collected by a microphone configured on the smart speaker; Perform speech-to-text recognition on the current voice content to obtain the current text content corresponding to the current voice content.
3. The method according to claim 1, wherein Parse the current text content through a target dialogue model to obtain the expected execution action information corresponding to the current voice content, including: When it is detected that the target dialogue model needs to be used for voice dialogue interaction, determine the reference dialogue guidance information corresponding to the target dialogue model through the target dialogue model. The reference dialogue guidance information includes the previous text content input to the target dialogue model and the expected execution action information corresponding to the previous voice content obtained by recognizing the previous text content through the target dialogue model. The previous text content is the text content corresponding to the previous voice content input to the smart speaker; Parse the reference dialogue guidance information and the current text content through the target dialogue model to obtain the first expected execution action information corresponding to the current voice content. The first expected execution action information is used to instruct the smart speaker to execute a first task operation, and the first task operation is to feedback a matching voice content for the current voice content.
4. The method according to claim 1, characterized in that Parse the current text content through a target dialogue model to obtain the expected execution action information corresponding to the current voice content, including: When it is detected that the target dialogue model needs to be used for device control, parse the current text content through the target dialogue model to obtain the second expected execution action information corresponding to the current voice content. The second expected execution action information is used to instruct the smart speaker to execute a second task operation, and the second task operation is to generate a matching device control instruction for the current voice content.
5. The method according to claim 4 or 5, characterized in that, Before parsing the current text content through the target dialogue model to obtain the expected execution action information corresponding to the current voice content, it further includes: Detect whether there is reference text description content in the current text content through the target dialogue model. The reference text description content is used to indicate whether to use the smart speaker for device control; If no reference text description content is detected, use the target dialogue model for voice dialogue interaction; If reference text description content is detected, use the target dialogue model for device control.
6. The method according to claim 1, characterized in that, Determine the current execution operation function for the expected execution action information corresponding to the current voice content, including: Parse the execution action type and execution action content from the expected execution action information corresponding to the current voice content; Select the current execution operation function that matches the expected execution action information corresponding to the current voice content from the preset execution operation function set according to the parsed execution action type and execution action content.
7. The method according to claim 1, characterized in that Control the smart speaker according to the current execution operation function, including: Call the executable function corresponding to the current execution operation function, and execute the task operation according to the execution action content corresponding to the expected execution action information corresponding to the current voice content; Control the smart speaker according to the output result after the execution of the executable function.
8. An intelligent speaker processing device, characterized in that, The device includes: A determination module, configured to determine the current text content corresponding to the current voice content input to the smart speaker; An analysis module, configured to analyze the current text content through a target dialogue model to obtain the expected execution action information corresponding to the current voice content, and the target question-and-answer model is pre-trained based on a natural language model; A matching module, configured to determine the current execution operation function adopted by the expected execution action information corresponding to the current voice content; A control module, configured to control the smart speaker according to the current execution operation function.
9. An electronic device, characterized in that, The electronic device includes: One or more processors; A storage device, configured to store one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the smart speaker processing method according to any one of claims 1-7.
10. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions are used to execute the smart speaker processing method according to any one of claims 1-7 when executed by a computer processor.