Voice control method, electronic equipment and computer readable storage medium

By receiving voice signals and generating personalized control instructions, the problem of the inability to personalize the control of smart devices in the prior art is solved, and the effect of smart devices providing users with personalized functional services is achieved.

CN120220672APending Publication Date: 2025-06-27GUANGDONG KETYOO INTELLIGENT TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510312896.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art cannot personalize the control of smart devices based on the voice signals sent by users and provide corresponding functional services.

Method used

By receiving voice signals, determining user information and target tasks, obtaining status information and environment information of the smart device, inputting it into a pre-trained task model, generating personalized control instructions, and sending control instructions to the smart device.

Benefits of technology

It realizes that users send control information through voice signals, and the intelligent device generates personalized control instructions based on user information, status information and environmental information, and provides personalized functional services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220672A_ABST
    Figure CN120220672A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a voice control method, electronic equipment and a computer readable storage medium, and relates to the field of Internet of Things. The method comprises the following steps: receiving a first voice signal, and determining user information and a target task for intelligent equipment according to the first voice signal; the user information is information related to a sender of the first voice signal; if it is determined that the target task is the task executable by the intelligent device, first information is obtained, the first information, the user information and the target task are input into a pre-trained task model, and a control instruction output by the task model is obtained; the first information comprises at least one of state information of the intelligent equipment and environment information of an environment where the intelligent equipment is located; the control instruction is an instruction matched with the first information, the user information and the target task; the control instruction is sent to the intelligent device so that the intelligent device can execute the control instruction, and various functions which are individually provided for a user by the intelligent device are achieved in a voice control mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet of Things technology. Specifically, this application relates to a voice control method, an electronic device, and a computer-readable storage medium. Background Art

[0002] Currently, with the rapid development of artificial intelligence technology, especially the wide application of dialogue large models (such as generative large language models), intelligent voice assistants have become an indispensable part of people's daily lives. They can understand and generate natural language, have smooth conversations with users, and thus provide various services. However, there is currently a problem that intelligent devices cannot be controlled according to the voice signals sent by users to provide corresponding functional services for users in a personalized manner. Summary of the Invention

[0003] Embodiments of this application provide a voice control method, an electronic device, and a computer-readable storage medium, which are used to solve the technical problem that intelligent devices cannot be controlled according to the voice signals sent by users to provide corresponding functional services for users in a personalized manner.

[0004] According to the first aspect of the embodiments of this application, a voice control method is provided, which is applied to a terminal. The method includes: receiving a first voice signal, and determining user information and a target task according to the first voice signal; the user information is information related to the sender of the first voice signal; If it is determined that the target task is a task executable by an intelligent device, obtain first information, input the first information, user information, and target task into a pre-trained task model, and obtain a control instruction output by the task model; the first information includes at least one of the status information of the intelligent device and the environmental information of the environment where the intelligent device is located; the control instruction is an instruction that matches the first information, user information, and target task; Send the control instruction to the intelligent device so that the intelligent device executes the control instruction; The task model is trained with sample user information, a sample task, and sample first information as training samples and a sample control instruction as a training label, and the sample control instruction is an instruction that matches the sample user information, the sample task, and the sample first information.

[0005] In a possible implementation, the task model includes a feature extraction layer and at least one task layer, and each task layer has its own corresponding task; Input the first information and the target task into the feature extraction layer, and obtain general features common to all task layers output by the feature extraction layer; Determine the target task layer corresponding to the target task from at least one task layer; Input user information and general features into the target task layer. The target task layer extracts task-related sub-features corresponding to the task layer from the user information and general features, and obtains a control instruction output by the target task layer based on the obtained sub-features.

[0006] In another possible implementation, extract a first speech feature from the first speech signal; the first speech feature is a feature characterizing the user identity. Calculate the similarity between the first speech feature and the second speech feature of any user in the speech database; the second speech features of each user are pre-stored in the speech database. Select the user with the highest similarity as the target user. According to the target user, obtain the user information of the target user from the information database. The user information of each user is pre-stored in the information database.

[0007] In yet another possible implementation, input the first speech signal into a pre-trained recognition model to obtain the text information output by the recognition model. Perform semantic analysis on the text information to determine the semantic analysis result; the semantic analysis result includes an intention and an entity. If a task template matching the intention is stored in the predefined task list, determine the target task according to the task template and the entity; the task list includes task templates for any executable tasks of the intelligent device and the terminal. The recognition model is trained with sample speech signals as training samples and the corresponding text information of the sample speech signals as training labels.

[0008] In yet another possible implementation, after sending a control instruction to the intelligent device, receive a response message; the response message is sent by the intelligent device after executing the target task. According to the response message, select a corresponding first reply text from the text database; reply texts corresponding to various response messages are pre-stored in the text database. Convert the first reply text into a second speech signal, display the first reply text, and play the second speech signal.

[0009] In yet another possible implementation, if it is determined that the target task is a task executable by the terminal, execute the target task. Determine the execution result, and select a corresponding second reply text from the text database according to the execution result; reply texts corresponding to various execution results are pre-stored in the text database. Convert the second reply text into a third speech signal, display the second reply text, and play the third speech signal. After receiving the first speech signal, extract an emotion feature from the first speech signal. Determine the target emotion corresponding to the emotion feature through a machine learning algorithm; Determine the operation instruction corresponding to the target emotion from the emotion library; operation instructions corresponding to each target emotion are pre-stored in the emotion library; the operation instruction is used to instruct the intelligent device to perform corresponding operations; Send a control instruction to the intelligent device, including: Send a control instruction and an operation instruction to the intelligent device.

[0010] In a possible implementation manner, according to the semantic analysis result, determine the target keyword corresponding to the text information; the target keyword refers to a keyword related to the theme of the text information; Obtain the historical semantic analysis results within a preset time period from the historical dialogue library; Obtain the target historical semantic analysis result from the historical semantic analysis results within the preset time period; the target historical semantic analysis result refers to the historical semantic analysis result that contains the target keyword; Determine the target historical semantic analysis result according to the semantic analysis result and the historical semantic analysis result; Generate a third reply text according to the historical semantic analysis result and a predefined reply template; the reply template is used to generate a question instructing the user to answer; Convert the third reply text into a fourth voice signal, display the third reply text and play the fourth voice signal.

[0011] In yet another possible implementation manner, the user information includes at least one of the following: The height of the user; The weight of the user; The historical operation record of the user for the intelligent device.

[0012] In yet another possible implementation manner, the intelligent device is an intelligent clothes dryer, and the target tasks include at least one of the following: Adjust the height of the drying rack of the intelligent clothes dryer; Detect whether there are clothes on the drying rack of the intelligent clothes dryer; Start corresponding functions according to the environmental information of the environment where the clothes dryer is located; Judge whether the intelligent clothes dryer is abnormal.

[0013] According to the second aspect of the embodiments of the present application, a voice control device is provided. The terminal includes: A receiving module, configured to receive a first voice signal, and determine user information and a target task according to the first voice signal; the user information is information related to the sender of the first voice signal; An acquisition module, configured to obtain first information if it is determined that the target task is a task executable by the intelligent device, input the first information, user information, and the target task into a pre-trained task model, and obtain a control instruction output by the task model; the first information includes at least one of the status information of the intelligent device and the environmental information of the environment where the intelligent device is located; the control instruction is an instruction matching the first information, user information, and the target task. A sending module, configured to send the control instruction to the intelligent device so that the intelligent device executes the control instruction. The task model is trained with sample user information, sample tasks, and sample first information as training samples and sample control instructions as training labels, and the sample control instructions are instructions matching the sample user information, sample tasks, and sample first information.

[0014] According to the third aspect of the embodiments of the present application, an electronic device is provided. The electronic device includes a memory, a processor, and a computer program stored on the memory. When the processor executes the program, the steps of the method provided in the first aspect are implemented.

[0015] According to the fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method provided in the first aspect are implemented.

[0016] According to the fifth aspect of the embodiments of the present application, a computer program product is provided. The computer program product includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. When a processor of a computer device reads the computer instructions from the computer-readable storage medium and the processor executes the computer instructions, the computer device is caused to execute the steps of the method provided in the first aspect.

[0017] The beneficial effects brought by the technical solutions provided by the embodiments of the present application are: The voice control method provided by the embodiments of the present application receives a first voice signal, determines the user information of the voice sender and the target task according to the first voice signal. If it is determined that the target task is a task executable by the intelligent device, it obtains first information representing the status information of the intelligent device and the environmental information of the environment where the intelligent device is located, and inputs the first information, user information and target task into a pre-trained task model together to obtain a control instruction output by the task model that matches the first information, user information and target task, and then sends the control instruction to the intelligent device so that the intelligent device executes the control instruction to complete the target task. The embodiments of the present application perform corresponding control on the intelligent device through the received voice signal, realizing that the user issues control information through the voice signal, and generating personalized control instructions based on the user information, the current status information of the intelligent device and the environmental information, achieving the purpose of instructing the intelligent device to provide various functions for the user in a personalized manner. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description in the embodiments of the present application.

[0019] Figure 1 It is a schematic diagram of the system architecture for implementing the voice control method provided by the embodiments of the present application; Figure 2 It is a schematic flowchart of a voice control method provided by the embodiments of the present application; Figure 3 It is a schematic flowchart of the method for obtaining the control instruction in a voice control method provided by the embodiments of the present application; Figure 4 It is a schematic flowchart of the method for determining user information in a voice control method provided by the embodiments of the present application; Figure 5 It is a schematic flowchart of the method for determining the target task in a voice control method provided by the embodiments of the present application; Figure 6 It is a schematic flowchart of a voice control method provided by the embodiments of the present application; Figure 7 It is a schematic flowchart of another voice control method provided by the embodiments of the present application; Figure 8 It is a schematic diagram of the structure of a voice control device provided by the embodiments of the present application; Figure 9 It is a schematic diagram of the structure of an electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] The embodiments of the present application will be described below with reference to the accompanying drawings in the present application. It should be understood that the embodiments described below in conjunction with the drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions of the embodiments of the present application.

[0021] Those skilled in the art of the present technology can understand that, unless specifically stated, the singular forms "a", "an", "" and "the" used herein may also include the plural forms. It should be further understood that the terms "comprising" and "including" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements, and / or components, but do not exclude the implementation of other features, information, data, steps, operations, elements, components, and / or their combinations supported by the art of the present technology. It should be understood that when we say an element is "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used here can include wireless connection or wireless coupling. The term "and / or" used here indicates at least one of the items defined by the term, for example, "A and / or B" can be implemented as "A", or implemented as "B", or implemented as "A and B".

[0022] To make the purpose, technical solutions, and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the drawings.

[0023] The technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application will be described below through the description of several exemplary embodiments. It should be noted that the following embodiments can refer to, draw on, or combine with each other. For the same terms, similar features, and similar implementation steps in different embodiments, they will not be described repeatedly.

[0024] Figure 1 FIG. is a schematic diagram of the system architecture for implementing the voice control method provided by the embodiment of the present application, where the system architecture includes: a terminal 120, a server 140, and an intelligent device 160.

[0025] The terminal 120 installs and runs an application program of the voice control method. The terminal 120 is used to send a control instruction to the intelligent device according to the received first voice signal.

[0026] The terminal 120 is connected to the server 140 through a wireless network or a wired network.

[0027] The server 140 includes at least one of a server, multiple servers, a cloud computing platform, and a virtualization center. Schematically, the server 140 includes a processor 144 and a memory 142. The memory 142 includes a display module 1421, a control module 1422, and a receiving module 1423. The server 140 is used to provide background services for the application programs of the method. Optionally, the server 140 undertakes the main computing work, and the terminal 120 undertakes the secondary computing work; or, the server 140 undertakes the secondary computing work, and the terminal 120 undertakes the main computing work; or, the server 140 and the terminal 120 perform collaborative computing using a distributed computing architecture.

[0028] Optionally, the device types of the terminal include at least one of a smart phone, a tablet computer, an e-book reader, a Moving Picture Experts Group Audio Layer III (MP3) player, a Moving Picture Experts Group Audio Layer IV (MP4) player, a laptop computer, and a desktop computer.

[0029] Those skilled in the art can know that the number of the above terminals can be more or less. For example, the above terminal can be only one, or the above terminals can be dozens or hundreds, or even more. The embodiments of the present application do not limit the number and device types of the terminals.

[0030] The related technologies are described below: Currently, the interaction experience between users and intelligent devices still needs to be improved. For example, there are still deficiencies in the accuracy, flexibility, and user experience of voice recognition. At the same time, there is a lack of visual voice interaction, and it cannot provide personalized function services for different users, and cannot recognize and respond to the emotions of users.

[0031] In view of at least one of the above technical problems or areas for improvement in the related art, the voice control method provided in the embodiments of the present application receives a first voice signal, determines the user information and the target task of the voice sender according to the first voice signal. If it is determined that the target task is a task executable by the intelligent device, the first information representing the status information of the intelligent device and the environmental information of the environment where the intelligent device is located is obtained, and the first information, the user information, and the target task are input into a pre-trained task model together to obtain a control instruction output by the task model that matches the first information, the user information, and the target task. Then, the control instruction is sent to the intelligent device so that the intelligent device executes the control instruction to complete the target task. The embodiments of the present application perform corresponding control on the intelligent device through the received voice signal, realizing that the user issues control information through the voice signal, and generating personalized control instructions based on the user information, the current status information of the intelligent device, and the environmental information, achieving the purpose of instructing the intelligent device to provide various functions for the user in a personalized manner.

[0032] An embodiment of the present application provides a voice control method, which is applied to a terminal, such as Figure 2 shown, and the method includes: S101, receive a first voice signal, and determine user information and a target task according to the first voice signal.

[0033] In the embodiment of the present application, the first voice signal refers to the sound signal emitted by the user when interacting with the device or system in a voice manner. The first voice signal usually contains the user's requirements, intentions, or instructions, and the terminal can receive the first voice signal emitted by the user through a microphone.

[0034] In the embodiment of the present application, an intelligent device generally refers to a device that can interact, control, or automate operations through the Internet, artificial intelligence, or other intelligent technologies. They can perform certain tasks, usually can sense the environment, collect data, make decisions, or respond to user instructions. The intelligent device can communicate with other devices and often has self-learning and adaptation functions. The intelligent device can be a smart home device, such as a smart thermostat, smart lights, a smart door lock, etc. The intelligent device can also be a smart home appliance, such as a smart refrigerator, a smart air conditioner, a smart washing machine, a smart clothes dryer, etc. The terminal refers to the terminal that establishes a binding relationship with the intelligent device, that is, the intelligent device can be controlled through the terminal.

[0035] In the embodiments of the present application, the user information is information related to the issuer of the first voice signal. For the same intelligent device, there may be different users operating it. Therefore, the user information of each user, such as the historical operation records of each user and the height of each user, etc., will be stored in advance. Thus, when the first voice signal is received, the user information of the issuer of the first voice signal is determined according to the first voice signal, and then the operation habits or operation preferences of the user when issuing the target task can be determined in combination with the user information, so as to complete the target task in a personalized manner.

[0036] In the embodiments of the present application, the target task is parsed from the first voice signal. The target task refers to the task that the user commands the terminal or intelligent device to execute. For example, when the intelligent device is an intelligent clothes dryer, when the target task is a task executed by the terminal, the target task can be providing personalized recommendations, quick answers (such as product usage instructions, common problem descriptions, troubleshooting step guides), virtual image setting, accessory purchase, device binding, installation and repair reporting, and genuine product verification, etc.; when the target task is a task executed by the intelligent clothes dryer, the target task can be adjusting the height of the clothes hanger, starting or closing the drying function of the clothes dryer, detecting whether there are clothes on the clothes hanger, and detecting whether the clothes dryer has a fault.

[0037] In the embodiments of the present application, the first voice signal is received, semantic analysis is performed on the first voice signal to determine the target task, and voice recognition is performed on the first voice signal to determine the identity of the issuer of the first voice information, so as to determine the corresponding user information.

[0038] S102, if it is determined that the target task is a task executable by the intelligent device, then the first information is obtained, and the first information, the user information, and the target task are input into a pre-trained task model to obtain a control instruction output by the task model.

[0039] In the embodiments of the present application, the first information includes at least one of the status information of the intelligent device and the environmental information of the environment where the intelligent device is located. The status information and environmental information of the intelligent device are detected by various sensors installed on the intelligent device.

[0040] In the embodiments of the present application, the status information of the intelligent device may be whether it is in a working state, the current working mode it is in, the battery status of the device, the battery temperature, the connection status between intelligent devices, etc. The above status information is information collected by sensors installed on the intelligent device or detected by the intelligent device itself.

[0041] In the embodiments of the present application, the environmental information of the intelligent device includes the temperature of the environment where the intelligent device is located, the humidity of the environment where the intelligent device is located, the light intensity of the environment where the intelligent device is located, the carbon dioxide concentration of the environment where the intelligent device is located, the noise of the environment where the intelligent device is located, the air pressure of the environment where the intelligent device is located, whether the intelligent device is in a moving state, etc., and can be collected by sensors such as a temperature sensor, a humidity sensor, an image sensor, a weight sensor, and a photosensitive sensor installed on the intelligent device.

[0042] In the embodiments of the present application, when the intelligent device is an intelligent clothes dryer, the first information may be the weight of the clothes, the type of the clothes, the humidity of the clothes, the temperature of the clothes, and the light intensity and light area of the environment of the clothes dryer detected by the sensor.

[0043] In the embodiments of the present application, since the control instruction output by the task model is for the intelligent device, and there is a situation where the target task parsed from the user's first voice signal is not a task that the intelligent device can execute. Therefore, before obtaining the first information, it is necessary to determine whether the target task is a task that the intelligent device can execute. For example, after determining the target task, compare the target task with the tasks in the task list for recording the tasks that the intelligent device can execute. After determining that there is a task associated with the target task in the task list, obtain the first information.

[0044] In the embodiments of the present application, the first information is obtained from the first information library, which is used to store the latest environmental information and status information of the intelligent device. When any sensor on the intelligent device detects a change in the status information and environmental information of the intelligent device, the newly detected status information and / or environmental information will be sent to the first information library, and the first information library will update the corresponding stored status information and environmental information.

[0045] In the embodiments of the present application, the control instruction is an instruction that matches the first information, the user information, and the target task, and the control instruction is used to instruct the intelligent device to execute the target task.

[0046] In the embodiments of the present application, the task model is constructed based on a neural network structure. The task model is used to determine a specific control instruction for controlling the intelligent device according to the user information, the first information, and the target task, so that the intelligent device can personalizedly complete the target task sent by the user through the first voice information.

[0047] In one example, the intelligent device is an intelligent drying rack, the target task is to adjust the height of the drying rack of the intelligent drying machine, the user information is that the user's height is X meters, the first information includes the weight and length of the clothes. Input the above target task, clothes weight, clothes length and height into the task model, and obtain a control instruction output by the task model indicating that the intelligent drying rack moves the drying rack upward by Y meters. In order to avoid the drying machine affecting the user, within the range where the intelligent drying machine can be raised, the larger X is, the larger the value of Y will be accordingly.

[0048] In the embodiment of the present application, the task model is trained with sample user information, sample tasks and sample first information as training samples and sample control instructions as training labels. The sample control instruction is an instruction that matches the sample user information, sample task and sample first information.

[0049] In the embodiment of the present application, for each type of sample task, a large number of different sample user information, sample first information and corresponding control instructions are collected. An appropriate neural network structure can be selected according to the task requirements to construct the task model, such as network structures like long short-term memory network, attention mechanism, deep convolutional neural network, etc. Using the collected large number of sample user information, sample tasks and sample first information as training samples, and the control instructions corresponding to the above sample user information, sample tasks and sample first information as training labels, the corresponding neural network structure is trained until convergence, so as to obtain a trained task model.

[0050] In one example, when the intelligent device is an intelligent drying machine, the sample task is to adjust the height of the drying rack of the intelligent drying machine, and the task model is a linear regression model. A large amount of information related to the adjustment of the drying rack height is collected, and control instructions indicating that the drying rack moves upward by a corresponding height are obtained under different clothes weights, different clothes lengths, different clothes types and different user heights. Take the clothes weight, clothes length and clothes type in each set of the above collected data as sample first information, take the user's height information as sample user information, and take the control instruction indicating that the drying rack moves upward by a corresponding height as sample control instruction. Use the above large number of training samples and corresponding training labels to train the linear regression model, so as to determine the linear relationship between the independent variable (the height of the drying rack) and the dependent variables (user height, clothes weight, clothes length and clothes type), and thus obtain a trained task model.

[0051] S103. Send a control instruction to the intelligent device so that the intelligent device executes the control instruction.

[0052] In the embodiment of the present application, a control instruction is sent to the intelligent device, and the intelligent device that receives the control instruction will complete the target task according to the control instruction.

[0053] In one example, when the intelligent device is an intelligent clothes dryer, the terminal sends a control instruction to the intelligent clothes dryer to move the clothes dryer upward by 50 cm. Then, after receiving the control instruction, the intelligent clothes dryer moves the clothes dryer upward by 50 cm.

[0054] The voice control method provided by the embodiments of the present application receives a first voice signal, determines the user information of the voice sender and the target task according to the first voice signal. If it is determined that the target task is a task executable by the intelligent device, it obtains first information representing the state information of the intelligent device and the environmental information of the environment where the intelligent device is located, and inputs the first information, user information and target task into a pre-trained task model together to obtain a control instruction output by the task model that matches the first information, user information and target task. Then, it sends the control instruction to the intelligent device so that the intelligent device executes the control instruction to complete the target task. The embodiments of the present application perform corresponding control on the intelligent device through the received voice signal, realizing that the user issues control information through the voice signal, and generating personalized control instructions based on the user information, the current state information of the intelligent device and the environmental information personality, achieving the purpose of instructing the intelligent device to provide various functions for the user in a personalized manner.

[0055] On the basis of the above embodiments, as an optional embodiment, the task model includes a feature extraction layer and at least one task layer, and each task layer has its corresponding task.

[0056] In the embodiments of the present application, the feature extraction layer is used to extract general features common to each task layer from the input first information and target task. The general features extracted by the feature extraction layer are used as the input of subsequent task layers. The feature extraction layer usually includes structures such as a convolutional layer, a pooling layer and a fully connected layer, and the above structures can effectively extract low-dimensional representations or high-dimensional features input to the feature extraction layer.

[0057] In the embodiments of the present application, for different tasks of the intelligent device, multiple task layers are set, each task layer corresponds to one task, and each task has its unique goal. Each task layer performs specialized processing on the corresponding task. The task layer extracts features from the general features output by the feature extraction layer and the user information to obtain sub-features related to the corresponding task, and processes the specific requirements and outputs of the corresponding task according to the sub-features, thereby improving the accuracy of each task layer for its corresponding task.

[0058] The embodiments of the present application provide a method for obtaining a control instruction, as Figure 3 shown, the specific content is as follows: S201, input the first information and the target task into the feature extraction layer to obtain general features common to all task layers output by the feature extraction layer; S202. Determine the target task layer corresponding to the target task from at least one task layer; S203. Input the user information and general features into the target task layer. The target task layer extracts task-related sub-features corresponding to the task layer from the user information and general features, and obtains a control instruction output by the target task layer based on the obtained sub-features.

[0059] In S201 of the embodiment of the present application, the first information and the target task input to the task model will be first input to the feature extraction layer, so as to obtain general features useful for all tasks output by the feature extraction layer.

[0060] In S202 of the embodiment of the present application, after the general features are extracted by the feature extraction layer, since the task model has multiple task layers, it is necessary to determine the corresponding target task layer according to the target task. For example, the intelligent device is an intelligent clothes dryer, the task model has task layer A, task layer B, and task layer C. The task corresponding to task layer A is to adjust the height of the clothes drying rack, the task corresponding to task layer B is to detect the abnormal state of the clothes dryer, and the task corresponding to task C is to adaptively start the drying function of the clothes dryer. For example, if the current target task is to detect the abnormal state of the clothes dryer, then task layer B will be used as the target task layer.

[0061] In S203 of the embodiment of the present application, since the tasks corresponding to each task layer are different, after the general features and user information are input to the task layer, the task layer will further extract the corresponding task-related sub-features, and then obtain the output control instruction based on the obtained sub-features.

[0062] In the above solution, the general features common to each task layer are extracted by the feature extraction layer, avoiding the overhead of each task learning features separately. This not only improves the computing efficiency but also reduces the training time and resource consumption. For different target tasks of intelligent devices, corresponding task layers are respectively set for processing to realize the combined use of voice control and multi-tasks, improve the convenience and efficiency of using the functions of intelligent devices, and by using the task model with multiple task layers, the model can process multiple target tasks simultaneously to meet the requirements of complex applications.

[0063] Based on the above embodiments, as an optional embodiment, the method for determining user information is as Figure 4 shown, and the specific content is as follows: S301. Extract first voice features from the first voice signal; S302. Calculate the similarity between the first voice features and the second voice features of any user in the voice library; S303. Select the user with the highest similarity as the target user; S304. Obtain the user information of the target user from the information database according to the target user.

[0064] In S301 of the embodiment of the present application, the first voice feature is a feature characterizing the user identity. When extracting features from the first voice signal, the voice signal is first preprocessed, such as denoising, removing silent segments, and voice framing. During the process of extracting the first voice feature from the first voice signal, the preprocessed first voice signal can be mapped to a first low-dimensional vector space through a deep neural network, so as to obtain the first voice feature that can effectively represent the identity feature of the sender of the first voice signal.

[0065] In S302 of the embodiment of the present application, the second voice features of each user are pre-stored in the voice database, that is, the second voice features of all users who have the permission to use the intelligent device are stored in the voice database. Therefore, after obtaining the first voice feature, in order to determine which specific user the sender of the currently received first voice signal is, it is necessary to calculate the similarity between the first voice feature and the second voice features of each user in the voice database.

[0066] In the embodiment of the present application, the similarity can be determined by calculating the cosine similarity between the first voice feature and the second voice feature. If the cosine similarity is close to 1, it means that the two feature vectors are very similar and the similarity is higher; if it is close to 0, it means there is no similarity and the similarity is lower. The similarity can also be determined by calculating the Euclidean distance between the two vectors. The smaller the distance, the more similar the two features are and the higher the similarity.

[0067] In S303 of the embodiment of the present application, select the user corresponding to the second voice feature with the highest similarity to the first voice feature from the calculated similarities as the target user.

[0068] In S304 of the embodiment of the present application, the user information of each user is pre-stored in the information database. Therefore, after determining the target user, directly obtain the corresponding user information from the information database.

[0069] In the above solution, the first voice feature representing the user identity information is extracted through the first voice signal; the similarity between the first voice feature and the second voice features of each user pre-stored in the voice database is calculated; the user with the highest similarity is selected as the target user, so as to obtain the user information of the target user; the identity of the voice sender is recognized through the first voice signal, and specifically, the user information of the user who currently issues the target task is determined, providing reference information for subsequent generation of personalized control instructions.

[0070] Based on the above embodiments, as an alternative embodiment, the method for determining the target task is as Figure 5As shown below, the specific content is as follows: S401: Input the first voice signal into a pre-trained recognition model to obtain the text information output by the recognition model. S402: Perform semantic analysis on the text information to determine the semantic analysis result. S403: If a task template matching the intent is stored in the predefined task list, determine the target task according to the task template and the entity. In S401 of the embodiment of the present application, the first voice signal is preprocessed, and then the preprocessed first voice signal is input into the pre-trained recognition model. The recognition model extracts features from the first voice signal and uses the extracted features as text information, thereby obtaining the text information output by the recognition model.

[0071] In the embodiment of the present application, preprocessing of the first voice signal includes sampling, quantization, filtering, etc. After preprocessing, feature representations available for the recognition model are extracted from the original model. The feature representations include: spectral features, Mel-frequency cepstral coefficients, etc. Mel-frequency cepstral coefficients are a feature representation method based on the auditory characteristics of the human ear. It maps the spectrum of the voice signal to the Mel scale, and performs operations such as logarithmic transformation and discrete cosine transformation to obtain a series of coefficients that can reflect the spectral characteristics of the voice signal. Mel-frequency cepstral coefficients have been widely used in speech recognition because they can better reflect the phonetic features in speech (such as the difference between vowels and consonants).

[0072] In the embodiment of the present application, sampling in the preprocessing process is the process of converting the continuous first voice signal into a discrete signal. During the sampling process, the first voice signal is recorded at a certain time interval to form a series of discrete sample values. The selection of the sampling rate (the number of samples collected per second) is crucial for retaining the integrity and accuracy of the first voice signal. Commonly used sampling rates include 8kHz, 16kHz, 44.1kHz, etc.

[0073] In the embodiment of the present application, quantization in the preprocessing process is the process of converting the discrete signal values obtained by sampling into a finite number of numerical values (or called quantization levels). During the quantization process, each sample value is mapped to the closest quantization level. The purpose of quantization is to reduce the storage and transmission requirements of data, but it may introduce quantization noise. The number of quantization levels (or called quantization bits) determines the quantization accuracy. For example, 8-bit quantization means that each sample value can be represented by 256 different levels.

[0074] In the embodiments of the present application, filtering in the preprocessing process is a process of performing frequency-domain or time-domain processing on speech signals, aiming to remove noise, enhance signal quality, or extract specific frequency components. Filtering can be achieved through various filters, such as low-pass filters, high-pass filters, band-pass filters, and band-stop filters, etc. In speech recognition, filtering is usually used in the preprocessing stage to improve the signal-to-noise ratio and recognizability of speech signals.

[0075] In the embodiments of the present application, the semantic analysis result includes an intention and an entity. The intention refers to the core purpose of the text information or the task that the user wants to complete. Conducting semantic analysis on the text information to obtain the intention is to determine the user's needs and provide a direction for subsequent task execution; the entity refers to the specific object or concept mentioned in the text information, such as information about time, height, location, etc. Obtaining the entity is to provide specific parameters or context information for the intention.

[0076] In S402 of the embodiments of the present application, semantic analysis is performed on the text information to obtain the entity and intention in the text information.

[0077] Keywords or phrases in the text information can be determined by predefined rules or template matching, thereby determining the intention; the semantic representation of the text can also be learned through a neural network model, classifying the intention of the text information, and determining the intention of the text information.

[0078] In the embodiments of the present application, entities in the text information can be matched by predefined rules or dictionaries; entities in the text information can also be predicted by using a trained neural network model.

[0079] In an example, the text information is segmented into words or phrases, and the part of speech of each word is labeled. The grammatical structure of the text information is analyzed to identify components such as the subject, predicate, and object, identify the semantic roles of each word in the sentence (such as time, location, etc.), extract entities (person names, place names, time, etc.) in the text information, understand the core purpose or user intention of the text information, determine the intention of the text information, analyze the sentiment tendency of the text information (such as positive or negative), and finally deepen the semantic understanding by combining the context information to obtain the final output of the semantic analysis result.

[0080] In S403 of the embodiments of the present application, the predefined task list includes task templates for any executable tasks of the smart device and the terminal, that is, the predefined task list records the task templates executable by the smart device and the terminal respectively, and each task template corresponds to an intention. Therefore, after determining the intention according to semantic analysis, it is necessary to judge whether there is a corresponding template in the task list. If there is a corresponding task template, it means that the requirements sent by the current user through the first voice signal can be met by the smart device or the terminal. Therefore, the specific user requirements are further determined according to the task template and the entity to obtain the target task.

[0081] In the embodiments of the present application, each task template corresponding to an intention is composed of a task name and a part to be filled. For example, if the intention is to start the drying function, the corresponding task template is: (1) + start the drying function, where (1) is used to limit the conditions for starting the drying function. Since only the above task template can be obtained according to the intention, it is also necessary to determine whether there is a description of the conditions for starting the drying function according to the entity obtained through semantic analysis. If the obtained entity does not have content describing the conditions for starting the drying function, then (1) is empty, and the obtained target task is to start the drying function; if the entity has text limiting the specific drying time, such as ten o'clock, then ten o'clock is filled into (1), so as to obtain the target task, start the drying function at ten o'clock.

[0082] In the embodiments of the present application, the recognition model is trained with sample voice signals as training samples and the corresponding text information of the sample voice signals as training labels.

[0083] In the embodiments of the present application, when training the recognition model, it is necessary to select a suitable large language model for training in combination with specific application scenarios and requirements. The key factors for selecting a large language model include: application scenarios, requirements, performance, cost, and deployment conditions. For application scenarios that require real-time response, multi-round dialogue, high-quality text generation, and personalized interaction functions, GPT-4 can be used for training to obtain the recognition model; for application scenarios with requirements for high accuracy, multi-language support, sentiment analysis, and question classification, BERT can be used for training to obtain the recognition model; for application scenarios that require low latency, high accuracy, offline support, and lightweight, OpenAI can be used for training to obtain the recognition model.

[0084] In the embodiments of the present application, in the training stage, a large number of sample voice signals are collected as training samples, and the corresponding text information of the sample voice signals is used as training labels to train the recognition model. During the training process, the recognition model will learn how to extract features from the voice signals and convert them into text information.

[0085] In an embodiment of the present application, during the testing phase of the recognition model, voice signals that did not participate in training are used to test the performance of the recognition model. By comparing the predicted text information output by the model with the true text annotation, the accuracy of the model can be evaluated. The model is tested using a test data set to evaluate the accuracy and generalization ability of the model. According to the test results, the model is optimized and adjusted to improve its performance. The optimization process may include: adjusting model parameters, improving the network, and increasing training data, etc.

[0086] In the above solution, the text information is determined through the recognition model and the first voice signal, so that the user's intention and entity can be determined by semantic analysis of the text information. When it is determined that the user's intention can be satisfied by the terminal or intelligent device, the target task that the user wants to execute will be further determined according to the intention and entity. Through the above operations, the user's intention and needs can be accurately understood, accurate services can be provided, and the user's usage efficiency is improved. And there is no need to manually input text or operate the device to enable the intelligent device to provide corresponding services.

[0087] Based on the above embodiments, as an optional embodiment, after sending a control instruction to the intelligent device, a response message is received; the response message is sent by the intelligent device after executing the target task; according to the response message, the corresponding first reply text is selected from the text library; various response message corresponding reply texts are pre-stored in the text library; the first reply text is converted into a second voice signal, the first reply text is displayed and the second voice signal is played.

[0088] In an embodiment of the present application, after the intelligent device completes the target task, it will return a response message to the terminal according to the execution result. The response message is a pre-set message, and each execution result corresponds to a response message. For example, if the intelligent device determines that the target task is successfully executed, the content represented by the corresponding response message is that the target task is successfully executed; if the intelligent device determines that the task execution fails, the content represented by the corresponding response message is that the target task execution fails; if the intelligent device determines that there is an external obstacle during the task execution (such as an obstacle preventing the clothes hanger from rising to the preset height), the content represented by the corresponding response message is to instruct the user to remove the existing obstacle.

[0089] In an embodiment of the present application, since various response message corresponding reply texts are pre-stored in the text library, after the terminal receives the response message, the corresponding first reply text is directly obtained from the text library, and then the first reply text is converted into a second voice signal. The text information is displayed on the visualization interface of the terminal, and at the same time, the second voice signal is played.

[0090] In the embodiments of the present application, using speech synthesis technology, the generated text information is converted into a voice signal so that the intelligent voice assistant can communicate with the user by voice. The speech synthesis technology pursues natural and realistic voice effects to improve the user's interaction experience.

[0091] In the above solution, since there will be corresponding various types of feedback, that is, response information, for the execution of each target task, therefore, the first reply text corresponding to each response information is pre-stored in the text library in advance. When the response information sent by the intelligent device is received, the corresponding first reply text is directly obtained from the text library, and the reply text is converted into voice information. While playing the voice information, the first reply text is displayed. Visual voice interaction is realized.

[0092] On the basis of the above embodiments, as an alternative embodiment, after determining the target task, if it is determined that the target task is a task executable by the terminal, then execute the target task; determine the execution result, and select the corresponding second reply text from the text library according to the execution result; various execution results corresponding reply texts are pre-stored in the text library; the second reply text is converted into a third voice signal, the second reply text is displayed and the third voice signal is played.

[0093] In the embodiments of the present application, determine the target task. If it is determined that the target task is a task executable by the terminal, then directly execute the determined target task. For example, if the intelligent device is an intelligent clothes dryer and the target task is to provide the reason for the stop of the intelligent clothes dryer, the terminal directly obtains the pre-stored reason for the stop of the clothes dryer, and then broadcasts it to the user in the form of text display and voice broadcast, thereby completing the execution of the target task.

[0094] In the embodiments of the present application, since various execution results corresponding reply texts are pre-stored in the text library, after the terminal determines the execution result according to the execution situation of the target task, the corresponding second reply text is directly obtained from the text library, and then the second reply text is converted into a third voice signal. The text information is displayed on the visual interface of the terminal, and at the same time, the third voice signal is played.

[0095] In the above solution, for the target task executable by the terminal, directly execute the target task, determine the corresponding reply text according to the execution result, and generate the corresponding voice information, so as to realize various scenario functions in a voice control manner and provide a natural and fluent visual interaction.

[0096] On the basis of the above embodiments, as an alternative embodiment, a voice control method is also provided, as Figure 6 shown, and the specific content is as follows: S501, extract the emotional features from the first voice signal; S502. Determine the target emotion corresponding to the emotion feature through a machine learning algorithm. S503. Determine the operation instruction corresponding to the target emotion from the emotion library according to the target emotion. S504. Send a control instruction and an operation instruction to the intelligent device.

[0097] In S501 of the embodiment of the present application, emotion features related to emotion are extracted from the first voice signal. Emotion features usually include the fundamental frequency, tone, speech rate, pitch, energy, phonemes, etc. of the voice. Emotion features can be extracted through digital signal processing techniques, such as short-time energy, short-time zero-crossing rate, linear predictive coding, and Mel-frequency cepstral coefficients, etc.

[0098] In S502 of the embodiment of the present application, a machine learning algorithm is used to analyze the emotion features to understand the emotional state of the sender of the first voice signal. Emotion analysis usually involves classification tasks, and classification tasks are usually trained using supervised learning algorithms, such as support vector machines, random forests, neural networks, and deep learning, etc. The target emotion is usually classified into positive, negative, or neutral emotions. Through the above machine learning algorithm, the target emotion corresponding to the emotion feature can be determined, that is, it is determined whether the emotion corresponding to the first voice signal is positive, negative, or neutral.

[0099] In the embodiment of the present application, operation instructions corresponding to each target emotion are pre-stored in the emotion library, and the operation instructions are used to instruct the intelligent device to perform corresponding operations.

[0100] In S503 of the embodiment of the present application, in order to take into account the user's emotions while providing executable functional services to the user, operation instructions corresponding to each target emotional state are preset and stored in the emotion library. Therefore, after determining the target emotion, the corresponding operation instruction is obtained from the emotion library. For example, when the detected target emotion is a positive emotion, the operation instruction corresponding to the positive emotion obtained from the emotion library is to turn on the pre-set lights; when the detected target emotion is a negative emotion, the operation instruction corresponding to the negative emotion obtained from the emotion library is to play soothing music.

[0101] In S504 of the embodiment of the present application, the operation instruction and the control instruction are sent to the intelligent device together, so that after the intelligent device receives the control instruction and the operation instruction, it executes the operation instruction while executing the target task according to the control instruction.

[0102] In one example, the control instruction instructs the intelligent clothes dryer to raise the drying rack by 80 cm, and the operation instruction instructs the intelligent clothes dryer to turn on the pre-set lights. After the intelligent clothes dryer receives the control instruction and the operation instruction, the intelligent clothes dryer performs the operation of raising the drying rack by 80 cm while turning on the pre-set lights.

[0103] In the above solution, emotional features are extracted from the first voice signal, and the target emotion corresponding to the emotional features is determined through a machine learning algorithm. After determining the target emotion, the operation instruction corresponding to the target emotion is determined from the emotion library according to the target emotion, and then while sending the control instruction to the intelligent device, the operation instruction is sent to the intelligent device together. It realizes the adjustment of the execution mode of the target task according to the emotional state of the user, and realizes the personalized execution of the intelligent device. Improve user satisfaction.

[0104] On the basis of the above embodiments, as an alternative embodiment, the intelligent device is an intelligent clothes dryer, and the target task includes at least one of the following: Adjust the height of the drying rack of the intelligent clothes dryer; Detect whether there is clothing on the drying rack of the intelligent clothes dryer; Start the corresponding function according to the environmental information of the environment where the clothes dryer is located; Judge whether the intelligent clothes dryer is abnormal.

[0105] In the embodiment of the present application, when the intelligent device is an intelligent clothes dryer, the target task may be to adjust the height of the drying rack of the intelligent clothes dryer, and the task layer corresponding to the target task can be obtained by training a linear regression model. The principle of the regression model is based on statistical and machine learning methods, and a mathematical model is established by minimizing the error between the predicted value and the actual value. In linear regression, this error is usually measured by the least squares method, and the goal of the model is to find a set of regression coefficients and intercepts that minimize the sum of the squared errors between the predicted value and the actual value.

[0106] In the embodiment of the present application, during the process of adjusting the height of the drying rack of the intelligent clothes dryer, we use a regression model to predict the height that the drying rack should be adjusted. By selecting appropriate independent variables, collecting sufficient data and training the model, we can obtain a regression model that can accurately predict the height of the drying rack. This model can automatically calculate and adjust the height of the drying rack according to the new input data.

[0107] In the embodiments of the present application, by using a linear regression model to train the task layer corresponding to the above target task, a mathematical relationship between independent variables and dependent variables can be established based on a large number of training samples and training labels, so as to accurately predict the height of the clothes dryer. At the same time, the regression model can process various types of independent variables and dependent variables, and as the number of training samples and training labels increases and the model is continuously optimized, the prediction performance of the task layer is also continuously improved.

[0108] In the embodiments of the present application, when the intelligent device is an intelligent clothes dryer, the target task can also be to detect whether there are clothes on the drying rack of the intelligent clothes dryer. The input of the task layer includes image data of the drying rack area. The image can be a grayscale image or a color image, which is selected according to actual needs.

[0109] In the embodiments of the present application, the task layer corresponding to the above target task is obtained by training a convolutional neural network. The convolutional neural network extracts features, and then maps the features to the final classification result (there are clothes or there are no clothes) through a fully connected layer. During the training process, the task layer will learn how to associate the input first information with the classification result, that is, the model will adjust the weights and bias terms of the convolutional kernels to minimize the error between the prediction result and the actual result. Then, through continuous iterative training, it can finally accurately determine whether there are clothes on the drying rack and the state of the clothes.

[0110] In the embodiments of the present application, the convolutional neural network includes an input layer, a convolutional layer, and a pooling layer. The convolutional layer uses multiple convolutional kernels to perform convolutional operations on the image input to the input layer to extract local features in the image. The convolutional layer usually uses an activation function to increase the non-linearity of the model. The pooling layer is used to perform downsampling on the output of the convolutional layer to reduce the dimension of the data while retaining important features. Common pooling operations include max pooling and average pooling. Multiple convolutional + pooling layers are used to extract deeper features, and usually multiple convolutional layers and pooling layers are stacked. The fully connected layer part includes a flattening layer and a fully connected layer. The flattening layer flattens the output of part of the convolutional neural network into a one-dimensional vector for input to the fully connected layer. The fully connected layer uses multiple neurons to perform weighted summation on the flattened features and outputs the classification result through an activation function (such as softmax). The softmax activation function can convert the output into a probability distribution, which is convenient for judging the state of the clothes. The output layer usually has multiple neurons corresponding to different classification results. For example, when judging whether there are clothes on the drying rack and the state of the clothes, the output layer may have two neurons (there are clothes / there are no clothes on the drying rod of the clothes dryer) or multiple neurons (such as dry clothes, wet clothes, thick clothes, thin clothes, etc.).

[0111] In the above solution, the convolutional neural network can effectively extract the features for judging whether there are clothes on the clothes hanger, and the convolutional neural network has a certain robustness to the translation and rotation transformation of the image. Even if the clothes on the clothes hanger change slightly all the time, it can still accurately identify them. By using the fully connected layer, the features are mapped to the final classification result, realizing the classification task. And by introducing the activation function, the fully connected layer can realize the non-linear combination of features, improving the expression ability of the model.

[0112] In the embodiments of the present application, the target task further includes starting corresponding functions according to the environmental information of the clothes dryer. For example, according to environmental factors such as indoor light and temperature, the photo and drying functions of the clothes dryer are automatically adjusted to provide a comfortable drying environment.

[0113] In the embodiments of the present application, the target task further includes judging whether the intelligent clothes dryer has an abnormality. By detecting the running state of the intelligent clothes dryer in real time, the abnormal conditions of the intelligent clothes dryer (such as being blocked, overweight, motor overheating, etc.) are discovered and reported in time to ensure the normal operation of the intelligent clothes dryer.

[0114] In the embodiments of the present application, the voice control method provided in the embodiments of the present application is deployed to the actual application scenario. In the actual application process, the model will receive the user's first voice signal and convert the first voice signal into text information, and the text information can be used in various application scenarios such as voice assistants, voice search, and voice control.

[0115] In the embodiments of the present application, by introducing the context management mechanism, when the user asks a question related to the previous conversation, the context information can be used to generate a more accurate reply, and multi-round conversations are supported, enabling the user to communicate continuously and deeply with the voice assistant.

[0116] In the embodiments of the present application, real-time communication technologies such as RTC (Real-time communication, real-time audio and video) technology are adopted to realize real-time communication between the user and the terminal. The RTC technology can better adapt to the changes in the user's network conditions and provide better real-time transmission performance.

[0117] In the embodiments of the present application, the network transmission protocol and algorithm can also be optimized to reduce the data transmission delay and packet loss rate, thereby ensuring that the real-time interaction between the user and the terminal is smoother and more natural.

[0118] Based on the above embodiments, as an alternative embodiment, referring to Figure 7 As shown, when there is no task template matching the intention stored in the predefined task list, a voice control method is provided, and the specific content is as follows: S601. Determine the target keywords corresponding to the text information according to the semantic analysis result; S602. Obtain the historical semantic analysis results within a preset time period from the historical conversation library; S603. Obtain the target historical semantic analysis result from the historical semantic analysis results within the preset time period; S604. Determine the target historical semantic analysis result according to the semantic analysis result and the historical semantic analysis result; S605. Generate a third reply text according to the historical semantic analysis result and a predefined reply template; the reply template is used to generate questions instructing the user to answer; S606. Convert the third reply text into a fourth voice signal, display the third reply text, and play the fourth voice signal.

[0119] In the embodiment of the present application, if there is no task template matching the intention stored in the predefined task list, it means that the requirements of the current user sent through the first voice signal cannot be met by the terminal and the intelligent device, or the information provided is not detailed enough, and the terminal cannot recognize the actual requirements of the user according to the first voice signal.

[0120] In S601 of the embodiment of the present application, the target keywords refer to the keywords related to the theme of the text information. According to the semantic analysis result, keyword components such as entities, verbs, and nouns are extracted to determine the target keywords.

[0121] In S602 of the embodiment of the present application, according to the user information, obtain the historical semantic analysis results of the current user within a preset time period from the historical conversation library.

[0122] In S603 of the embodiment of the present application, the target historical semantic analysis result refers to the historical semantic analysis result containing the target keywords. According to the target keywords, filter out the target historical semantic analysis results containing the target keywords from the historical semantic analysis results within the preset time period.

[0123] In S604 of the embodiment of the present application, the target semantic analysis result is obtained by combining the requirement information of the current semantic analysis result and the requirement information of the historical semantic analysis result. Extract the current requirement information and historical requirement information of the user from the semantic analysis result and the historical semantic analysis result respectively, and combine the current requirement information and the historical requirement to obtain the target semantic analysis result.

[0124] In S605 of the embodiment of the present application, in order to ensure effective communication with the user, the terminal needs to generate a third reply text for further inquiring about the user's needs. Therefore, after determining the target semantic analysis result of the user based on the target historical semantic analysis result, the third reply text is generated through a reply template for indicating the question for the user to answer. The reply template is a predefined template, which consists of a part to be filled and a fixed text part. For example, the reply template is as follows: "Do you need to obtain relevant information about (2)?", where (2) belongs to the part to be filled. After determining the target semantic analysis result, (2) is filled according to the target semantic analysis result, so as to obtain the third reply text.

[0125] In S606 of the embodiment of the present application, after generating the third reply text, the third reply text is converted into a fourth voice signal, and the text information is displayed on the visual interface of the terminal, and at the same time, the fourth voice signal is played. Using the speech synthesis technology, the generated third reply text is converted into a fourth voice signal to confirm the specific needs of the user in a timely manner.

[0126] In the embodiment of the present application, after the terminal determines that there is no task template matching the intention stored in the task list, according to the user information, the historical semantic analysis result of the current user within a preset time period is obtained from the historical conversation library, the keywords in the current semantic analysis result are determined, and the target historical semantic analysis result related to the current conversation is retrieved from the historical semantic analysis result within the preset time period, that is, the historical semantic analysis result containing the keywords. The target historical semantic analysis result is combined with the current semantic analysis result to determine the target semantic analysis result. According to the target semantic analysis result and the preset inquiry template, the third reply text is constructed, the third reply text is converted into a fourth voice signal, the third reply text is displayed and the fourth voice signal is played, which solves the problem that the user's intention is not clear and the target needs of the user cannot be completely determined. The third reply text can be generated for further confirmation, so as to improve the efficiency of communicating with the user.

[0127] The embodiment of the present application provides a voice control device, as Figure 8 shown, the voice control device 80 may include: a receiving module 801, an obtaining module 802, and a sending module 803.

[0128] Specifically, the receiving module 801 is configured to receive a first voice signal and determine user information and a target task according to the first voice signal; the user information is information related to the sender of the first voice signal; An acquisition module 802, configured to acquire first information if it is determined that the target task is a task executable by the intelligent device, input the first information, user information, and target task into a pre-trained task model, and obtain a control instruction output by the task model; the first information includes at least one of the status information of the intelligent device and the environmental information of the environment where the intelligent device is located; the control instruction is an instruction matching the first information, user information, and target task. A sending module 803, configured to send the control instruction to the intelligent device so that the intelligent device executes the control instruction. The task model is trained with sample user information, sample tasks, and sample first information as training samples and sample control instructions as training labels, and the sample control instructions are instructions matching the sample user information, sample tasks, and sample first information.

[0129] The device according to the embodiment of the present application can execute the method provided by the embodiment of the present application, and its implementation principle is similar. The actions performed by each module in the device according to the embodiments of the present application correspond to the steps in the method according to the embodiments of the present application. For the detailed function descriptions of each module of the device, reference may specifically be made to the descriptions in the corresponding method shown above, and details are not described herein again.

[0130] The voice control device provided by the embodiment of the present application receives a first voice signal, determines the user information and target task of the voice sender according to the first voice signal. If it is determined that the target task is a task executable by the intelligent device, it acquires first information representing the status information of the intelligent device and the environmental information of the environment where the intelligent device is located, inputs the first information, user information, and target task into a pre-trained task model together, obtains a control instruction output by the task model and matching the first information, user information, and target task, and then sends the control instruction to the intelligent device so that the intelligent device executes the control instruction to complete the target task. The embodiment of the present application controls the intelligent device through the received voice signal, realizes that the user issues control information through the voice signal, and generates personalized control instructions based on the user information, the current status information of the intelligent device, and the environmental information personality, achieving the purpose of instructing the intelligent device to provide various functions for the user in a personalized manner.

[0131] An electronic device (computer device / equipment / system) is provided in the embodiment of the present application, including a memory, a processor, and a computer program stored on the memory. The processor executes the above computer program to implement the steps of the voice control method. Compared with the related art, it can be realized that: through the received voice signal, the intelligent device is controlled accordingly, the user issues control information through the voice signal, and personalized control instructions are generated based on the user information, the current status information of the intelligent device, and the environmental information personality, achieving the purpose of instructing the intelligent device to provide various functions for the user in a personalized manner.

[0132] In an alternative embodiment, an electronic device is provided, such as Figure 9 shown Figure 9 The electronic device 4000 shown includes: a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as connected through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, and the transceiver 4004 may be used for data interaction between the electronic device and other electronic devices, such as data sending and / or data receiving, etc. It should be noted that in practical applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation to the embodiments of the present application.

[0133] The processor 4001 may be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in combination with the disclosure of the present application. The processor 4001 may also be a combination that implements computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0134] The bus 4002 may include a path for transmitting information between the above components. The bus 4002 may be a PCI (Peripheral Component Interconnect, peripheral component interconnect standard) bus or an EISA (Extended Industry Standard Architecture, extended industry standard architecture) bus, etc. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 9 only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0135] The memory 4003 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, which is not limited herein.

[0136] The memory 4003 is used to store the computer program for implementing the embodiments of the present application and is controlled by the processor 4001 to execute. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.

[0137] Among them, the electronic device package may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 9 The illustrated electronic device is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0138] The embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps and corresponding contents shown in the foregoing method embodiments can be implemented. Compared with the prior art, it can achieve: It should be noted that the computer-readable medium described above in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0139] The embodiments of the present application also provide a computer program product, including a computer program, which can implement the steps and corresponding contents of the foregoing method embodiments when executed by a processor. Compared with the prior art, it can be realized that: through the received voice signal, the intelligent device is controlled accordingly, enabling the user to send control information through the voice signal, and based on the user information, the current state information of the intelligent device, and the environmental information personality, personalized control instructions are generated, achieving the purpose of instructing the intelligent device to provide various functions for the user in a personalized manner.

[0140] The terms "first", "second", "third", "fourth", "1", "2", etc. (if any) in the description and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than the illustrated or described order.

[0141] It should be understood that although the flowcharts of the embodiments of the present application indicate each operation step by arrows, the execution order of these steps is not limited to the order indicated by the arrows. Unless otherwise clearly stated in this article, in some implementation scenarios of the embodiments of the present application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage among these sub-steps or stages can also be executed at different times respectively. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present application do not limit this.

[0142] The above are only optional implementation manners of some implementation scenarios of the present application. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of the present application, using other similar implementation means based on the technical idea of the present application also belongs to the protection scope of the embodiments of the present application.

Claims

1. A voice control method, characterized in that: Applied to a terminal, the method comprises: receiving a first voice signal, and determining user information and a target task according to the first voice signal; the user information is information related to a sender of the first voice signal; If it is determined that the target task is a task executable by the smart device, first information is obtained, the first information, the user information and the target task are input into a pre-trained task model, and a control instruction output by the task model is obtained; the first information includes at least one of the state information of the smart device and the environmental information of the environment in which the smart device is located; Sending the control instruction to the smart device so that the smart device executes the control instruction; The task model is trained using sample user information, sample tasks and sample first information as training samples, and sample control instructions as training labels.

2. The method according to claim 1, characterized in that The task model includes a feature extraction layer and at least one task layer, and each task layer has its own corresponding task; The step of inputting the first information, the user information, and the target task into a pre-trained task model to obtain a control instruction output by the task model includes: Inputting the first information and the target task into the feature extraction layer to obtain common features common to all task layers output by the feature extraction layer; Determining a target task layer corresponding to the target task from the at least one task layer; The user information and the general features are input into the target task layer, the target task layer extracts sub-features related to the task corresponding to the task layer from the user information and the general features, and obtains control instructions output by the target task layer based on the obtained sub-features.

3. The method according to claim 1, characterized in that The determining corresponding user information according to the first voice signal includes: Extracting a first voice feature from the first voice signal; the first voice feature is a feature that characterizes the identity of the user; Calculating the similarity between the first voice feature and a second voice feature of any user in a voice database, wherein the second voice features of each user are pre-stored in the voice database; Select the user with the highest similarity as the target user; According to the target user, obtaining user information of the target user from an information database; The information base stores user information of each user in advance.

4. The method according to claim 1, characterized in that: The determining the target task according to the first speech signal comprises: Inputting the first speech signal into a pre-trained recognition model to obtain text information output by the recognition model; Performing semantic analysis on the text information to determine a semantic analysis result; the semantic analysis result includes intent and entity; If a task template matching the intention is stored in the predefined task list, a target task is determined according to the task template and the entity; the task list includes task templates of any executable task of the smart device and the terminal; The recognition model is trained using sample speech signals as training samples and text information corresponding to the sample speech signals as training labels.

5. The method according to claim 1, characterized in that The step of sending the control instruction to the smart device further includes: Receiving response information; the response information is sent by the smart device after executing the target task; According to the response information, a corresponding first reply text is selected from a text library; the text library pre-stores reply texts corresponding to various types of response information; The first reply text is converted into a second voice signal, the first reply text is displayed and the second voice signal is played.

6. The method according to claim 1, characterized in that The determination of the target task also includes: If it is determined that the target task is a task executable by the terminal, executing the target task; Determine the execution result, and select a corresponding second reply text from a text library according to the execution result; the text library pre-stores reply texts corresponding to various execution results; Converting the second reply text into a third voice signal, displaying the second reply text and playing the third voice signal; The receiving of the first voice signal further comprises: Extracting emotional features from the first speech signal; Determine the target emotion corresponding to the emotion feature through a machine learning algorithm; According to the target emotion, determining the operation instruction corresponding to the target emotion from the emotion library; the emotion library pre-stores the operation instruction corresponding to each target emotion; the operation instruction is used to instruct the smart device to perform the corresponding operation; The sending the control instruction to the smart device includes: The control instruction and the operation instruction are sent to the smart device.

7. The method according to claim 4, characterized in that If no task template matching the intention is stored in the predefined task list, the method further includes: Determine the target keyword corresponding to the text information according to the semantic analysis result; the target keyword refers to a keyword related to the subject of the text information; Obtain historical semantic analysis results within a preset time period from the historical conversation library; Obtaining a target historical semantic analysis result from the historical semantic analysis results within the preset time period; the target historical semantic analysis result refers to a historical semantic analysis result containing the target keyword; Determining a target historical semantic analysis result according to the semantic analysis result and the historical semantic analysis result; generating a third reply text according to the historical semantic analysis result and a predefined reply template; the reply template is used to generate a question for instructing the user to answer; The third reply text is converted into a fourth voice signal, the third reply text is displayed and the fourth voice signal is played.

8. The method according to claim 1, characterized in that: If the smart device is a smart clothes drying machine, the target task includes at least one of the following: Adjusting the height of the clothes drying rack of the smart clothes drying machine; Detecting whether there are clothes on the clothes drying rack of the smart clothes drying machine; Activate corresponding functions according to environmental information of the environment in which the clothes drying machine is located; Determine whether an abnormality occurs in the intelligent clothes drying machine.

9. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • Voice interruption processing method and device, electronic equipment and storage medium

    CN120748401A