Distributed voice control method and electronic device
By splitting the voice control model in smart devices into feature extraction by the first terminal and recognition operation by the second terminal, the high load and low efficiency problems caused by frequent model training in existing technologies are solved, and more efficient voice control is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2021-10-22
- Publication Date
- 2026-05-22
AI Technical Summary
In existing technologies, smart devices frequently retrain machine learning models to adapt to new types or manufacturers' devices, resulting in problems such as large development workloads for mobile phone manufacturers, heavy model loads, high processing latency, and low voice control efficiency.
The voice control model is split into a first model and a second model. The first model extracts feature information on the first terminal, and the second model recognizes operation information on the second terminal, thus decoupling the feature extraction and operation recognition processes.
It reduces the computational load on the first terminal, improves the efficiency of voice control, reduces model training and maintenance costs, and enhances the user experience.
Smart Images

Figure CN116030790B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminal technology, and in particular to distributed voice control methods and electronic devices. Background Technology
[0002] With the widespread adoption of smart devices, more and more users are using them in various smart scenarios. Among these smart scenarios is voice control. In voice control scenarios, a single electronic device can be used to control other devices in a distributed voice control system. For example, in... Figure 1 In the scenario shown, the user inputs the voice message "Turn on the TV" into the phone. The phone parses the operation information represented by the voice message (i.e., the user wants to turn on the TV), generates a control signal, and sends the control signal to the TV to control the TV to turn on.
[0003] In some solutions, mobile phones can use machine learning models to parse users' voice information. However, since different devices may come from different manufacturers, when a new type or manufacturer of device establishes a wireless connection with the phone, the phone manufacturer usually needs to retrain the machine learning model so that the model can correctly parse the voice information used to control the new type or manufacturer of device. Therefore, in existing technologies, frequent model retraining leads to a large workload for phone manufacturers, requiring them to continuously retrain and maintain the entire model. Furthermore, the complexity and heavy load of the model running on the phone result in high processing latency and low voice control efficiency. Summary of the Invention
[0004] This application provides a distributed voice control method and electronic device, which can improve the efficiency of voice control.
[0005] To achieve the above objectives, the embodiments of this application provide the following technical solutions:
[0006] The first aspect provides a distributed voice control method, which can be applied to a first terminal or a component (such as a chip system) capable of realizing the functions of the first terminal. The first terminal responds to voice information input by a user, inputs the voice information into a first model, and obtains feature information corresponding to the voice information through the first model. The first model exists in the first terminal. The first terminal sends the feature information to a second terminal so that the second terminal inputs the feature information into a second model, determines the operation information corresponding to the voice information through the second model, and performs a corresponding operation according to the operation information. The second model exists in the second terminal.
[0007] Compared to existing technologies where the first terminal (e.g., a mobile phone) needs to complete the process from voice feature extraction to operation information recognition, resulting in high computational load and low voice control efficiency, the technical solution of this application decouples the feature extraction and operation information recognition processes in voice control scenarios such as smart home devices. For example, the complete model used for voice control can be split into at least a first model and a second model. The first model exists in the first terminal, which can extract the feature information corresponding to the voice information. The second model exists in the second terminal, which can recognize operation information through the second model (e.g., various smart home devices controlled by a mobile phone). Since the first terminal no longer executes all the steps in voice control, such as no longer performing operation information recognition, the computational load is reduced, improving the operating speed of the first terminal and thus increasing the efficiency of voice control.
[0008] In one possible design, the first model is a model trained based on at least one set of first sample data, the first sample data including: first speech information, the feature information of the first speech information being known, and / or...
[0009] The second model is a model trained based on at least one second sample data, which includes: first feature information, and the operation information corresponding to the first feature information is known.
[0010] In one possible design, the first terminal and at least one second terminal are in the same local area network;
[0011] Alternatively, the first terminal and at least one second terminal are in different local area networks.
[0012] In one possible design, the first terminal sends feature information to the second terminal, including: the first terminal broadcasting feature information to the second terminal.
[0013] In one possible design, the feature information corresponding to the speech information includes the spectrogram of the speech information and the phonemes of the spectrogram.
[0014] The second aspect provides a distributed voice control method, the method comprising:
[0015] The second terminal receives feature information corresponding to the voice information from the first terminal; the feature information is obtained by the first terminal inputting the voice information into the first model and through the first model, and the first model exists in the first terminal.
[0016] The second terminal inputs feature information into the second model and determines the operation information corresponding to the voice information through the second model. The second model exists in the second terminal.
[0017] The second terminal performs the corresponding operation based on the operation information.
[0018] In one possible design, the second terminal performs corresponding operations based on the operation information, including:
[0019] If it is determined that the operation information corresponding to the voice information is the same as the operation information matched by the second terminal, then the second terminal executes the target operation according to the operation information corresponding to the voice information; and / or,
[0020] If it is determined that the operation information corresponding to the voice information is not the same as the operation information matched by the second terminal, the second terminal discards the operation information.
[0021] In one possible design, the first model is a model trained based on at least one first sample data, the first sample data including: first speech information, the feature information of the first speech information being known; and / or, the second model is a model trained based on at least one second sample data, the second sample data including: first feature information, the operation information corresponding to the first feature information being known.
[0022] In one possible design, the first terminal and the second terminal are on the same local area network (LAN), or the first terminal and the second terminal are on different local area networks (LANs).
[0023] In one possible design, the feature information corresponding to the speech information includes the spectrogram of the speech information and the phonemes of the spectrogram.
[0024] The third aspect provides a speech recognition method that can be applied to a first terminal or components (such as a chip system) that implement the functions of the first terminal. Taking the implementation of the method on a first terminal as an example, the method includes:
[0025] The first terminal receives first voice information in the first language input by the user;
[0026] The first terminal responds to the first voice information by inputting the first voice information into the first model and obtaining the feature information corresponding to the first voice information through the first model; the first model exists in the first terminal.
[0027] The first terminal sends the feature information to the second terminal, so that the second terminal inputs the feature information into the second model and determines the subtitle information corresponding to the first voice information through the second model. The second model exists in the second terminal.
[0028] In one possible design, the first model is a model trained based on at least one first sample data, the first sample data including: first speech information, the feature information of the first speech information being known; and / or, the second model is a model trained based on at least one second sample data, the second sample data including: first feature information, the operation information corresponding to the first feature information being known.
[0029] In one possible design, the subtitle information is subtitle information in a second language.
[0030] In one possible design, the first language is different from the second language.
[0031] This method can be applied in speech-to-text scenarios, such as remote conferencing. The second terminal may need to generate subtitles from the speech of the speaker using the first terminal and display them on the screen for a clearer understanding of the speaker's speech. Furthermore, if the second terminal has its speech translation function enabled, it can translate the first speech information (e.g., English) into subtitles in the corresponding language (e.g., Chinese) based on the characteristics of the first speech information, thus enabling the user of the second terminal to better understand the meaning of the speaker's speech.
[0032] Furthermore, since the speech-to-text conversion operation is jointly performed by the first terminal and the second terminal, the first terminal does not need to be responsible for converting speech information into corresponding operation information. Therefore, the computational load of the first terminal is reduced, which can improve the running speed of the first terminal and thus improve the efficiency of speech-to-text conversion.
[0033] In one possible design, the second terminal includes a terminal with voice translation enabled.
[0034] In one possible design, the first terminal sends the feature information to the second terminal, including: the first terminal broadcasting the feature information.
[0035] The fourth aspect provides a speech recognition method that can be applied to a second terminal or components (such as a chip system) that implement the functions of the second terminal. Taking the implementation of this method in a second terminal as an example, the method includes:
[0036] The second terminal receives feature information corresponding to the first voice information; the first voice information is voice information in a first language;
[0037] The second terminal inputs the feature information into the second model and determines the subtitle information corresponding to the first voice information through the second model; the second model exists in the second terminal.
[0038] In one possible design, the first model is a model trained based on at least one first sample data, the first sample data including: first speech information, the feature information of the first speech information being known; and / or, the second model is a model trained based on at least one second sample data, the second sample data including: first feature information, the operation information corresponding to the first feature information being known.
[0039] In one possible design, the subtitle information is subtitle information in a second language.
[0040] In one possible design, the first language is different from the second language.
[0041] In one possible design, the method further includes: determining second speech information in a second language corresponding to the first speech information, and playing the second speech information in the second language. The first language is different from the second language. Similar to simultaneous interpretation, in this scheme, the second terminal can translate the first speech information (English speech information) of the speaker using the first terminal into second speech information (Chinese speech information), play the second speech information, and simultaneously display subtitles in the corresponding language (e.g., Chinese subtitles). Alternatively, the second terminal can also play bilingual speech information and display bilingual subtitles. Alternatively, the second terminal can play monolingual speech information and display bilingual subtitles, or vice versa. This application does not limit the technical solution in this regard.
[0042] In one possible design, the second terminal includes a terminal with voice translation enabled.
[0043] In one possible design, the first terminal sends the feature information to the second terminal, including: the second terminal broadcasting the feature information.
[0044] The fifth aspect provides a first terminal, comprising:
[0045] The processing module is used to respond to the user's input voice information, input the voice information into a first model, and obtain the feature information corresponding to the voice information through the first model; the first model exists in the first terminal;
[0046] The communication module is used to send feature information to the second terminal, so that the second terminal inputs the feature information into the second model, determines the operation information corresponding to the voice information through the second model, and performs the corresponding operation according to the operation information. The second model exists in the second terminal.
[0047] In one possible design, the first model is a model trained based on at least one first sample data, the first sample data including: first speech information, the feature information of the first speech information being known; and / or, the second model is a model trained based on at least one second sample data, the second sample data including: first feature information, the operation information corresponding to the first feature information being known.
[0048] In one possible design, the first terminal and at least one second terminal are in the same local area network;
[0049] Alternatively, the first terminal and at least one second terminal are in different local area networks.
[0050] In one possible design, the communication module is used to send feature information to the second terminal, including: the first terminal broadcasting feature information.
[0051] In one possible design, the feature information corresponding to the speech information includes the spectrogram of the speech information and the phonemes of the spectrogram.
[0052] The sixth aspect provides a second terminal, comprising:
[0053] The communication module is used to receive feature information corresponding to voice information from the first terminal; the feature information is obtained by the first terminal inputting voice information into a first model and obtaining it through the first model; the first model exists in the first terminal;
[0054] The processing module is used to input feature information into the second model and determine the operation information corresponding to the voice information through the second model; the second model exists in the second terminal;
[0055] The processing module is used to perform corresponding operations based on the operation information.
[0056] In one possible design, the second terminal performs corresponding operations based on the operation information, including:
[0057] If it is determined that the operation information corresponding to the voice information is the operation information matched by the second terminal, then the second terminal executes the target operation according to the operation information corresponding to the voice information; and / or, if it is determined that the operation information corresponding to the voice information is not the operation information matched by the second terminal, then the second terminal discards the operation information.
[0058] In one possible design, the first model is a model trained based on at least one first sample data, the first sample data including: first speech information, the feature information of the first speech information being known; and / or, the second model is a model trained based on at least one second sample data, the second sample data including: first feature information, the operation information corresponding to the first feature information being known.
[0059] In one possible design, the first terminal and the second terminal are on the same local area network (LAN), or the first terminal and the second terminal are on different local area networks (LANs).
[0060] In one possible design, the feature information corresponding to the speech information includes the spectrogram of the speech information and the phonemes of the spectrogram.
[0061] The seventh aspect provides a first terminal, comprising:
[0062] The input module is used to receive the user's input of the first speech information in the first language;
[0063] The processing module is configured to respond to the first voice information by inputting the first voice information into a first model and obtaining feature information corresponding to the first voice information through the first model; the first model exists in the first terminal.
[0064] A communication module is used to send the feature information to a second terminal, so that the second terminal inputs the feature information into a second model and determines the subtitle information corresponding to the first voice information through the second model; the second model exists in the second terminal.
[0065] In one possible design, the first model is a model trained based on at least one first sample data, the first sample data including: first speech information, the feature information of the first speech information being known; and / or, the second model is a model trained based on at least one second sample data, the second sample data including: first feature information, the operation information corresponding to the first feature information being known.
[0066] In one possible design, the subtitle information is subtitle information in a second language.
[0067] In one possible design, the first language is different from the second language.
[0068] In one possible design, the second terminal includes a terminal with voice translation enabled.
[0069] In one possible design, the communication module is used to send the feature information to the second terminal, including broadcasting the feature information.
[0070] The eighth aspect provides a second terminal, comprising:
[0071] An input module is used to receive feature information corresponding to first voice information; the first voice information is voice information of a first language; the feature information is obtained by the first terminal inputting the first voice information into a first model and through the first model; the first model exists in the first terminal;
[0072] The processing module is used to input the feature information into the second model and determine the subtitle information corresponding to the first voice information through the second model. The second model exists in the second terminal.
[0073] In one possible design, the first model is a model trained based on at least one first sample data, the first sample data including: first speech information, the feature information of the first speech information being known; and / or, the second model is a model trained based on at least one second sample data, the second sample data including: first feature information, the operation information corresponding to the first feature information being known.
[0074] In one possible design, the subtitle information is subtitle information in a second language.
[0075] In one possible design, the first language is different from the second language.
[0076] In one possible design, the processing module is further configured to determine second speech information in a second language corresponding to the first speech information;
[0077] The output module is used to play second speech information in a second language. The first language is different from the second language.
[0078] In one possible design, the second terminal includes a terminal with voice translation enabled.
[0079] In one possible design, the communication module is used to send the feature information to the second terminal, including broadcasting the feature information.
[0080] A ninth aspect provides an electronic device having the function of implementing the distributed voice control method as described in any of the foregoing aspects and any possible implementations thereof. This function can be implemented in hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the foregoing function.
[0081] A tenth aspect provides a computer-readable storage medium including computer instructions that, when executed on an electronic device, cause the electronic device to perform a distributed voice control method as described in any of the foregoing aspects and any possible implementation thereof.
[0082] The eleventh aspect provides a computer program product that, when run on an electronic device, causes the electronic device to execute a distributed voice control method as described in any aspect and any possible implementation thereof.
[0083] The twelfth aspect provides a circuit system including processing circuitry configured to perform a distributed voice control method as described in any of the foregoing aspects and any possible implementation thereof.
[0084] The thirteenth aspect provides a first terminal, comprising: a display screen; one or more processors; one or more memories; the memories storing one or more programs, which, when executed by the processors, cause the first terminal to perform any of the methods described in any of the above aspects.
[0085] The fourteenth aspect provides a second terminal, comprising: a display screen; one or more processors; one or more memories; the memories storing one or more programs, which, when executed by the processors, cause the second terminal to perform the method as designed in any of the above aspects.
[0086] The fifteenth aspect provides a chip system including at least one processor and at least one interface circuit, the at least one interface circuit being used to perform transceiver functions and send instructions to at least one processor, wherein when at least one processor executes instructions, at least one processor executes a distributed voice control method as described in any of the foregoing aspects and any possible implementation thereof. Attached Figure Description
[0087] Figure 1 A flowchart illustrating the voice control method provided in an embodiment of this application;
[0088] Figure 2A , Figure 2B A flowchart illustrating the voice control method provided in an embodiment of this application;
[0089] Figure 3 This is a schematic diagram of the system architecture provided in the embodiments of this application;
[0090] Figure 4 , Figure 5 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application;
[0091] Figures 6-8 A flowchart illustrating the voice control method provided in an embodiment of this application;
[0092] Figure 9 A schematic diagram illustrating the training method of the first model provided in an embodiment of this application;
[0093] Figure 10 A schematic diagram illustrating the training method of the second model provided in an embodiment of this application;
[0094] Figure 11 , Figure 12 A flowchart illustrating the face recognition method provided in this application embodiment;
[0095] Figure 13 A flowchart illustrating the speech information translation method provided in this application embodiment;
[0096] Figure 14 A flowchart illustrating the voice control method provided in an embodiment of this application;
[0097] Figure 15 A schematic diagram of the apparatus provided in the embodiments of this application;
[0098] Figure 16 This is a schematic diagram of a chip system provided in an embodiment of this application. Detailed Implementation
[0099] Figure 2A This paper illustrates an existing speech recognition process, taking a user controlling a TV volume via voice command on a mobile phone as an example. The mobile phone inputs the voice information "turn up the TV volume" into a voice activity detection (VAD) model. The VAD model extracts the human voice from the speech and uses this human voice as input to an automatic speech recognition (ASR) model. The ASR model converts the input sound signal into text and outputs it. The text is then processed by a natural language understanding (NLU) model or regular expression matching to convert it into the corresponding user operation information. Afterward, the mobile phone generates a control signal based on the user operation information (i.e., turning up the TV volume) and sends the control signal to the TV, which then turns up the volume accordingly.
[0100] exist Figure 2AIn the corresponding implementation, if a new type of device (such as a device from a different manufacturer than the phone) establishes a connection with the phone, then, considering factors such as compatibility between the phone and the new device, the phone manufacturer can retrain the NLU model used for speech recognition or update the regular expression matching. The retrained NLU model or regular expression matching can be packaged in the installation package of an application used for voice control (such as an application for managing smart homes), so that users can download the new version of the application to their phones by updating the application, and then use the relevant model to process artificial intelligence tasks (such as speech recognition tasks) through the new application. For example, the phone is currently connected to a TV and a speaker. The phone can control the TV and speaker through a smart home app. The phone detects a new type of device (such as a smart lamp) establishing a connection and reports the detection of the new type of device to the server. After learning from the server that a new type of device has established a connection with the phone, the phone manufacturer retrains the NLU model. After the phone manufacturer has trained the model, it can package the trained model in the installation package of the smart home app and store the updated smart home app on the server. Users can download the updated smart home app to their mobile phones and use the app to control newly added smart lamps on the network. For example, users can control the smart lamps to turn on, off, and adjust their brightness using voice commands.
[0101] From the user's perspective, current technologies require mobile phone manufacturers to frequently train models, meaning users need to frequently update applications, resulting in a poor user experience. From the mobile phone's perspective, current technologies typically require mobile phones to perform tasks including recognizing user input, leading to high phone load, high processing latency, and low efficiency in voice control.
[0102] Figure 2B This paper presents an alternative speech recognition scheme. This scheme uses a spoken language understanding (SLU) model to replace the aforementioned ASR and NLU models (or regular expression matching). The SLU model can directly convert sound signals into user operation information. While this scheme can directly convert sound signals into user operation information, it still requires retraining the SLU model when a new type of device is detected connecting to the phone, resulting in high maintenance costs. Secondly, as the types and number of devices connecting to the phone increase, the SLU model needs to recognize more and more operation information, requiring a complex model structure, which slows down the phone's operation. Furthermore, the SLU model requires accurate input of voice commands, making it prone to misrecognition during casual conversation.
[0103] The above Figure 2A , Figure 2BThe existing technical solutions require the mobile phone to complete numerous tasks, including recognizing operation information, resulting in a high load on the phone. Furthermore, each time a new type of device is detected and connects to the phone, the phone manufacturer needs to redevelop and train a new neural network to match the new device type. Therefore, existing voice recognition solutions place a high load on the mobile phone and have high processing latency, leading to low efficiency in voice control.
[0104] To improve the efficiency of voice control, this application provides a voice recognition method. This method is applicable to systems requiring voice control. Figure 3 The diagram shown is an example of a system architecture provided in an embodiment of this application. The system includes one or more electronic devices, such as electronic device 100 and electronic device 200 (e.g., smart home devices 1-3).
[0105] In this application, electronic devices can establish connections with each other. Optionally, the methods for establishing connections between devices include, but are not limited to, one or more of the following: establishing a communication connection by scanning a QR code or barcode; establishing a connection through communication protocols such as Wireless Fidelity (Wi-Fi) and Bluetooth; or establishing a connection through a near-field communication service. After establishing a communication connection, the devices can transmit data and / or signaling. This application does not limit the methods for establishing connections between electronic devices.
[0106] In some scenarios, a single device can control other connected devices via voice. Taking voice control of smart home devices via a mobile phone as an example, a user inputs the voice message "Turn up the TV volume" into mobile phone 100. Mobile phone 100 extracts the feature information corresponding to this voice message and sends it to the connected smart home devices 1-3. Smart home devices 1-3 process the feature information to obtain the corresponding operation information and determine whether a response is needed based on this information. It should be understood that operation information includes, but is not limited to, operation commands and control commands. Optionally, operation information may also include classification results obtained by the smart home devices based on the feature information, such as classifying different operation commands. The smart home devices can then perform corresponding operations based on the operation information. Different types of operation information (such as different control commands) are used to control the smart home devices to perform different operations.
[0107] Specifically, after receiving the feature information, smart home device 3 (TV) processes it to determine that the corresponding operation information for the voice message is "want to turn up the TV volume," and executes the operation corresponding to this operation information, i.e., turning up the volume. After receiving the feature information from mobile phone 100, smart home device 1 (table lamp) processes it to obtain the corresponding operation information for the voice message, and determines not to execute the corresponding operation based on the operation information. Optionally, the table lamp discards the operation information. Similarly, smart home device 2 (air conditioner) also does not execute the corresponding operation. In this process, the step of recognizing the operation information (operation command) is completed by each smart home device, without needing to be completed in the mobile phone, thereby reducing the computational load on the mobile phone and improving the efficiency of the smart voice control process.
[0108] Optionally, the mobile phone can extract the feature information corresponding to the voice information by inputting the voice information into a first model, and the first model outputting the feature information corresponding to the voice information. The first model is used to convert the voice information into the corresponding feature information.
[0109] Optionally, smart home devices process the feature information from the mobile phone to obtain operation information (such as control commands) corresponding to the voice information. This can be achieved by the smart home device inputting the feature information from the mobile phone into a second model, which then outputs the operation information corresponding to the voice information. The second model is used to convert the feature information into corresponding operation information. The first model, the second model, and the feature information will be described in detail below.
[0110] In the embodiments of this application, the above-mentioned electronic device may also be referred to as a terminal.
[0111] Optionally, the system also includes one or more servers 300. The servers can establish connections with electronic devices. In some embodiments, electronic devices can connect to each other via the servers. For example, in... Figure 1 In the system shown, mobile phone 100 can remotely control smart home devices through server 300.
[0112] In some embodiments, the first model and the second model can be trained by the server 300. After the server 300 has trained the first model and the second model, it can distribute the trained first model and the second model to various terminals. In other embodiments, the first model and the second model can be trained by the terminal, such as by a mobile phone.
[0113] Optionally, the first model and the second model can be models obtained based on any algorithm, such as models based on neural networks, which can be one or more combinations of Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), Deep Neural Networks (DNN), Multi-Layer Perceptron (MLP), and Gradient Boosting Decison Tree (GBDT).
[0114] For example, the aforementioned electronic devices 100 and 200 can be mobile phones, tablets, personal computers (PCs), personal digital assistants (PDAs), smartwatches, netbooks, wearable electronic devices, augmented reality (AR) devices, virtual reality (VR) devices, in-vehicle devices, smart cars, smart speakers, robots, headphones, cameras, and other devices that can be used for voice control or be controlled by voice. This application does not impose any special restrictions on the specific form of the electronic devices 100 and 200.
[0115] The terms "first" and "second" in the specification and drawings of this application are used to distinguish different objects or to distinguish different treatments of the same object. The terms "first" and "second" can distinguish identical or similar items with essentially the same function and effect. For example, "first device" and "second device" are only used to distinguish different devices and do not limit their order. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and that "first" and "second" do not necessarily imply differences. "At least one" refers to one or more, and "more" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can be represented as: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single or multiple.
[0116] Furthermore, the terms "comprising" and "having," and any variations thereof, used in the description of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the steps or units listed, but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to such process, method, product, or apparatus.
[0117] It should be noted that in the embodiments of this application, the words "exemplary" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of the words "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0118] In the specification and drawings of this application, the terms "of", "corresponding", and "corresponding" are sometimes used interchangeably. It should be noted that when the distinction is not emphasized, they have the same meaning.
[0119] Taking electronic device 100 as a mobile phone as an example, Figure 4 A schematic diagram of the structure of the electronic device 100 is shown.
[0120] Electronic device 100 may include processor 110, external memory interface 120, internal memory 121, universal serial bus (USB) interface 130, charging management module 140, power management module 141, battery 142, antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, sensor module 180, button 190, motor 191, indicator 192, camera 193, display screen 194, and subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0121] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0122] Processor 110 may include one or more processing units, such as application processors (APs), modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). These different processing units may be independent devices or integrated into one or more processors.
[0123] The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of instruction fetching and execution.
[0124] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0125] In some embodiments, processor 110 may include one or more interfaces.
[0126] In some embodiments of this application, the process by which electronic device 100 processes voice information to obtain feature information, and the process by which electronic device 200 processes feature information from electronic device 100 to obtain operation information corresponding to the voice information, may involve some or all of the data processing in the processor 110 of electronic device 100. Electronic device 100 is also referred to as a first terminal, and electronic device 200 is also referred to as a second terminal.
[0127] It is understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are merely illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0128] The charging management module 140 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 receives charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 receives wireless charging input via the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also supply power to the electronic device via the power management module 141.
[0129] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to power the processor 110, internal memory 121, display 194, camera 193, and wireless communication module 160, etc.
[0130] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.
[0131] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with tuning switches.
[0132] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G / 6G, applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.
[0133] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through an audio device (not limited to speaker 170A, receiver 170B, etc.) or displays images or videos through the display screen 194. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 110 and may be housed in the same device as the mobile communication module 150 or other functional modules.
[0134] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.
[0135] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, so that electronic device 100 can communicate with networks and other devices through wireless communication technology.
[0136] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0137] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. In some embodiments, the electronic device 100 may include one or N display screens 194, where N is a positive integer greater than 1.
[0138] Electronic device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.
[0139] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.
[0140] Camera 193 is used to capture still images or videos. An object passes through the lens to generate an optical image that is projected onto a photosensitive element. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP (Internet Service Provider) for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP (Digital Signal Processor) for processing. The DSP converts the digital image signal into image signals in standard formats such as RGB and YUV. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.
[0141] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when electronic device 100 selects a frequency, the DSP can perform Fourier transforms on the frequency energy.
[0142] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. Thus, electronic device 100 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0143] An NPU (Neural Processing Unit) is a computational processor for neural networks (NNs). By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.
[0144] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to perform data storage functions.
[0145] Internal memory 121 can be used to store computer executable program code, which includes instructions. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of electronic device 100 (such as audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. Processor 110 executes various functional applications and data processing of electronic device 100 by running instructions stored in internal memory 121 and / or instructions stored in memory located in the processor.
[0146] Electronic device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.
[0147] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.
[0148] The speaker 170A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or make hands-free calls through the speaker 170A.
[0149] The receiver 170B, also known as the "earpiece," is used to convert audio electrical signals into sound signals. When the electronic device 100 answers a telephone call or voice message, the receiver 170B can be brought close to the ear to listen to the voice.
[0150] Microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 170C, inputting the sound signal into microphone 170C. Electronic device 100 may have at least one microphone 170C. In some embodiments, electronic device 100 may have two microphones 170C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, electronic device 100 may also have three, four, or more microphones 170C, which can collect sound signals, reduce noise, identify the sound source, and perform directional recording, etc.
[0151] The 170D headphone jack is used to connect wired headphones. The 170D headphone jack can be a USB 130 interface or a 3.5mm Open Mobile Terminal Platform (OMTP) standard interface, a CTIA (Cellular Telecommunications Industry Association of the USA) standard interface.
[0152] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch-sensitive buttons. Electronic device 100 can receive button input and generate key signal inputs related to user settings and function control of electronic device 100.
[0153] Motor 191 can generate vibration alerts.
[0154] Indicator 192 can be an indicator light, used to indicate charging status, power changes, or to indicate messages, missed calls, notifications, etc.
[0155] The SIM card interface 195 is used to connect the SIM card.
[0156] For example, the above description uses electronic device 100 as an example to illustrate the structure of the electronic device in this application embodiment, but it does not constitute a limitation on the structure or form of the electronic device. This application embodiment does not limit the structure or form of the electronic device. For example, Figure 5 Another exemplary structure of an electronic device is shown. For example... Figure 5 As shown, the electronic device includes: a processor 501, a memory 502, and a transceiver 503. The implementations of the processor 501 and memory 502 can be found in the implementation of the processor and memory of the electronic device 100. The transceiver 503 is used for interaction between the electronic device and other devices (such as electronic device 100). The transceiver 503 can be a device based on protocols such as Wi-Fi, Bluetooth, or other communication protocols.
[0157] Optionally, the server architecture can be found in [reference needed]. Figure 5 The structure shown will not be described in detail here.
[0158] In other embodiments of this application, the electronic device or server may include more or fewer components than illustrated, or combine some components, or split some components, or replace some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0159] The technical solutions involved in the following embodiments can all be implemented in situations such as... Figure 4 , Figure 5 Implemented in the device with the structure shown.
[0160] For example, taking a smart home scenario as an example, such as Figure 6 The mobile phone includes a first model, and each smart home device includes a second model. The first model is trained and deployed on the mobile phone by the mobile phone manufacturer. The first model can be used to obtain multi-dimensional feature information corresponding to voice information. In this embodiment, the weights of the first model are usually fixed, eliminating the need for frequent updates.
[0161] The second model is trained independently by each smart home device manufacturer. This second model can be used to convert multi-dimensional feature information corresponding to voice information into corresponding operational information. In this embodiment, the smart home device manufacturers can update the second model according to actual usage needs. That is, when new devices are added later, typically only the manufacturer of the new device needs to retrain the second model used to identify operational information (such as classifying control commands) for that new device. Mobile phone manufacturers do not need to frequently update the first model used to identify feature information, thus reducing model training and maintenance costs for manufacturers such as mobile phone manufacturers. Furthermore, since the second model only involves identifying operational information of a specific device (such as classifying control commands), the model is small and easy to train and update.
[0162] Optionally, updating the second model includes updating the weights of the second model.
[0163] It should be noted that since different smart home devices may come from different manufacturers, and each manufacturer may use different algorithms to train the second model, the second model on different smart home devices may be different.
[0164] exist Figure 6In the scenario shown, the user inputs the voice message "turn up the TV volume" into the phone. After detecting the user's voice input, the phone can input the voice information into a first model, which then outputs the feature information of the voice information. Optionally, the feature information of the voice information can be output in the form of, but not limited to, a feature matrix. After obtaining the feature information, the phone can send it to various smart home devices connected to the phone (such as…). Figure 6 (The table lamp, air conditioner, and television are shown).
[0165] After receiving feature information from a mobile phone, smart home devices can input this feature information into a second model, which then outputs the corresponding operation information based on the voice message. For example... Figure 6 As shown, the television can recognize the operation information (the classification result of the control command) as "turn up the volume" based on the second model. Therefore, the television can execute the corresponding operation, i.e., adjust the volume, based on the recognized operation information. The desk lamp, however, cannot recognize the operation information, or the operation information output by the desk lamp through the second model does not match its own, or the operation information output by the desk lamp through the second model does not match its own executable operation information (control command). In this case, the desk lamp can determine from the operation information (such as the control command) output by the second model that the user's voice information is not for controlling itself, and the desk lamp does not need to respond to the user's voice information. Similarly, the air conditioner does not need to respond to the user's voice information.
[0166] In contrast to existing technologies where the mobile phone must complete the process from feature information extraction to operation information recognition, resulting in high computational load and low voice control efficiency, the aforementioned smart home device voice control scenario decouples the feature information extraction and operation information recognition processes. The feature information extraction process is performed by the mobile phone, while the operation information recognition is performed by each smart home device. Compared to existing technologies, the mobile phone no longer performs the operation information recognition operation, reducing computational load, improving mobile phone speed, and thus increasing voice control efficiency.
[0167] The following describes the specific interactions between devices during voice control in embodiments of this application. Figure 7 As shown, taking the example of a user controlling the volume of a TV via voice control on a mobile phone, the voice control method provided in this application includes:
[0168] S101: The phone detected user-input voice information.
[0169] For example, the user input voice message is "turn up the TV volume".
[0170] S102, The mobile phone converts voice information into feature information.
[0171] Typically, voice information is an analog signal. Mobile phones need to convert it into a digital signal through an encoding model and extract feature information. Subsequently, other devices can identify the operation information corresponding to the voice information based on the extracted feature information.
[0172] Feature information refers to the distinguishable components obtained from speech information. These distinguishable speech components can accurately describe the difference between a speech segment and other speech segments. Optionally, the distinguishable components in speech information include, but are not limited to, spectra and phonemes of the spectra. Phonemes of the spectra include, but are not limited to, formants in the spectra. The feature information in this application is not limited to the types listed above. Due to space limitations, this application will not exhaustively list all feature information. Any information in speech that can play a distinguishing role can be called feature information.
[0173] As one possible implementation, the mobile phone includes a first model. The mobile phone can input voice information into the first model, and the first model can calculate and output feature information corresponding to the voice information. The training method of the first model can be found in the following embodiments.
[0174] In this embodiment, the first model can be implemented as an encoding model (also called an encoding module, an encoding neural network, or other names), and the name does not constitute a limitation on the encoding model. The encoding model can be regarded as a functional module on a mobile phone, which is used to convert the information corresponding to the voice information into the feature information of the voice.
[0175] Alternatively, the first model may integrate other models besides the encoding model, such as the VAD model. Optionally, the first model may also integrate other functions; however, this application embodiment does not limit whether the first model integrates other functions or the specific types of other functions.
[0176] Similarly, the second model in this application embodiment can be implemented as a decoding model (also called a decoding module, decoding neural network, or other names). The decoding model can also be regarded as a functional module in a device (such as a television) that is used to convert the feature information of speech into corresponding operation information.
[0177] Optionally, in addition to the decoding model, the second model may also integrate other models or modules. This application embodiment does not limit whether the second model integrates other functions or the specific types of other functions.
[0178] In this embodiment, the first model may also be referred to as the first model file, or other names. The second model may also be referred to as the second model file, or other names. The names do not constitute a limitation on the first model and the second model.
[0179] S103, Mobile phone broadcast feature information.
[0180] Correspondingly, various devices connected to the mobile phone, such as televisions, receive feature information from the mobile phone.
[0181] In this embodiment, the mobile phone does not perform the operation information recognition operation; instead, the operation information recognition is completed by the various devices controlled by the mobile phone. Therefore, the mobile phone does not know which device the user's voice information is used to control. Thus, the mobile phone needs to broadcast feature information to all connected devices. Each of the other devices then identifies the operation information corresponding to the feature information and determines whether the user's voice information is used to control itself to perform an operation. If so, it responds to the user's voice information and executes the operation corresponding to the voice information; otherwise, it does not respond to the user's voice information and does not execute the operation corresponding to the voice information.
[0182] S104. The television converts the feature information into corresponding operation information.
[0183] As one possible implementation, the television includes a second model. After receiving voice feature information (such as a voice feature matrix) from the mobile phone, the television inputs the voice feature information into the second model, which then determines and outputs the operation information corresponding to the voice information. For example, in... Figure 6 In the scenario shown, the TV inputs the voice feature information (such as the feature matrix) into the second model, and the second model calculates and determines the operation information corresponding to the feature information as "turn up the TV volume".
[0184] S105. The television responds to the operation information and executes the operation corresponding to the operation information.
[0185] For example, as Figure 6 In the scenario shown, once the TV recognizes that the user's voice message corresponds to the operation message "turn up the TV volume", and this operation message is a matching operation message for the TV, it can respond to the operation message and execute the target operation corresponding to the operation message, namely, turn up the volume.
[0186] Next, we will explain the interaction between devices in the voice control method by combining the internal functional modules of the device. For example... Figure 8 As shown, the voice control method in this application embodiment includes:
[0187] S201. The mobile phone detects voice information and inputs the voice information into the VAD model.
[0188] For example, a user inputs the voice message "turn up the TV volume" into the phone. After the phone detects the voice message, it inputs the corresponding voice information into the VAD model.
[0189] The S202 and VAD models detect human voice information in the speech information and input the human voice information in the speech information into the encoding model of the mobile phone.
[0190] Considering that when a user inputs voice information, the phone may also collect other sounds from the environment at the same time, in order to reduce the amount of data processing in subsequent calculations and avoid interference from environmental noise, the phone can use a VAD model to identify human voice information and non-human voice information (noise) in the collected voice information. The VAD model can be any type of model capable of performing speech classification tasks.
[0191] Optionally, the acquired raw speech information can be divided into multiple segments (frames), such as 20ms or 25ms frames, and the speech information can be input into the VAD model, which outputs the classification results of the speech information. Optionally, the VAD model outputs the classification results of each frame as either human voice or non-human voice, and uses the speech belonging to human voice as the input of the subsequent coding model.
[0192] The training process of the VAD model involved in the embodiments of this application can be found in the prior art, and will not be repeated here.
[0193] In this embodiment of the application, the VAD model can also be regarded as a functional module on the mobile phone, which has the function of recognizing human voices and non-human voices.
[0194] S203. The coding model outputs the feature information corresponding to the speech information based on the human voice information in the speech information.
[0195] Optionally, the coding model divides the human voice information into multiple frames. For each frame, the coding model extracts feature information of the human voice information according to certain rules (such as, but not limited to, the mel frequency cepstrum coefficient (MFCC) rule). Optionally, the coding model can convert the extracted feature information into a feature vector.
[0196] For example, a method for extracting feature information using an encoding model is given. First, the speech information is preprocessed. Preprocessing includes, but is not limited to, dividing the speech information into multiple frames. Then, for each frame, the following operation is performed:
[0197] The spectrum corresponding to each frame is obtained through a Fast Fourier Transform (FFT), and then processed using a Mel filter bank to obtain the Mel spectrum corresponding to that frame. This transforms the linear natural spectrum into a Mel spectrum that reflects the characteristics of human hearing. Next, cepstral analysis is performed on the Mel spectrum corresponding to that frame to obtain the corresponding MFCC, which can be used as feature information corresponding to the speech information of that frame.
[0198] After obtaining the feature information of each frame of speech, the feature information of each frame can be combined to obtain the feature information (such as feature vector) corresponding to the speech information.
[0199] It should be noted that there are other methods for the encoding model to extract feature information, and it is not limited to the methods listed above.
[0200] S204. The mobile phone's communication module obtains the feature information corresponding to the voice information.
[0201] Optionally, the communication module enables the mobile phone to communicate with other electronic devices. For example, the communication module can connect to a network via wireless or wired communication to communicate with other personal terminals or network servers. Wireless communication can employ at least one of the following cellular communication protocols: 5G, Long Term Evolution (LTE), LTE-A Advanced, Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Universal Mobile Telecommunications System (UMTS), Wireless Broadband (WiBro), or Global System for Mobile Communications (GSM). Wireless communication may include, for example, short-range communication. Short-range communication may include at least one of Wi-Fi, Bluetooth, Near Field Communication (NFC), Magnetic Stripe Transmission (MST), or GNSS.
[0202] As one possible implementation, the processing module (such as the processor) in the mobile phone can obtain the output result of the above encoding module, that is, obtain the feature information corresponding to the voice information, and send the feature information corresponding to the voice information to the communication module, so that the communication module of the mobile phone can execute the following step S205.
[0203] S205. The mobile phone's communication module broadcasts the characteristic information corresponding to the voice information.
[0204] Correspondingly, the television's communication module receives voice feature information from the mobile phone.
[0205] S206. The television decoding model obtains the feature information corresponding to the voice information.
[0206] As one possible implementation, after the television's communication module receives the feature information corresponding to the voice information from the mobile phone, it sends the feature information to the television's processing module, and the processing module inputs the feature information into the decoding model.
[0207] S207. The television decoding model outputs the operation information corresponding to the feature information based on the feature information corresponding to the voice information.
[0208] Optionally, the decoding model can be a model used to perform classification tasks, whose output is the operation information corresponding to the speech information.
[0209] It should be noted that the decoding model in this embodiment differs from the decoder in traditional ASR. The decoder in traditional ASR converts the feature information corresponding to the speech information into text, which is then converted into corresponding operation information by subsequent functional modules. The decoding model in this embodiment, however, can convert and classify the feature information corresponding to the speech information into corresponding operation information. Therefore, the decoding model in this embodiment has higher conversion efficiency.
[0210] S208. Determine whether the operation information output by the decoding model is the operation information for television matching. If yes, proceed to step S209; otherwise, proceed to S210.
[0211] S209. Respond to the operation information and execute the operation corresponding to the operation information.
[0212] For example, as Figure 6 In the scenario shown, after the TV receives the feature information (such as the feature matrix) corresponding to the voice information from the mobile phone, it outputs operation information (such as the control command) "turn up the TV volume" through the second model (such as the decoding model), and performs the operation corresponding to the operation information, that is, turn up the volume.
[0213] S210. Do not respond to this operation information, and do not execute the operation corresponding to this operation information.
[0214] For example, as Figure 6 In the scenario shown, after the air conditioner receives the feature information corresponding to the voice message from the mobile phone, it outputs the corresponding operation information, "others" type operation information, through a second model (such as a decoding model). This operation information (such as a control command) indicates that the user's voice message was not intended to control the air conditioner. Therefore, based on this operation information, the air conditioner does not perform the corresponding operation. Similarly, after the desk lamp receives the feature information corresponding to the voice message, it outputs the corresponding operation information based on the feature information and determines that no operation needs to be performed.
[0215] By including an encoding neural network on the control device (such as a mobile phone) and a decoding neural network on the controlled device (such as a home appliance), the training of the decoding neural network can be delegated to various third-party manufacturers. Different manufacturers can train their own decoding neural networks. On the one hand, this eliminates the need for frequent training of the encoding neural network, significantly reducing development costs for new devices. On the other hand, since the mobile phone only performs feature extraction in speech recognition and no longer performs operation information recognition, the computational load and power consumption of the mobile phone can be reduced, increasing processing speed and thus reducing latency in the speech recognition process.
[0216] The training methods for the first and second models described above are as follows. The first model is trained based on at least one set of first sample data, which includes first speech information, the feature information of which is known. The second model is trained based on at least one set of second sample data, which includes first feature information, the operation information corresponding to which is known.
[0217] Figure 9 An example of a training method for the first model is given. For example... Figure 9 As shown in (1), the model for recognizing operational information is first trained by providing N (N is a positive integer) training samples. The training samples include speech information for which the operational information is known (i.e., the first speech information). Multiple types of speech data can be used to ensure a sufficiently rich corpus and improve recognition accuracy. Optionally, the training samples also include labels for the speech data, used to represent the operational information corresponding to the speech information. Training multiple samples yields a model capable of extracting feature information and recognizing the operational information corresponding to the speech information. This model can output the operational information corresponding to the speech information.
[0218] like Figure 9 In the training model scenario described in (1), the trained model includes 32 layers of neurons. Layers L1-L16 are used to extract feature information corresponding to the speech information, and layers L17-L32 are used to identify operation information corresponding to the speech information. For a neuron in a certain layer, it can connect to one or more neurons in the next layer, and output corresponding signals through these connections. For example... Figure 9 (1) shows the weights corresponding to the connections between some neurons in the trained model. For example, the connection between the first neuron in the L1 layer and the first neuron in the L2 layer corresponds to weight w11, the connection between the first neuron in the L1 layer and the second neuron in the L2 layer corresponds to weight w12, and so on.
[0219] Optionally, to improve the model's recognition accuracy, the model can be evaluated and tested. When the model's recognition rate reaches a certain threshold, it indicates that the model has been trained well. If the model's recognition rate is low, training can continue until the model's recognition accuracy reaches a certain threshold.
[0220] Optionally, the model training process can be carried out on the device side (such as a mobile phone) or the cloud side (such as a server). Training can be offline or online. This application does not limit the specific training method of the model.
[0221] like Figure 9As shown in (2), after training a complete model for extracting feature information and recognizing operational information (such as classifying control commands), the parts corresponding to layers L17-L32 used for recognizing operational information (such as recognizing control commands) are removed from the complete model to obtain the model for extracting feature information. Figure 9 As shown in (2), the first model for extracting feature information includes 16 layers, L1-L16. After inputting speech data (also known as speech information) into the first model, the first model can output the feature vector (also known as feature information) corresponding to the speech data.
[0222] In summary, taking the first model as the encoder and the second model as the decoder as an example, the encoder in this embodiment is the encoder part of the trained encoder-decoder model, which is equivalent to extracting the encoder part of an encoder-decoder model to form the first model.
[0223] like Figure 10 An example of training and using the second model is given. Figure 10 As shown in (1), the first model is obtained before training the second model. As one possible implementation, if the first model used to extract feature information corresponding to speech information is trained by a mobile phone, the mobile phone can upload the first model to a server. Subsequently, other devices can obtain the first model from the server and train the second model based on the first model. Alternatively, the device can obtain the first model from the mobile phone through other means; the embodiments of this application do not limit the specific way the device obtains the first model.
[0224] One possible implementation is to train the second model by using the output of the first model as its input, forming a neural network for training. In this network, the input of the first model serves as the input to the entire neural network, and the output of the second model serves as its output. During training, the weights of the first model remain unchanged. For example, ... Figure 10 (1) The output of the first model (i.e., the feature information corresponding to the speech information) can be used as a training sample, and the second model can be trained based on the training sample. The trained second model has the function of outputting operation information based on the input feature information. Taking the television recognizing the operation information corresponding to the speech information through the second model as an example, such as Figure 10 As shown in (2), the TV can input the feature information (such as the feature vector received from the mobile phone) corresponding to the voice information whose operation information is unknown into the second model, and then the second model outputs the operation information (such as turning up the volume of the TV) corresponding to the voice information.
[0225] In other embodiments, the device can also train a second model independently, that is, train a second model without needing to obtain the first model. In this implementation, the training samples are also feature vectors corresponding to speech information (an example of the first feature information), and the second model is obtained by training the training samples.
[0226] This method is not limited to voice control scenarios; it can be applied to other distributed task processing scenarios as well. Examples of such scenarios include, but are not limited to, remote conferencing scenarios (including but not limited to real-time translation scenarios) and facial recognition verification scenarios.
[0227] In facial recognition verification scenarios, taking facial recognition via a mobile phone as an example, for instance, such as... Figure 11 As shown, a face recognition model can be divided into at least a first model and a second model. The phone's camera module (e.g., a webcam) includes the first model. The first model is used to extract facial feature information from a face image. The phone's processing module includes the second model. The second model is used to output the face recognition result based on the facial feature information.
[0228] like Figure 12 An exemplary flow of the method of this application embodiment in a face recognition scenario is shown, the flow including the following steps:
[0229] S301, The camera module captures the facial image input by the user.
[0230] S302, The camera module inputs the face image into the first model, and the first model outputs the feature information of the face image.
[0231] S303, The camera module transmits the feature information of the face image to the processing module.
[0232] S304. The processing module inputs the feature information of the face image into the second model, and the second model outputs the face recognition result.
[0233] S305. The processing module determines whether the face is a valid face based on the face recognition result. If yes, proceed to S306; otherwise, proceed to S307.
[0234] S306. Perform the operation corresponding to the facial information.
[0235] For example, in a payment scenario, when a user inputs a facial image, the phone determines that the face is a valid face using the first model in the camera module and the second model in the processing module, and then executes the payment operation. In a screen unlock scenario, when a user inputs a facial image, the phone unlocks the screen using the first model in the camera module and the second model in the processing module.
[0236] S307. Do not perform operations corresponding to facial information.
[0237] The above explanation uses the camera as an example on a mobile phone. In other scenarios, the camera containing the first model can also be located in a module independent of the mobile phone, with the second model included in the mobile phone. In this way, the external camera of the mobile phone can work together with the mobile phone to complete the face recognition process. Furthermore, since the model has been split into the first model and the second model, the efficiency of face recognition can be improved.
[0238] Similarly, in other distributed intelligent scenarios, a device can split one or more models used to perform one or more tasks into multiple sub-models, and deploy these sub-models in multiple modules of the device, thereby distributing the model operation load of a single module across these multiple modules. This application does not limit the specific method of model splitting, nor does it limit which modules the model, after being split into multiple sub-models, is specifically distributed and deployed in.
[0239] In real-time translation scenarios during remote conferencing, the existing model can be split into at least a first model and a second model. The first model runs on the speaker's device, and the second model runs on the receiver's device. For example... Figure 13 An exemplary flow of the method of this application embodiment in a remote conference translation scenario is shown, the flow including the following steps:
[0240] S401, The audio acquisition module of mobile phone A acquires the first language information of the source language (i.e., the first language) and inputs the first voice information into the first model of mobile phone A.
[0241] Optionally, the audio acquisition module includes, but is not limited to, a microphone. Taking English to Chinese translation as an example, the first speech information of the source language is the English phrase "this meeting is". The audio acquisition module of mobile phone A acquires the speaker's English speech information and inputs the English speech information into the first model.
[0242] S402, The first model extracts the feature information of the first speech information.
[0243] For example, feature information of English speech information is extracted, namely the feature information corresponding to the English speech "this meeting is".
[0244] S403, the communication module of mobile phone A obtains the feature information of the first voice information.
[0245] As one possible implementation, the communication module of mobile phone A obtains the feature information of the first voice information from the first model, or the communication module of mobile phone A obtains the feature information of the first voice information from the processing module.
[0246] S404. The communication module of mobile phone A sends the feature information of the first voice message to the communication module of mobile phone B.
[0247] S405. The second model of mobile phone B obtains the feature information of the first voice message.
[0248] As a possible implementation, the second model obtains the feature information of the first voice message from the communication module. Alternatively, the processing module inputs the feature information into the second model, that is, the second model obtains the feature information of the first voice message from the processing module.
[0249] S406. The second model determines the subtitle information and / or the second voice message in the target language (the second language) corresponding to the first voice message according to the feature information of the first voice message.
[0250] Optionally, the first language is different from or the same as the second language.
[0251] Exemplarily, in the scenario where mobile phone B enables the voice translation (such as English to Chinese) function, after receiving the feature information of the first voice message, mobile phone B can automatically input the feature information into the second model, and through the second model, translate the feature information corresponding to the English voice message into Chinese subtitles. For example, according to the English "this meeting is", output the corresponding Chinese operation information (such as a control instruction) "此次会议是".
[0252] For another example, in the scenario where mobile phone B does not enable the cross-language translation function, the second model outputs the recognition result of the English operation information according to the feature information corresponding to the English voice message. For example, output the corresponding English operation information "thismeeting is".
[0253] S407. The processing module of mobile phone B controls to display the subtitle information in the second language and / or play the second voice message in the second language.
[0254] Exemplarily, the processing module controls the display module to display the translated Chinese subtitle "此次会议是", and the display module can also display the English subtitle "this meeting is". For another example, the processing module controls the audio output module (such as a speaker) to play the translated Chinese voice "此次会议是", and the speaker can also play the English voice "this meetingis".
[0255] In the remote conference translation scenario, mobile phone B only needs to run the process of the feature information of the source language - translation result, and does not need to run the process of the source language voice message - the feature information of the source language (i.e., extracting feature information), which reduces the computing amount of mobile phone B and can improve the translation efficiency.
[0256] For training methods of the first and second models in remote conferencing and facial recognition scenarios, please refer to [link to relevant documentation]. Figure 9 , Figure 10 The model training method will not be elaborated here. In one possible design, the first model is a model trained based on at least one first sample data, the first sample data including: first speech information, the feature information of the first speech information is known, and / or, the second model is a model trained based on at least one second sample data, the second sample data including: first feature information, the operation information corresponding to the first feature information is known.
[0257] As can be seen, the technical solutions in this application embodiment can break down complex parametric models (including but not limited to machine learning models) and non-parametric models into multiple sub-models with smaller granularity. These sub-models can then be run on different modules of the same device, on different devices within the same network (as in the aforementioned speech recognition scenario), or on multiple devices in different networks (as in the aforementioned remote conferencing scenario). This reduces the computational load on individual modules or devices, thereby improving the overall processing efficiency of the task. Furthermore, this application embodiment does not limit the granularity of the sub-model splitting, the splitting method, or which modules or devices the split models are deployed on; these can be flexibly determined according to the scenario, device type, and other characteristics.
[0258] Furthermore, the above only lists a few possible application scenarios. The technical solutions of this application embodiment can also be applied to other scenarios. Due to space limitations, not all possible scenarios will be exhaustively listed here. For example, this application embodiment can be applied to the bone voiceprint recognition scenario. Currently, bone voiceprint technology, one of the biometric technologies, has high recognition rate, speed, and convenience. Its principle for identifying a person is: collecting the person's voice information and verifying the legitimacy of the person's identity based on the voice information. Since each person's bone structure is unique, the echo of sound reflected between bones is also unique. The echo of sound reflected between bones can be called a bone voiceprint. Similar to how fingerprints can be used to identify different people, bone voiceprints can be used to identify the identity of different users.
[0259] In this embodiment of the application, in the bone conduction voiceprint recognition scenario, the model used for bone conduction voiceprint recognition can be split into two parts. One part (the first model) is set in a Bluetooth headset, and the other part (the second model) is set in a mobile phone. After the headset collects the user's voice information (such as the user inputting "unlock screen"), it can extract the feature information of the sound signal (also known as voice information) through the set first model and send the feature information to the mobile phone. The mobile phone uses the second model to identify whether the voice is the legitimate user's voice. If so, it performs the corresponding operation (such as unlocking the screen).
[0260] Figure 14 The flowchart of the distributed voice control method provided in this application embodiment is illustrated. The method is applied to a first terminal and includes:
[0261] S1401, the first terminal responds to the voice information input by the user, inputs the voice information into the first model, and obtains the feature information corresponding to the voice information through the first model.
[0262] The first model exists in the first terminal, and the second model exists in the second terminal.
[0263] For example, taking a mobile phone as the first terminal, such as... Figure 6 As shown, the mobile phone receives the user's voice input "turn up the TV volume" and outputs the feature information (i.e., feature matrix) of the voice information through the first model.
[0264] S1402, the first terminal sends feature information to the second terminal so that the second terminal inputs the feature information into the second model, determines the operation information corresponding to the voice information through the second model, and performs the corresponding operation according to the operation information.
[0265] Still with Figure 6 For example, the second terminal includes a desk lamp, air conditioner, and television connected to the mobile phone. After obtaining the feature information corresponding to the voice information, the mobile phone broadcasts the feature information to the desk lamp, air conditioner, and television. The desk lamp, air conditioner, and television recognize the operation information (such as control commands) through the second model. Among them, if the operation information recognized by the television matches the television, the television will execute the target operation corresponding to the operation information "turn up the television volume", that is, turn up its own playback volume.
[0266] It should be noted that some operations in the processes of the above method embodiments are optionally combined, and / or the order of some operations is optionally changed. Furthermore, the execution order between the steps of each process is merely exemplary and does not constitute a limitation on the execution order between steps; other execution orders are also possible. It is not intended to indicate that the execution order is the only possible order in which these operations can be performed. Those skilled in the art will conceive of various ways to reorder the operations described herein. Additionally, it should be noted that this document combines other methods described herein (e.g., Figure 7 Corresponding methods Figure 8 The details of other processes described in the corresponding method also apply in a similar manner to the above-mentioned combination. Figure 12 The method described.
[0267] Alternatively, some steps in the method embodiments can be equivalently replaced with other possible steps. Alternatively, some steps in the method embodiments can be optional and can be deleted in certain use cases. Alternatively, other possible steps can be added to the method embodiments.
[0268] Other embodiments of this application provide an apparatus, which can be the aforementioned electronic device (such as a foldable screen phone). The apparatus may include a display screen, a memory, and one or more processors. The display screen, memory, and processors are coupled. The memory stores computer program code, which includes computer instructions. When the processor executes the computer instructions, the electronic device can perform various functions or steps performed by the mobile phone in the above method embodiments. The structure of the electronic device can be referred to... Figure 4 or Figure 5 The electronic device shown.
[0269] The core structure of this electronic device can be represented as follows: Figure 15 The structure shown may include: a processing module 1301, an input module 1302, a storage module 1303, and a display module 1304. Figure 15 The components shown are merely exemplary; electronic devices may include more or fewer components than illustrated, or combine or separate certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of both.
[0270] The processing module 1301 may include at least one of a central processing unit (CPU), an application processor (AP), or a communication processor (CP). The processing module 1301 can perform operations or data processing related to the control and / or communication with at least one of the other components of the user's electronic device. Specifically, the processing module 1301 can be used to control the content displayed on the main screen according to certain triggering conditions, or to determine the content displayed on the screen according to preset rules. The processing module 1301 is also used to process input instructions or data and determine the display style based on the processed data.
[0271] In the embodiments of this application, if Figure 15 The structure shown is a first electronic device (first terminal) or chip system. The processing module 1301 is used to respond to the voice information input by the user, input the voice information into a first model, and obtain the feature information corresponding to the voice information through the first model.
[0272] In the embodiments of this application, if Figure 15The structure shown is a second electronic device (second terminal) or chip system. The processing module 1301 is used to input the feature information into the second model and determine the operation information corresponding to the voice information through the second model.
[0273] The processing module is used to perform corresponding operations based on the operation information.
[0274] In one possible design, the second terminal performs a corresponding operation based on the operation information, including:
[0275] If it is determined that the operation information corresponding to the voice information is the operation information matched by the second terminal, then the second terminal performs the target operation according to the operation information corresponding to the voice information, and / or, if it is determined that the operation information corresponding to the voice information is not the operation information matched by the second terminal, then the second terminal discards the operation information.
[0276] Input module 1302 is used to acquire user input instructions or data and transmit the acquired instructions or data to other modules of the electronic device. Specifically, the input method of input module 1302 may include touch, gesture, proximity to the screen, or voice input. For example, the input module may be the screen of the electronic device, acquire user input operations, generate input signals based on the acquired input operations, and transmit the input signals to processing module 1301.
[0277] The storage module 1303 may include volatile memory and / or non-volatile memory. The storage module is used to store at least one related instruction or data from other modules of the user terminal device; specifically, the storage module may store a first model and a second model.
[0278] Display module 1304 may include, for example, a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic light-emitting diode (OLED) display, a microelectromechanical system (MEMS) display, or an electronic paper display. It is used to display user-viewable content (e.g., text, images, videos, icons, symbols, etc.).
[0279] Optional, Figure 15 The structure shown may also include an output module (not shown in the diagram). Figure 15 (As shown in the diagram). Output modules can be used to output information. For example, playing or outputting voice information. Output modules include, but are not limited to, modules such as speakers.
[0280] Optional, Figure 15The illustrated structure may also include a communication module 1305 for supporting communication between the electronic device and other electronic devices. For example, the communication module may be connected to a network via wireless or wired communication to communicate with other personal terminals or network servers. Wireless communication may employ at least one of the following cellular communication protocols: 5G, Long Term Evolution (LTE), LTE-A Advanced, Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Universal Mobile Telecommunications System (UMTS), Wi-Fi, or Global System for Mobile Communications (GSM). Wireless communication may include, for example, short-range communication. Short-range communication may include at least one of Wi-Fi, Bluetooth, Near Field Communication (NFC), Magnetic Stripe Transmission (MST), or GNSS.
[0281] In the embodiments of this application, if Figure 15 The structure shown is a first electronic device or chip system, with a communication module 1305 used to send the feature information to the second terminal.
[0282] Optionally, sending the feature information to the second terminal includes broadcasting the feature information.
[0283] In the embodiments of this application, if Figure 15 The structure shown is a second electronic device or chip system, with a communication module 1305 used to receive feature information corresponding to voice information from the first terminal.
[0284] It should be noted that the descriptions of each step in the method embodiments of this application can be referenced to the corresponding modules of the device, and will not be repeated here.
[0285] This application also provides a chip system, such as... Figure 16 As shown, the chip system includes at least one processor 1401 and at least one interface circuit 1402. The processor 1401 and the interface circuit 1402 are interconnected via lines. For example, the interface circuit 1402 can be used to receive signals from other devices (e.g., the memory of an electronic device). As another example, the interface circuit 1402 can be used to send signals to other devices (e.g., the processor 1401). Exemplarily, the interface circuit 1402 can read instructions stored in memory and send those instructions to the processor 1401. When the instructions are executed by the processor 1401, the electronic device can perform the steps in the above embodiments. Of course, the chip system may also include other discrete devices, which are not specifically limited in this application embodiment.
[0286] This application also provides a computer storage medium that includes computer instructions. When the computer instructions are executed on the electronic device, the electronic device causes the electronic device to perform various functions or steps performed by the mobile phone in the above method embodiment.
[0287] This application also provides a computer program product that, when run on a computer, causes the computer to perform the various functions or steps performed by the mobile phone in the above method embodiments.
[0288] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0289] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another apparatus, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0290] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0291] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0292] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, in essence, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0293] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A distributed voice control method, characterized in that, The method includes: The first terminal responds to the voice information input by the user, inputs the voice information into the first model, and obtains the feature information corresponding to the voice information through the first model. The first model exists in the first terminal. The first terminal sends the feature information to the second terminal, so that the second terminal inputs the feature information into the second model, determines the operation information corresponding to the voice information through the second model, and performs the corresponding operation according to the operation information. The second model exists in the second terminal. The first model and the second model are sub-models of a model for voice control, which is an encoder-decoder model. The first model is the encoding part of the voice control model, and the second model is the decoding part. The first model is used to acquire multi-dimensional feature information corresponding to the voice information. The first model is an encoding model, trained by the manufacturer of the first terminal, and its weights are fixed. The second model is used to convert the multi-dimensional feature information corresponding to the voice information into corresponding operation information. The second model is a decoding model, trained by the manufacturer of the second terminal, and its weights are updatable.
2. The method according to claim 1, characterized in that, The first model is trained based on at least one first sample data, the first sample data including: first speech information, the feature information of the first speech information being known; and / or, The second model is a model trained based on at least one second sample data, which includes: first feature information, and the operation information corresponding to the first feature information is known.
3. The method according to claim 1 or 2, characterized in that, The first terminal sends the feature information to the second terminal, including: the first terminal broadcasting the feature information.
4. A distributed voice control method, characterized in that, The method includes: The second terminal receives feature information corresponding to the voice information from the first terminal; the feature information is obtained by the first terminal inputting the voice information into a first model and obtaining it through the first model, and the first model exists in the first terminal. The second terminal inputs the feature information into the second model and determines the operation information corresponding to the voice information through the second model. The second model exists in the second terminal. The second terminal performs the corresponding operation based on the operation information; The first model and the second model are sub-models of a model for voice control, which is an encoder-decoder model. The first model is the encoding part of the voice control model, and the second model is the decoding part. The first model is used to acquire multi-dimensional feature information corresponding to the voice information. The first model is an encoding model, trained by the manufacturer of the first terminal, and its weights are fixed. The second model is used to convert the multi-dimensional feature information corresponding to the voice information into corresponding operation information. The second model is a decoding model, trained by the manufacturer of the second terminal, and its weights are updatable.
5. The method according to claim 4, characterized in that, The second terminal performs corresponding operations based on the operation information, including: If it is determined that the operation information corresponding to the voice information is the operation information matched by the second terminal, then the second terminal executes the target operation according to the operation information corresponding to the voice information; and / or, If it is determined that the operation information corresponding to the voice information is not the operation information matched by the second terminal, then the second terminal discards the operation information.
6. The method according to claim 4 or 5, characterized in that, The first model is trained based on at least one first sample data, the first sample data including: first speech information, the feature information of the first speech information being known; and / or, The second model is a model trained based on at least one second sample data, which includes: first feature information, and the operation information corresponding to the first feature information is known.
7. A first terminal, characterized in that, include: Display screen; One or more processors; One or more memory units; The memory stores one or more programs that, when executed by the processor, cause the first terminal to perform the method as described in any one of claims 1 to 3.
8. A second terminal, characterized in that, include: Display screen; One or more processors; One or more memory units; The memory stores one or more programs that, when executed by the processor, cause the second terminal to perform the method as described in any one of claims 4 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed on a terminal, cause the terminal to perform the method as described in any one of claims 1 to 3, or to perform the method as described in any one of claims 4 to 6.
10. A computer program product, characterized in that, When the computer program product is run on a terminal, the terminal performs the method as described in any one of claims 1 to 3, or performs the method as described in any one of claims 4 to 6.