Voice control method and apparatus

By using a semantic reasoning model to determine and select the target device for voice control commands, the problem of simplicity and user-friendliness in traditional control methods is solved, thus improving the user experience.

CN115083401BActive Publication Date: 2026-04-07GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-10
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Traditional control panel-based operation methods cannot provide users with a simple and user-friendly control experience. The homogenization of functions in smart devices makes it difficult to correctly understand the user's intentions and select the appropriate device.

Method used

The target device for executing voice control commands is determined by a semantic reasoning model. It is then determined whether the voice data contains the name of the target device. If not, the target device is selected by calling the semantic reasoning model based on the identity information.

Benefits of technology

It enables the selection of the execution device based on user intent, thereby improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115083401B_ABST
    Figure CN115083401B_ABST
Patent Text Reader

Abstract

This application discloses a voice control method and apparatus. The method includes: an electronic device receiving voice data from a target user, the voice data including voice control instructions used to instruct the electronic device to execute operation commands; determining whether the voice data includes the name of a target device; if the voice data includes the name of the target device, converting the voice control instructions into device control instructions; otherwise, obtaining the target user's identity information based on the voice data, and using the identity information to call a semantic reasoning model to determine the target device; finally, sending the device control instructions to the target device. This application uses a semantic reasoning model to determine the target device for executing voice control instructions, thereby achieving the selection of the execution device based on the user's intent and improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a voice control method and apparatus. Background Technology

[0002] With the continuous advancement of IoT technology, the interconnection, interoperability, and integration of everything are current technological hotspots and future trends. More and more smart terminal devices, such as mobile phones, speakers, tablets, televisions, and air conditioners, are appearing in the public eye, evolving from single devices to distributed devices across multiple scenarios. Traditional control panel-based operation methods no longer provide users with a simple, user-friendly, and intelligent control experience. Significant breakthroughs in voice recognition and semantic understanding technologies have led to their rapid popularization. Furthermore, due to its simplicity, convenience, and intelligence, voice-based interaction has become the mainstream method for controlling distributed devices.

[0003] Currently, the processing power of smart devices is gradually increasing, and the differences in capabilities between devices are gradually decreasing or disappearing, resulting in functional homogenization. For example, smart TVs and mobile phones are functionally identical in terms of video and music services. Therefore, how to correctly understand the user's intentions to select the device that provides the service is an urgent problem to be solved. Summary of the Invention

[0004] This application provides a voice control method and apparatus, which uses a semantic reasoning model to determine the target device for executing voice control commands, thereby enabling the selection of the execution device based on the user's intent and improving the user experience.

[0005] In a first aspect, embodiments of this application provide a voice control method applied to an arbitration device, the method comprising:

[0006] Receive voice data from a target user, the voice data including voice control instructions, the voice control instructions being used to instruct the electronic device to execute operation commands;

[0007] Determine whether the voice data includes the name of the target device;

[0008] When the voice data includes the name of the target device, the voice control command is converted into a device control command; otherwise, the target user's identity information is obtained based on the voice data, and the semantic reasoning model is invoked based on the identity information to determine the target device.

[0009] Send the device control command to the target device.

[0010] Secondly, embodiments of this application provide a voice control device applied to arbitration equipment, the device comprising:

[0011] A transceiver unit is used to receive voice data from a target user, the voice data including voice control instructions, the voice control instructions being used to instruct an electronic device to execute an operation command;

[0012] The processing unit is configured to determine whether the voice data includes the name of the target device; if the voice data includes the name of the target device, the voice control command is converted into a device control command; otherwise, the target user's identity information is obtained based on the voice data, and the semantic reasoning model is invoked based on the identity information to determine the target device.

[0013] The transceiver unit is also used to send the device control command to the target device.

[0014] Thirdly, embodiments of this application provide an electronic device, including a processor, a memory, a communication interface, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the processor, and the programs include instructions for performing steps in any method of the first aspect of this application.

[0015] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program for electronic data interchange, wherein the computer program causes a computer to perform some or all of the steps described in any method of the first aspect of this application.

[0016] Fifthly, embodiments of this application provide a computer program product, wherein the computer program product includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps described in any method of the first aspect of this application. The computer program product may be a software installation package.

[0017] As can be seen, in this embodiment, the electronic device receives voice data from the target user, including voice control instructions used to instruct the electronic device to execute operation commands. It then determines whether the voice data includes the name of the target device. If the voice data does, the voice control instructions are converted into device control instructions. Otherwise, the target user's identity information is obtained from the voice data, and a semantic reasoning model is invoked to determine the target device based on the identity information. Finally, the device control instructions are sent to the target device. This application uses a semantic reasoning model to determine the target device for executing voice control instructions, thereby enabling the selection of the execution device based on the user's intent and improving the user experience. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0020] Figure 2 This is a schematic diagram of the software structure of an electronic device provided in an embodiment of this application;

[0021] Figure 3a This is a schematic diagram of a device control system provided in an embodiment of this application;

[0022] Figure 3b A schematic diagram of the structure of an arbitration device provided in an embodiment of this application;

[0023] Figure 4 This application provides a schematic diagram of a device wake-up process.

[0024] Figure 5 This application provides a flowchart illustrating a voice control method.

[0025] Figure 6 This is a flowchart illustrating a training method for a first semantic reasoning model provided in an embodiment of this application;

[0026] Figure 6a This is a flowchart illustrating another voice control method provided in an embodiment of this application;

[0027] Figure 7 This is a schematic diagram of the structure of a voice control device provided in an embodiment of this application. Detailed Implementation

[0028] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0029] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.

[0030] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application can be combined with other embodiments.

[0031] The voice control method provided in this application can be applied to handheld devices, in-vehicle devices, wearable devices, augmented reality (AR) devices, virtual reality (VR) devices, projection devices, projectors, or other devices connected to a wireless modem. It can also be various specific forms of user equipment (UE), terminal device, smartphone, smart screen, smart TV, smartwatch, laptop, smart speaker, camera, game controller, mouse, microphone, station (STA), access point (AP), mobile station (MS), personal digital assistant (PDA), personal computer (PC), or relay device, etc. This application does not impose any restrictions on the specific types of terminal devices and servers.

[0032] For example, the terminal device may be a station (ST) in a WLAN, a cellular phone, a cordless phone, a Session Initiation Protocol (SIP) phone, a Wireless Local Loop (WLL) station, a Personal Digital Assistant (PDA) device, a handheld device with wireless communication capabilities, a computing device or other processing device connected to a wireless modem, an in-vehicle device, a vehicle networking terminal, a computer, a laptop computer, a handheld communication device, a handheld computing device, a satellite wireless device, a wireless modem card, a set-top box (STB), customer premises equipment (CPE), and / or other devices for communication on a wireless device, as well as next-generation communication devices, such as mobile terminals in 5G networks or mobile terminals in future evolved Public Land Mobile Network (PLMN) networks.

[0033] As an example and not a limitation, when the terminal device is a wearable device, the term "wearable device" can also refer to any device that utilizes wearable technology to intelligently design and develop everyday wearables, such as glasses, gloves, watches, clothing, and shoes. Wearable devices are portable devices worn directly on the body or integrated into a user's clothing or accessories. Wearable devices are not merely hardware devices; they achieve powerful functions through software support, data interaction, and cloud interaction. Broadly defined, wearable smart devices include those with comprehensive functions, large sizes, and the ability to perform complete or partial functions without relying on a smartphone, such as smartwatches or smart glasses, as well as those focused on a specific application function that require interaction with other devices such as smartphones, such as various smart bracelets and smart jewelry for vital sign monitoring.

[0034] The first part describes the software and hardware operating environment of the technical solution disclosed in this application.

[0035] For example, Figure 1A schematic diagram of the structure of electronic device 100 is shown. Electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, antenna 1, antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, a compass 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.

[0036] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0037] Processor 110 may include one or more processing units, such as application processors (APs), modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). Different processing units may be independent components or integrated into one or more processors. In some embodiments, electronic device 100 may also include one or more processors 110. The controller can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution. In other embodiments, processor 110 may also include a memory for storing instructions and data. For example, the memory in processor 110 may be a cache memory. This memory can store instructions or data that processor 110 has just used or is reusing. If processor 110 needs to reuse the instruction or data, it can directly retrieve it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the electronic device 100 in processing data or executing instructions.

[0038] In some embodiments, the processor 110 may include one or more interfaces. These interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a SIM card interface, and / or a USB interface, etc. The USB interface 130 is a USB standard-compliant interface, specifically a Mini USB interface, a Micro USB interface, a USB Type-C interface, etc. The USB interface 130 can be used to connect a charger to charge the electronic device 100, and can also be used for data transfer between the electronic device 100 and peripheral devices. The USB interface 130 can also be used to connect headphones for audio playback.

[0039] It is understood that the interface connection relationships between the modules illustrated in the embodiments of this application are merely illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.

[0040] The charging management module 140 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 receives charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 receives wireless charging input via the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also supply power to the electronic device via the power management module 141.

[0041] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, external memory, display screen 194, camera 193, and wireless communication module 160. The power management module 141 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.

[0042] The wireless communication function of electronic device 100 can be implemented through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor, and baseband processor.

[0043] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with tuning switches.

[0044] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.

[0045] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), and UWB. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0046] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0047] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a mini light-emitting diode (miniled), a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or more display screens 194.

[0048] Electronic device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display screen 194 and application processor.

[0049] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.

[0050] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the electronic device 100 may include one or more cameras 193.

[0051] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when electronic device 100 selects a frequency, the DSP can perform Fourier transforms on the frequency energy.

[0052] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. Thus, electronic device 100 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG 2, MPEG 3, MPEG 4, etc.

[0053] An NPU (Neural Processing Unit) is a computational processor for neural networks (NNs). By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.

[0054] The external storage interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.

[0055] Internal memory 121 can be used to store one or more computer programs, which include instructions. Processor 110 can execute the instructions stored in internal memory 121, thereby causing electronic device 100 to perform the methods for displaying page elements provided in some embodiments of this application, as well as various applications and data processing. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system; the program storage area may also store one or more applications (such as a gallery, contacts, etc.). The data storage area may store data created during the use of electronic device 100 (such as photos, contacts, etc.). In addition, internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as one or more disk storage components, flash memory components, universal flash storage (UFS), etc. In some embodiments, processor 110 can execute instructions stored in internal memory 121 and / or instructions stored in memory disposed in processor 110, thereby causing electronic device 100 to perform the methods for displaying page elements provided in embodiments of this application, as well as other applications and data processing. Electronic device 100 can implement audio functions such as music playback and recording through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.

[0056] The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0057] The pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. In some embodiments, the pressure sensor 180A can be disposed on the display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, and capacitive pressure sensors. A capacitive pressure sensor may include at least two parallel plates with conductive materials. When a force is applied to the pressure sensor 180A, the capacitance between the electrodes changes. The electronic device 100 determines the pressure intensity based on the change in capacitance. When a touch operation is applied to the display screen 194, the electronic device 100 detects the intensity of the touch operation based on the pressure sensor 180A. The electronic device 100 can also calculate the touch position based on the detection signal from the pressure sensor 180A. In some embodiments, touch operations applied to the same touch position but with different touch operation intensities can correspond to different operation commands. For example, when a touch operation with an intensity less than a first pressure threshold is applied to the SMS application icon, a command to view an SMS message is executed. When a touch operation with an intensity greater than or equal to the first pressure threshold is applied to the SMS application icon, a command to create a new SMS message is executed.

[0058] The gyroscope sensor 180B can be used to determine the motion attitude of the electronic device 100. In some embodiments, the gyroscope sensor 180B can determine the angular velocity of the electronic device 100 around three axes (i.e., the X, Y, and Z axes). The gyroscope sensor 180B can be used for image stabilization. For example, when the shutter is pressed, the gyroscope sensor 180B detects the angle of the electronic device 100's shake, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to counteract the shake of the electronic device 100 through reverse movement, thus achieving image stabilization. The gyroscope sensor 180B can also be used in navigation and motion-sensing game scenarios.

[0059] The 180E accelerometer can detect the magnitude of acceleration of electronic device 100 in various directions (typically three axes). When electronic device 100 is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the posture of electronic devices and applied to applications such as screen orientation switching and pedometers.

[0060] The ambient light sensor 180L is used to sense the brightness of ambient light. The electronic device 100 can adaptively adjust the brightness of the display screen 194 based on the sensed ambient light brightness. The ambient light sensor 180L can also be used to automatically adjust the white balance when taking pictures. The ambient light sensor 180L can also work with the proximity sensor 180G to detect whether the electronic device 100 is in a pocket to prevent accidental touches.

[0061] The fingerprint sensor 180H is used to collect fingerprints. The electronic device 100 can utilize the characteristics of the collected fingerprints to achieve fingerprint unlocking, accessing application locks, taking photos with fingerprints, answering calls with fingerprints, etc.

[0062] Temperature sensor 180J is used to detect temperature. In some embodiments, electronic device 100 uses the temperature detected by temperature sensor 180J to execute a temperature handling strategy. For example, when the temperature reported by temperature sensor 180J exceeds a threshold, electronic device 100 performs thermal protection by reducing the performance of a processor located near temperature sensor 180J to reduce power consumption. In other embodiments, when the temperature is below another threshold, electronic device 100 heats battery 142 to prevent abnormal shutdown of electronic device 100 due to low temperature. In still other embodiments, when the temperature is below yet another threshold, electronic device 100 boosts the output voltage of battery 142 to prevent abnormal shutdown due to low temperature.

[0063] Touch sensor 180K, also known as a "touch panel," can be located on display screen 194. The touch sensor 180K and display screen 194 together form a touchscreen, also known as a "touch screen." Touch sensor 180K detects touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 194. In other embodiments, touch sensor 180K may also be located on the surface of electronic device 100, in a different position than display screen 194.

[0064] For example, Figure 2 A software architecture block diagram of the electronic device 100 is shown. The layered architecture divides the software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer. The application layer may include a series of application packages.

[0065] like Figure 2 As shown, the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, and SMS.

[0066] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.

[0067] like Figure 2 As shown, the application framework layer may include a window manager, content provider, view system, phone manager, resource manager, notification manager, etc.

[0068] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.

[0069] Content providers store and retrieve data, making that data accessible to applications. This data may include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc.

[0070] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.

[0071] The phone manager is used to provide communication functions for electronic device 100. For example, it manages call status (including connection and disconnection).

[0072] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.

[0073] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.

[0074] The Android Runtime consists of core libraries and a virtual machine. The Android runtime is responsible for the scheduling and management of the Android system.

[0075] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.

[0076] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0077] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.

[0078] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.

[0079] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.

[0080] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.

[0081] A 2D graphics engine is a graphics engine for 2D drawing.

[0082] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.

[0083] The second part introduces the example application scenarios disclosed in the embodiments of this application as follows.

[0084] For example, the technical solutions of the embodiments of this application can be applied to, for example, Figure 3a The device control system shown may include a voice acquisition device, a first type of electronic device, a second type of electronic device, and an arbitration device. The arbitration device may connect to multiple first-type electronic devices, second-type electronic devices, and voice acquisition devices, and the multiple first-type electronic devices and multiple second-type electronic devices may communicate with each other via wireless network or wired data connection.

[0085] Among them, voice acquisition devices possess basic voice input, voice output, or voice recognition capabilities, serving only as the voice input interface for user voice control commands without executing those commands (or voice control instructions, voice control commands). The first type of electronic devices possesses both voice reception and service provision capabilities, such as mobile phones, smart speakers, and televisions. They can function as voice receiving devices in a distributed device control system or as controlled devices executing user voice control commands. The second type of electronic devices provides services but lacks the ability to provide voice input, acting only as controlled devices executing user voice control commands, such as air conditioners, refrigerators, and washing machines. Arbitration devices are primarily responsible for target device wake-up arbitration and user intent recognition. These arbitration devices can be devices from the first or second type of electronic devices that possess hardware resources supporting the aforementioned capabilities, or they can be independent servers or remote cloud servers, etc.

[0086] Specifically, the voice acquisition device collects the user's voice data and transmits it to the arbitration device. The arbitration device then selects an electronic device from among multiple connected first-class and second-class electronic devices that matches the user's intent to execute the user's command. For example, the user speaks "play the news broadcast" through the microphone on their mobile phone. The phone then processes the speech, identifies the user's true intent, and selects an electronic device from smart TVs, mobile phones, laptops, and tablets that matches the user's intent to play the news broadcast.

[0087] For example, a voice assistant can be installed in the electronic device of the aforementioned voice control system to enable the electronic device to perform voice control functions. The voice assistant is generally in a dormant state. Before using the voice control function of the electronic device, the user needs to wake up the voice assistant using voice. The voice data used to wake up the voice assistant can be called a wake-up word (or wake-up voice). This wake-up word can be pre-registered in the electronic device. In this embodiment, waking up the voice assistant means that the electronic device activates the voice assistant in response to the user's spoken wake-up word. The voice control function means that after the voice assistant of the electronic device is activated, the user can trigger the electronic device to automatically execute the event corresponding to the voice command by speaking a voice command (e.g., a piece of voice data).

[0088] Furthermore, the aforementioned voice assistant can be an embedded application within an electronic device (i.e., a system application of the electronic device) or a downloadable application. An embedded application is an application provided as part of the implementation of an electronic device (such as a mobile phone). A downloadable application is an application that can provide its own Internet Protocol Multimedia Subsystem (IMS) connectivity. Downloadable applications can be pre-installed on electronic devices or are third-party applications downloaded and installed by the user.

[0089] For example, such as Figure 3b As shown, Figure 3b This is a schematic diagram of the structure of an arbitration device provided in an embodiment of this application. Figure 3b As shown, the arbitration device includes a voice wake-up module, a user identification module, a semantic reasoning model, an optimization and update module, and a user registration module.

[0090] Before using the arbitration device for voice control, users can pre-register or obtain their biometric information (such as voiceprint, fingerprint, face, iris, etc.) and / or user account information from connected electronic devices. For example, if a user logs into multiple electronic devices connected to the arbitration device using the same user account, the arbitration device can obtain information such as the user's voice, face, fingerprint, and iris from these multiple electronic devices. The obtained biometric information and / or user account information are then stored to provide identification features for subsequent tasks.

[0091] Furthermore, the user identification module identifies users by using biometric information stored in the user registration module, thus providing identity verification for subsequent personalized control. The user identification module roughly categorizes users into registered users and unregistered users, with registered users specifically identified as the specific users within the user registration module.

[0092] The main function of the semantic reasoning model is to perform semantic parsing and device inference. For example, when a user says "open the news broadcast," the semantic reasoning model needs to infer which electronic device the user wants to play the news broadcast on. Simultaneously, for personalized reasoning, this module includes at least one general semantic reasoning model and a user-specific semantic reasoning model. The general semantic reasoning model is a reasoning model trained using a general predictive tool for unregistered users, while the user-specific semantic reasoning model is a reasoning model fine-tuned using the user's predictive tool based on the general semantic reasoning model. The semantic reasoning model understands the user's intent from implicit control commands, such as "open the news broadcast," and accurately infers the electronic device the user wishes to control. This module uses a large-scale corpus to train a semantic reasoning model, which learns common sense from our social life from the corpus to assist in understanding the user's intent. This can be abstractly expressed as:

[0093] P(Device|Device Control Command) = model(Voice Control Command)

[0094] Here, `model` is the semantic reasoning model, which takes voice control commands as input and approximates the preceding probability model. For example, for the command "open the news broadcast," for an unregistered user, since there is no user-specific semantic reasoning model for that user, based on our everyday semantic information, the value of P(TV|open the news broadcast) should be much larger than the values ​​of P(mobile phone|open the news broadcast) and P(tablet|open the news broadcast).

[0095] Furthermore, the optimization and update module can continuously learn and optimize the user-specific semantic reasoning model based on the registered user's operating habits, thereby enabling the user-specific semantic reasoning model to more accurately identify the user's intent.

[0096] Finally, when the user's voice data includes the wake-up word of an electronic device, such as when the voice data received by the arbitration device includes "Xiao Bu Xiao Bu" and the wake-up word of each electronic device connected to the arbitration device is "Xiao Bu Xiao Bu", the arbitration device can wake up the target device after selecting the target device and relay the user's subsequent control commands to the target device.

[0097] For example, the device wake-up process is explained using an arbitration device as the server and a first-class electronic device such as a smart speaker, smartphone, or smart TV. Please refer to [link / reference]. Figure 4First, the user inputs a wake-up voice command, such as "Xiao Bu Xiao Bu." Second, electronic devices with voice input capabilities (smart speakers, smartphones, and smart TVs) receive this wake-up voice command. These devices are equipped with a smart voice assistant and are in sleep mode. Third, the electronic device matches the wake-up voice command with a pre-stored wake-up word. If the match is successful, the device uploads the signal strength of the received wake-up voice command, its service capability information, and its device identification information to the server. Next, the server receives the above information to complete the registration of the electronic device and responds to the wake-up voice command according to preset wake-up rules (such as proximity, recent usage time, or highest historical usage frequency) to determine whether to wake up the smart speaker and issue a control command to it. Finally, the smart speaker receives the user's control command, activates its own voice assistant, and sends a prompt message to the user (such as "Here, Master").

[0098] It should be understood that the equipment control system may also include other electronic devices, which are not specifically limited here.

[0099] Part Three, the scope of protection disclosed in the embodiments of this application is described below.

[0100] Please see Figure 5 , Figure 5 This application provides a flowchart illustrating a voice control method, applied to the above-mentioned... Figure 3b Arbitration equipment in, such as Figure 5 As shown, this voice control method includes the following operations.

[0101] S510. Receive voice data from the target user, the voice data including voice control instructions, the voice control instructions being used to instruct the electronic device to execute operation commands.

[0102] In this embodiment, after the voice acquisition device acquires the user's voice data, it can send the voice data to the arbitration device so that the arbitration device can select a target device to execute the voice control commands in the voice data.

[0103] For example, when the first type of electronic device and / or the second type of electronic device are in a sleep state, the user can first output a wake-up voice to wake up the first type of electronic device and / or the second type of electronic device to execute the user's subsequent voice control commands. Specifically, before speaking the voice control command, the user can first speak the wake-up voice. After the voice acquisition device collects the user's wake-up voice, it can send the wake-up voice to the arbitration device. The arbitration device can send the wake-up command to the electronic device that needs to be woken up according to the wake-up rules. Then, after receiving the voice operation command from the voice acquisition device, it outputs the target device of the voice operation command according to the semantic reasoning model. Finally, the voice operation command is converted into a device operation command and sent to the target device so that the target device can execute the user's operation command.

[0104] It should be noted that the electronic device receiving the wake-up command sent by the arbitration device and the target device may or may not be the same electronic device. For example, after the server receives the user's wake-up phrase "Hello Xiaobu," because the smart speaker is relatively close to the user (the smart speaker receives the highest signal energy from the wake-up phrase), the server can send a wake-up command to the smart speaker to wake it up, and then receive the user's subsequent commands through the smart speaker. When the user says "Play music," after the server receives this voice control command "Play music," the semantic reasoning module determines that the target device is a mobile phone (the user frequently uses their mobile phone to play music in daily life). The server then sends the "Play music" control command to the mobile phone. After receiving the control command, the mobile phone provides audio playback services to execute the control command.

[0105] For example, when the target user's voice data includes wake-up commands and voice control instructions, the arbitration device can determine the target device for executing the voice control instructions based on the semantic reasoning module, and then simultaneously send the wake-up command and device control command to the target device. Upon receiving the wake-up command and device control command, the target device can wake up first and then execute the control command. For instance, when a mobile phone captures a user saying "Hello Xiaobu, please play the news broadcast," the target device determined by the semantic reasoning module in the mobile phone is a smart TV (the user frequently uses a smart TV to play the news broadcast in daily life). The mobile phone then sends a wake-up command and the control command "play the news broadcast" to the smart TV. Upon receiving these commands, the smart TV wakes up first and then provides video service to play the news broadcast.

[0106] S520. Determine whether the voice data includes the name of the target device.

[0107] In practical applications, when performing voice control, users may specify the target device to execute the voice command. For example, for the voice command "Play the news broadcast on TV," the user specifies the TV as the target device. Alternatively, users may not specify the target device, for example, for the voice command "Play the news broadcast." Therefore, after receiving the voice data from the target user, the arbitration device needs to determine whether the voice data is complete, that is, whether the voice data includes the name of the target device.

[0108] S530: When the voice data includes the name of the target device, the voice control command is converted into a device control command; otherwise, the target user's identity information is obtained based on the voice data, the semantic reasoning model is invoked based on the identity information to determine the target device, and the voice control command is converted into a device control command.

[0109] Upon receiving the voice data, the arbitration device needs to parse it. If the parsed voice data contains a voice control command and a target device, the arbitration device can directly convert the voice control command into a device control command and send it to the target device for execution. For example, if a user says, "Play the news broadcast on TV," and the arbitration device receives the voice data and parses it to find the device name "TV," the arbitration device can directly output the device control command corresponding to "TV" for "Play the news broadcast" to the TV.

[0110] Furthermore, if the voice data does not include the name of the target device, an arbitration device is needed to select the target device. Each user has their own expression habits in daily life, and different users may use different electronic devices to perform the same function or service. For example, user A's target device for "playing the news broadcast" is a mobile phone, while user B's target device is a television. Therefore, the arbitration device can obtain the target user's identity based on the voice data and, based on the user's identity, invoke the semantic reasoning model corresponding to the target user to select the target device.

[0111] Optionally, the identity information includes registered users and unregistered users;

[0112] The step of obtaining the target user's identity information based on the voice data includes:

[0113] Extract the voiceprint feature information corresponding to the voice data; match the voiceprint feature information with at least one pre-stored voiceprint feature information respectively; if the voiceprint feature information matches the pre-stored target voiceprint feature information, determine the identity information of the target user as the registered user, and the target voiceprint feature information is any one of the at least one voiceprint feature information; if the voiceprint feature information does not match any of the pre-stored voiceprint feature information, determine the identity information of the target user as the unregistered user.

[0114] Before using voice control, users can register their identity information with the arbitration device. Specifically, users can record their voice into the arbitration device or send the recorded voice to the arbitration device through a voice acquisition device. The arbitration device then processes the voice recorded by each user, extracts the voiceprint feature information of each user, and establishes a mapping relationship between the user and the voiceprint feature information.

[0115] For example, the voiceprint features extracted from speech data can be obtained using Linear Predictive Coding (LPC) features, MFCC features, Perceptual Linear Predictive (PLP) features, etc. This application does not limit the type of acoustic features.

[0116] Specifically, the voiceprint features extracted from the voice data are matched against the voiceprint features of registered users stored in the arbitration device. If the voiceprint features of the voice data match the voiceprint features stored in the arbitration device, it indicates that the target user is a registered user; if the voiceprint features of the voice data do not match the voiceprint features stored in the arbitration device, it indicates that the target user is an unregistered user.

[0117] Optionally, the semantic reasoning model includes multiple first semantic reasoning models and second semantic reasoning models, wherein the first semantic reasoning model is trained using the registered user corpus as training samples, and the second semantic reasoning model is trained using general corpus as training samples.

[0118] Each user corresponds to a first semantic reasoning model, and all unregistered users correspond to a second semantic reasoning model. Each first semantic reasoning model is trained based on multiple semantic manipulation commands of the corresponding registered user, while the second semantic reasoning model is trained based on multiple semantic manipulation commands commonly used in daily life.

[0119] Optionally, the step of determining the target device by invoking the semantic reasoning model based on the identity information includes: if the user is a registered user, inputting the voice control command into the first semantic reasoning model to obtain the target device; if the user is an unregistered user, inputting the voice control command into the second semantic reasoning model to obtain the target device.

[0120] Specifically, if the target user is a registered user, it indicates that the arbitration device has a first semantic reasoning model for that target user. Therefore, the first semantic reasoning model for the target user is determined based on the mapping relationship between voiceprint feature information and the first semantic reasoning model. This mapping relationship between voiceprint feature information and the first semantic reasoning model can be pre-stored or constructed during the training of the first semantic reasoning model; this embodiment does not limit this. Then, the voice control command is input into the first semantic reasoning model to obtain the target device for the voice control command. If the target user is an unregistered user, the second semantic reasoning model can be used directly to obtain the target device for the voice control command.

[0121] Optional, such as Figure 6 As shown, the training method for the first semantic reasoning model includes the following steps:

[0122] S610. Obtain the training dataset, which includes multiple voice data of registered users.

[0123] The training dataset can be audio data from users' daily voice operations. After a user registers on the arbitration device, the device stores the voice data received from the voice acquisition device, the first type of electronic device, and the second type of electronic device about the registered user in a database. Each voice data entry includes a voice control command and the target device specified by the user. For example, if the voice data does not include the name of the target device, the executing device corresponding to each voice control command is recorded.

[0124] Optionally, obtaining the training dataset includes: determining the target registered user corresponding to the target voiceprint feature information based on the mapping relationship between the voiceprint feature information and the registered user; obtaining multiple pieces of raw voice data, the raw voice data including the voice control commands of the target registered user; determining the execution device that executes the voice control commands of the target registered user; and labeling the execution device as the target device of the voice control commands to obtain the multiple pieces of voice data.

[0125] Specifically, after identifying the target user as the target registered user based on the mapping relationship between voiceprint feature information and registered users, the arbitration device can retrieve the target registered user's original voice data from the database. Then, the original voice data is parsed to obtain voice control commands and corresponding execution devices, which are then labeled as the target device for the voice control commands. Finally, the voice control commands labeled with the target device are used as training data to train the first semantic reasoning model to be trained.

[0126] S620. Perform feature extraction on the multiple voice data to obtain multiple audio features.

[0127] After acquiring the speech data for training, it is necessary to extract audio features from the speech data to train the first semantic reasoning model. These audio features can be Mel-frequency cepstral coefficients (MFCCs) and filter bank features, etc.

[0128] S630. Input the multiple audio features into the first semantic reasoning model to be trained for training until the training termination condition is met, and obtain the first semantic reasoning model.

[0129] The first semantic reasoning model to be trained can be a machine learning algorithm for classification, such as the K-means algorithm, the K-Nearest Neighbor (KNN) classification algorithm, decision trees, etc., or a neural network algorithm, such as a recurrent neural network (RNN), a convolutional neural network (CNN), a long short-term memory network (LSTM), and various variant neural network algorithms.

[0130] Optionally, the step of inputting the plurality of audio features into the first semantic reasoning model to be trained for training until the training termination condition is met to obtain the first semantic reasoning model includes: inputting the plurality of audio features into the first semantic reasoning model to be trained to obtain the output device corresponding to each speech data; constructing a loss function based on the output device and the labeled target device; updating the parameters corresponding to minimizing the loss function to the parameters of the first semantic reasoning model to be trained to obtain the first semantic reasoning model.

[0131] During model training, to ensure the model's output closely approximates the desired predicted value, the model's predictions are compared to the target value. The weight vector of the intelligent algorithm is updated based on the difference. For example, if the model's prediction is too high, the weight vector is adjusted to predict a lower value. This adjustment continues until the model can predict the target value accurately or very closely. The loss function is a crucial equation used to measure the difference between the predicted and target values. A higher loss function output indicates a greater difference, so model training becomes a process of minimizing this loss. Ultimately, the parameters corresponding to the minimum loss function are determined as the parameters for training the model.

[0132] Specifically, the audio features of each voice control command are input into the first semantic reasoning model to be trained, obtaining the first posterior probability of each first-class electronic device and each second-class electronic device connected to the arbitration device. This first posterior probability represents the probability that the device is the target device. The electronic device corresponding to the highest first posterior probability is output as the target device. Then, the output electronic device is compared with the labeled target device. If the output electronic device is not the labeled target device, the parameters of the first semantic reasoning model to be trained are adjusted to reduce the first posterior probability of the output electronic device. Then, the electronic device with the highest first posterior probability is compared with the labeled target device again until the electronic device with the highest first posterior probability output by the first semantic reasoning model to be trained is the labeled target device.

[0133] For example, if the electronic device with the highest first posterior probability is not the labeled target device, the loss function can be 1; if the electronic device with the highest first posterior probability is the labeled target device, the loss function can be 0.

[0134] S540, Send the device control command to the target device.

[0135] Once the target device is identified, the arbitration device can send the device control command converted from the voice control command to the target device, so that the target device can provide the corresponding service to execute the control command.

[0136] For example, in a voice control scenario, when user A says "Play the news broadcast on TV", the arbitration device receives the voice data, parses it to find the target device name, and then directly outputs the device control command corresponding to the specific device to the TV. After receiving the device control command, the TV provides the corresponding function service to execute the control command.

[0137] For example, if user B is an unregistered user, when user B says "play the news broadcast", since most people are used to watching the news broadcast on TV in their daily lives, the arbitration device, according to the second semantic reasoning model, determines that the target device for executing user B's "play the news broadcast" is the TV.

[0138] For example, both User A and User B are registered users. If User A prefers to receive services via mobile phone and User B prefers to receive services via television, when User A and User B say "play the news broadcast" in the same scenario, the arbitration device determines, based on User A's first semantic reasoning model, that the target device for executing User A's "play the news broadcast" is a mobile phone; and based on User B's first semantic reasoning module, determines that the target device for executing User B's "play the news broadcast" is a television.

[0139] The following example illustrates the decision-making process for the target device, using the arbitration device as the server, the first type of electronic device as a smart speaker, smartphone, and smart TV, and the target device as a smart TV.

[0140] Please see Figure 6a , Figure 6a This is a flowchart illustrating another voice control method provided in an embodiment of this application. Figure 6a As shown, the user inputs voice data, including the voice command "play the news broadcast." A smart speaker with its intelligent voice assistant enabled receives this voice data and uploads it to the server. Next, the server identifies the voice data to determine if it includes the name of the target device. If it does, the server directly converts the voice command into a device control command and sends it to the target device. If the voice command does not include the target device name, the server determines, based on a semantic reasoning model, that the electronic device executing the voice command is a smart TV. The server then converts the voice command into a device control command and sends it to the smart TV. This device control command is used to control the smart TV to provide corresponding functional services. Finally, upon receiving the device control command, the smart TV uses the video playback service to execute the "play the news broadcast" operation.

[0141] As can be seen, the voice control method proposed in this application involves an electronic device receiving voice data from a target user. The voice data includes voice control instructions, which are used to instruct the electronic device to execute operation commands. The method then determines whether the voice data includes the name of the target device. If the voice data does, the voice control instructions are converted into device control instructions. Otherwise, the target user's identity information is obtained from the voice data, and a semantic reasoning model is invoked to determine the target device based on the identity information. Finally, the device control instructions are sent to the target device. This application uses a semantic reasoning model to determine the target device for executing the voice control instructions, thereby enabling the selection of the execution device based on the user's intent and improving the user experience.

[0142] It is understood that, in order to achieve the above-mentioned functions, electronic devices include hardware and / or software modules that perform the respective functions. Based on the algorithmic steps of the examples described in the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in conjunction with the embodiments, but such implementation should not be considered beyond the scope of this application.

[0143] This embodiment can divide the electronic device into functional modules according to the above method example. For example, each function can be divided into its own functional modules, or two or more functions can be integrated into one processing module. The integrated modules can be implemented in hardware. It should be noted that the module division in this embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.

[0144] When dividing each function into modules according to its corresponding function. Figure 7 A schematic diagram of a voice control device is shown, such as Figure 7 As shown, the voice control device 700 is applied to an arbitration device, and the voice control device 700 may include a transceiver unit 701 and a processing unit 702.

[0145] The transceiver unit 701 can be used to support electronic devices in performing the above-described S510, S540, etc., and / or other processes used in the technology described herein.

[0146] The processing unit 702 can be used to support electronic devices in performing the above-described S520, S530, etc., and / or other processes used in the technology described herein.

[0147] It should be noted that all relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.

[0148] The electronic device provided in this embodiment is used to execute the above-described voice control method, and therefore can achieve the same effect as the above-described implementation method.

[0149] When using integrated units, the electronic device may include a processing module, a storage module, and a communication module. The processing module can be used to control and manage the operation of the electronic device; for example, it can support the electronic device in executing the steps performed by the transceiver unit 701 and the processing unit 702. The storage module can support the electronic device in executing stored program code and data. The communication module can support communication between the electronic device and other devices.

[0150] The processing module can be a processor or a controller. It can implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor can also be a combination of functions that implement computing capabilities, such as a combination of one or more microprocessors, a combination of digital signal processing (DSP) and a microprocessor, etc. The storage module can be a memory. The communication module can specifically be a radio frequency circuit, a Bluetooth chip, a Wi-Fi chip, or other devices that interact with other electronic devices.

[0151] In one embodiment, when the processing module is a processor and the storage module is a memory, the electronic device involved in this embodiment can be a device having... Figure 1 The device with the structure shown.

[0152] This embodiment also provides a computer storage medium storing computer instructions. When the computer instructions are executed on an electronic device, the electronic device performs the aforementioned method steps to implement the voice control method in the above embodiment.

[0153] This embodiment also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned steps to implement the voice control method described in the above embodiment.

[0154] In addition, embodiments of this application also provide an apparatus, which may specifically be a chip, component or module. The apparatus may include a connected processor and a memory; wherein the memory is used to store computer execution instructions, and when the apparatus is running, the processor may execute the computer execution instructions stored in the memory to cause the chip to execute the voice control methods in the above-described method embodiments.

[0155] In this embodiment, the electronic device, computer storage medium, computer program product or chip are all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects of the corresponding method provided above, and will not be repeated here.

[0156] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0157] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0158] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0159] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0160] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, in essence, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0161] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A voice control method, characterized in that, Applied to arbitration equipment, the method includes: Receive voice data from a target user, the voice data including voice control instructions, the voice control instructions being used to instruct the electronic device to execute operation commands; Determine whether the voice data includes the name of the target device; When the voice data includes the name of the target device, the voice control command is converted into a device control command; otherwise, the target user's identity information is obtained based on the voice data, and the semantic reasoning model is invoked based on the identity information to determine the target device. Send the device control command to the target device; The identity information includes registered users and unregistered users; The step of obtaining the target user's identity information based on the voice data includes: Extract the voiceprint feature information corresponding to the speech data; The voiceprint feature information is matched with at least one pre-stored voiceprint feature information respectively; If the voiceprint feature information matches the pre-stored target voiceprint feature information, the identity information of the target user is determined to be the registered user, and the target voiceprint feature information is any one of the at least one voiceprint feature information. If the voiceprint feature information does not match any of the pre-stored voiceprint feature information, the identity information of the target user is determined to be the unregistered user; The semantic reasoning model includes a first semantic reasoning model and a second semantic reasoning model. The first semantic reasoning model is trained using registered user corpus as training samples, and the second semantic reasoning model is trained using general corpus as training samples. The step of determining the target device by invoking a semantic reasoning model based on the identity information includes: If the user is a registered user, the voice control command is input into the first semantic reasoning model to obtain the target device; each registered user corresponds to a first semantic reasoning model, and each first semantic reasoning model is trained based on multiple semantic control commands of the corresponding registered user; If the user is an unregistered user, the voice control command is input into the second semantic reasoning model to obtain the target device; all unregistered users correspond to the second semantic reasoning model, which is trained from multiple semantic control commands commonly used in daily life.

2. The method according to claim 1, characterized in that, The training methods for the first semantic reasoning model include: Obtain a training dataset, which includes multiple voice data points from the registered user; Perform feature extraction on the multiple voice data to obtain multiple audio features; The multiple audio features are input into the first semantic reasoning model to be trained for training until the training termination condition is met, thus obtaining the first semantic reasoning model.

3. The method according to claim 2, characterized in that, The acquisition of the training dataset includes: Based on the mapping relationship between the voiceprint feature information and the registered user, the target registered user corresponding to the target voiceprint feature information is determined; Acquire multiple raw voice data, including the voice control commands of the target registered user; Determine the device that executes the voice control commands of the target registered user; The execution device is labeled as the target device of the voice control command, and the multiple voice data are obtained.

4. The method according to claim 2, characterized in that, The step of inputting the multiple audio features into the first semantic reasoning model to be trained for training until the training termination condition is met, thereby obtaining the first semantic reasoning model, includes: The multiple audio features are input into the first semantic reasoning model to be trained to obtain the output device corresponding to each piece of speech data. Construct a loss function based on the output device and the labeled target device; The parameters corresponding to minimizing the loss function are updated to the parameters of the first semantic reasoning model to be trained, thus obtaining the first semantic reasoning model.

5. A voice control device, characterized in that, Applied to arbitration equipment, the device includes: A transceiver unit is used to receive voice data from a target user, the voice data including voice control instructions, the voice control instructions being used to instruct an electronic device to execute an operation command; The processing unit is configured to determine whether the voice data includes the name of the target device; if the voice data includes the name of the target device, the voice control command is converted into a device control command; otherwise, the target user's identity information is obtained based on the voice data, and the semantic reasoning model is invoked based on the identity information to determine the target device. The transceiver unit is also used to send the device control command to the target device; The identity information includes registered users and unregistered users; The step of obtaining the target user's identity information based on the voice data includes: Extract the voiceprint feature information corresponding to the speech data; The voiceprint feature information is matched with at least one pre-stored voiceprint feature information respectively; If the voiceprint feature information matches the pre-stored target voiceprint feature information, the identity information of the target user is determined to be the registered user, and the target voiceprint feature information is any one of the at least one voiceprint feature information. If the voiceprint feature information does not match any of the pre-stored voiceprint feature information, the identity information of the target user is determined to be the unregistered user; The semantic reasoning model includes multiple first semantic reasoning models and second semantic reasoning models. The first semantic reasoning model is trained using registered user corpus as training samples, and the second semantic reasoning model is trained using general corpus as training samples. In determining the target device by invoking a semantic reasoning model based on the identity information, the processing unit is specifically used for: If the user is a registered user, the first semantic reasoning model corresponding to the target voiceprint feature information is determined according to the mapping relationship between the voiceprint feature information and the first semantic reasoning model. The voice control command is input into the first semantic reasoning model to obtain the target device. Each registered user corresponds to one first semantic reasoning model, and each first semantic reasoning model is trained based on multiple semantic control commands of the corresponding registered user. If the user is an unregistered user, the voice control command is input into the second semantic reasoning model to obtain the target device; all unregistered users correspond to the second semantic reasoning model, which is trained from multiple semantic control commands commonly used in daily life.

6. An electronic device, characterized in that, The method includes a processor, a memory, a communication interface, and one or more programs, said programs being stored in the memory and configured to be executed by the processor, said programs including instructions for performing the steps of the method as described in any one of claims 1-4.

7. A computer-readable storage medium, characterized in that, A computer program for storing electronic data interchange is provided, wherein the computer program causes a computer to perform the method as described in any one of claims 1-4.

Citation Information

Patent Citations

  • System and method for achieving intelligent home device control on smart watch

    CN103730116A

  • Method and device for controlling household electrical appliance through sound box, and sound box

    CN106886166A

  • Context-based device arbitration

    CN111344780A

  • Voice instruction identification method and device, storage medium and electronic device

    CN112116910A