Methods and related devices for distance relationship determination, device control, and model training
By acquiring the sound acquisition data of multiple devices and using the distance comparison model to determine the relative distance relationship between the device and the sound source target, the accuracy and efficiency of distance relationship determination in multi-device voice interaction are solved, and more efficient device wake-up and control are achieved.
Patent Information
- Application Number
- CN202110273250.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-10
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2041-03-10
AI Technical Summary
In the prior art, in multi-device voice interaction scenarios, it is difficult to efficiently and accurately determine the relative distance relationship between the sound source target and the device, resulting in insufficient accuracy and efficiency of the nearest wake-up device.
By acquiring the sound acquisition data of multiple devices, using a pre-trained distance comparison model, the relative distance identification between the device and the sound source target is determined, and the device distance relationship sequence is formed, and the relative distance relationship prediction is achieved.
The prediction efficiency and applicability of relative distance relationships in multi-device voice interaction scenarios are improved, and the hardware requirements and algorithm complexity of conventional microphone array sound source positioning algorithms are overcome, achieving higher accuracy and convenience.
Smart Images

Figure CN115083436B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of voice assistants, and specifically relates to a method and related device for determining distance relationships, device control, and model training. Background Art
[0002] With the advent of the third wave of artificial intelligence technology, voice assistants have gradually entered all aspects of life and are installed on intelligent devices such as mobile phones, watches, speakers, and TVs. Due to the rich variety of devices, there may be multiple devices with voice assistant functions in the same space.
[0003] Currently, in the near wake-up product solutions, most use microphone arrays to measure the distance of the sound source, and then determine the device closest to the sound source by comparing distances, and wake up the device to execute user instructions. Summary of the Invention
[0004] This application provides a method and related device for determining distance relationships, device control, and model training, aiming to improve the comprehensiveness, efficiency, and application convenience of the arbitration device in calculating the distance between the sound source target and the device in the near wake-up product solution.
[0005] In a first aspect, this application provides a method for determining distance relationships, which is applied to an arbitration device. The method includes:
[0006] Obtain a plurality of sound collection data corresponding to a plurality of devices one by one. Each sound collection data in the plurality of sound collection data includes reference voice data obtained by the corresponding device collecting the sound of the sound source target and the device identifier of the device;
[0007] According to the reference voice data and device identifiers in the plurality of sound collection data and a pre-trained distance comparison model, determine at least one relative distance identifier corresponding to at least one device in the plurality of devices. Each relative distance identifier in the at least one relative distance identifier is used to indicate the position of the corresponding device in the device distance relationship sequence. The device distance relationship sequence is a sequence formed by sorting the plurality of devices according to a preset sorting strategy for distances. The distance refers to the distance between the device and the sound source target.
[0008] It can be seen that in this example, the arbitration device first obtains multiple voice collection data of multiple devices. Secondly, according to the reference voice data and device identifiers in the multiple voice collection data and a pre-trained distance comparison model, at least one relative distance identifier corresponding to at least one of the multiple devices is determined. Since each relative distance identifier is used to indicate the position of the corresponding device in the device distance relationship sequence, and the device distance relationship sequence is a sequence formed by sorting the multiple devices according to a preset sorting strategy of distances, and the distance refers to the distance between the device and the sound source target, it can be seen that the model can predict the relative position of the distance between the device and the sound source target in the distance sequence. For example, if the device distance relationship sequence formed by devices 1, 2, and 3 in the order of increasing distance from the sound source target is device 3 → device 2 → device 1, the prediction result can be indicated by the relative distance identifier of device 3 being 1 to indicate that the distance between device 3 and the sound source target is the closest. Compared with the existing solution of the model independently predicting the absolute distance, the present application realizes predicting the relative distance relationship including global information in the multi-device voice interaction scenario through the distance comparison model. At the same time, since each voice collection data includes the reference voice data obtained by the corresponding device collecting the sound of the sound source target and the device identifier of the device, it can be seen that there is no limitation on the number of channels for collecting the reference voice data, so the hardware requirements and algorithm complexity problems of the conventional microphone array sound source localization algorithm can be overcome, which is beneficial to improving the efficiency and applicability of predicting the relative distance relationship.
[0009] In a second aspect, the present application provides a device control method applied to a target device. The method includes:
[0010] Obtaining indication information of an arbitration device, where the indication information is generated by the arbitration device when determining that the target device among the multiple devices is used to execute a voice command associated with the sound of a sound source target according to at least one relative distance identifier corresponding to at least one of the multiple devices. The at least one relative distance identifier is obtained by the arbitration device performing the following operations: obtaining multiple voice collection data corresponding to the multiple devices, where each voice collection data in the multiple voice collection data includes reference voice data obtained by the corresponding device collecting the sound of the sound source target and the device identifier of the device; and determining the at least one relative distance identifier according to the reference voice data and device identifiers in the multiple voice collection data and a pre-trained distance comparison model. Each relative distance identifier in the at least one relative distance identifier is used to indicate the position of the corresponding device in the device distance relationship sequence, and the device distance relationship sequence is a sequence formed by sorting the multiple devices according to a preset sorting strategy of distances, and the distance refers to the distance between the device and the sound source target;
[0011] Performing an operation indicated by the voice command associated with the sound of the sound source target according to the indication information.
[0012] It can be seen that in this example, the target device first obtains the indication information of the arbitration device. Secondly, it performs the operations indicated by the voice commands associated with the sound source target according to the indication information. Since the indication information is generated by the arbitration device when determining the target device among multiple devices for executing the voice commands associated with the sound source target based on at least one relative distance identifier corresponding one-to-one to at least one of the multiple devices, and the relative distance identifier is used to indicate the position of the corresponding device from the sound source target in the distance sequence, and the distance sequence is used to indicate the sequence formed by sorting multiple distances according to a preset sorting strategy, and the multiple distances include the distances between each of the multiple devices and the sound source target. Compared with the existing solution of determining the nearby wake-up device based on the absolute distance between the device and the sound source target, this application realizes predicting the relative distance relationship containing global information in the multi-device voice interaction scenario through a distance comparison model, and determining the target device to be woken up according to this relative distance relationship. At the same time, since each sound acquisition data includes the reference voice data obtained by the corresponding device collecting the sound of the sound source target and the device identifier of the device, it can be seen that there is no limitation on the number of channels for collecting the reference voice data. Therefore, it can overcome the problems of high hardware requirements and complex algorithms of the conventional microphone array sound source localization algorithm, which is beneficial to improving the efficiency and applicability of predicting the relative distance relationship.
[0013] In a third aspect, this application provides a training method for a distance comparison model, including:
[0014] Obtain training data, where the training data includes multiple voice data sets. Each voice data set in the multiple voice data sets contains multiple reference voice data corresponding one-to-one to multiple devices. Each reference voice data in the multiple reference voice data is the voice data obtained by the corresponding device collecting the sound of the sound source target, and the multiple voice data sets correspond to the voice data sets collected in different sound acquisition environments. The sound acquisition environment at least includes the position where the sound source target is located;
[0015] Train a preset distance comparison model according to the reference voice data of the multiple voice data sets and a preset loss function to obtain a trained distance comparison model. The loss function is used to characterize the loss of the distance comparison model from the dimension of the prediction accuracy of the relative distance relationship between two devices in the device pairing group and the sound source target in the same sound acquisition environment. The device pairing group is composed of any two devices among the multiple devices.
[0016] It can be seen that in this example, the device first obtains training data. Secondly, the preset distance comparison model is trained according to the reference voice data of the multiple voice data sets and the preset loss function, and the trained distance comparison model is obtained. Since the training data includes multiple voice data sets, each voice data set contains multiple reference voice data corresponding one-to-one to multiple devices, each reference voice data is the voice data obtained by the corresponding device collecting the sound of the sound source target, and the multiple voice data sets correspond to the voice data sets collected in different sound collection environments. At the same time, the loss function is used to characterize the loss of the distance comparison model from the dimension of the prediction accuracy of the relative distance relationship between two devices in the device pairing group and the sound source target in the same sound collection environment. The device pairing group is composed of any two devices among the multiple devices. Therefore, the distance prediction model has the ability to predict the relative distance relationship between the device and the sound source target, and there is no limit to the number of channels of the reference voice data. Therefore, it can overcome the problems of high hardware requirements and complex algorithms of the conventional microphone array sound source localization algorithm, which is beneficial to improving the efficiency and applicability of relative distance relationship prediction.
[0017] In a fourth aspect, the present application provides a distance relationship determination device, which is applied to an arbitration device. The device includes:
[0018] An acquisition unit, configured to obtain multiple sound collection data corresponding one-to-one to multiple devices, where each sound collection data in the multiple sound collection data includes reference voice data obtained by the corresponding device collecting the sound of the sound source target and the device identifier of the device;
[0019] A determination unit, configured to determine at least one relative distance identifier corresponding one-to-one to at least one device among the multiple devices according to the reference voice data and device identifier in the multiple sound collection data and a pre-trained distance comparison model. Each relative distance identifier in the at least one relative distance identifier is used to indicate the position of the corresponding device in the device distance relationship sequence. The device distance relationship sequence is a sequence formed by sorting the multiple devices according to a preset sorting strategy for distances. The distance refers to the distance between the device and the sound source target.
[0020] In a fifth aspect, the present application provides a device control device, which is applied to an electronic device. The device includes:
[0021] An acquisition unit, configured to acquire indication information of an arbitration device, where the indication information is generated by the arbitration device when determining a voice command for a target device among the multiple devices to perform voice association of a sound source target according to at least one relative distance identifier corresponding to at least one of the multiple devices one by one, and the at least one relative distance identifier is obtained by the arbitration device performing the following operations: acquiring multiple sound collection data corresponding to the multiple devices one by one, where each sound collection data in the multiple sound collection data includes reference voice data obtained by the corresponding device collecting the sound of the sound source target and the device identifier of the device; and determining the at least one relative distance identifier according to the reference voice data and device identifiers in the multiple sound collection data and a pre-trained distance comparison model, where each relative distance identifier in the at least one relative distance identifier is used to indicate the position of the corresponding device in a device distance relationship sequence, and the device distance relationship sequence is a sequence formed by sorting the multiple devices according to a preset sorting strategy of distances, and the distance refers to the distance between the device and the sound source target;
[0022] An execution unit, configured to execute an operation indicated by the voice command for voice association of the sound source target according to the indication information.
[0023] In a sixth aspect, the present application provides a training device for a distance comparison model, including:
[0024] An acquisition unit, configured to acquire training data, where the training data includes multiple voice data sets, and each voice data set in the multiple voice data sets contains multiple reference voice data corresponding to multiple devices one by one, each reference voice data in the multiple reference voice data is voice data obtained by the corresponding device collecting the sound of the sound source target, and the multiple voice data sets correspond to voice data sets collected in different sound collection environments, and the sound collection environment at least includes the position where the sound source target is located;[[ID=IO]]
[0025] A training unit, configured to train a preset distance comparison model according to the reference voice data in the multiple voice data sets and a preset loss function to obtain a trained distance comparison model, where the loss function is used to characterize the loss of the distance comparison model from the dimension of the prediction accuracy of the relative distance relationship between two devices in a device pairing group and the sound source target in the same sound collection environment, and the device pairing group is composed of any two devices among the multiple devices.
[0026] In a seventh aspect, the present application provides an electronic device, including one or more processors;
[0027] One or more memories, configured to store programs,
[0028] The one or more memories and the program are configured so that the one or more processors control the electronic device to execute instructions such as the steps in any method of the first aspect, the second aspect, or the third aspect of the embodiments of the present application.
[0029] In an eighth aspect, the present application provides a chip, comprising: a processor for calling and running a computer program from a memory, so that a device equipped with the chip executes part or all of the steps described in any method of the first aspect, second aspect, or third aspect of the embodiments of the present application.
[0030] In the ninth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program for electronic data exchange, wherein the computer program enables a computer to execute part or all of the steps described in any method of the first aspect, second aspect, or third aspect of the embodiments of the present application.
[0031] In a tenth aspect, the present application provides a computer program, wherein the computer program is operable to cause a computer to execute some or all of the steps described in any of the methods of the first, second, or third aspects of the embodiments of the present application. The computer program can be a software installation package. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0033] Figure 1a This is a schematic diagram of user control in a multi-device scenario provided by an embodiment of the present application;
[0034] Figure 1b This is an architecture diagram of a device control system 10 provided in an embodiment of the present application;
[0035] Figure 1c This is a functional interface diagram of an intelligent voice assistant provided in an embodiment of the present application;
[0036] Figure 1d This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application;
[0037] Figure 2 1 is a flow chart of a distance relationship determination method provided in an embodiment of the present application;
[0038] Figure 3It is a schematic flowchart of a device control method provided by an embodiment of the present application;
[0039] Figure 4 It is a schematic flowchart of a training method for a distance comparison model provided by an embodiment of the present application;
[0040] Figure 5 It is a block diagram of the functional units of a distance relationship determination device provided by an embodiment of the present application;
[0041] Figure 6 It is a block diagram of the functional units of another distance relationship determination device provided by an embodiment of the present application;
[0042] Figure 7 It is a block diagram of the functional units of a device control device provided by an embodiment of the present application;
[0043] Figure 8 It is a block diagram of the functional units of another device control device provided by an embodiment of the present application;
[0044] Figure 9 It is a block diagram of the functional units of a training device for a distance comparison model provided by an embodiment of the present application;
[0045] Figure 10 It is a block diagram of the functional units of another training device for a distance comparison model provided by an embodiment of the present application. Detailed implementation manners
[0046] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0047] The terms "first", "second", etc. in the specification and claims of the present application and the above accompanying drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0048] Reference to "embodiment" in this document means that the specific features, structures, or characteristics described in connection with the embodiment may be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0049] Currently, as Figure 1a shown, in the space where the user is located, there are a smart speaker (0.5 m away from the user), a smart TV 1 (0.6 m away from the user), a computer (1.2 m away from the user), and a smart TV 2 (0.55 m away from the user). When the user wants to listen to music and issues the command "Play music", the current intelligent voice assistant can only measure the distance between the device containing the microphone array and the sound source. If the computer does not contain a microphone array, the intelligent voice assistant cannot calculate the distance and thus cannot accurately implement the control of waking up nearby devices.
[0050] In view of the above problems, embodiments of the present application provide a method and related device for distance relationship determination, device control, and model training, which will be described in detail below with reference to the accompanying drawings.
[0051] Please refer to Figure 1b , Figure 1b which is a device control system 10 provided by an embodiment of the present application. The device control system 10 includes an electronic device 100 with sound collection capabilities (such as a smart TV, a smart speaker, a smart phone, etc.), an arbitration device 200 installed with an intelligent voice assistant, and a server 300. The arbitration device can be any one of the electronic devices 100, or any one of the mobile devices, such as the user's mobile phone, or a dedicated control box in the smart home scenario, or a server in the cloud, or a device group composed of multiple devices that jointly complete data processing. The arbitration device 200 is communicatively connected to both the electronic device 100 and the server 300 to form a device control network in the smart home scenario.
[0052] Among them, the intelligent voice assistant can be installed on various devices such as mobile phones to support the device control method of the present application. The specific function names and interface interaction methods it exhibits can be diverse and are not uniquely limited here. For example, it is installed on a mobile phone and presents a setting function interface of the "Breeno" intelligent assistant as shown in Figure 1c . The legend includes the function settings of one-key commands, specifically including the function of navigating home, the nearby function, the reminder function when arriving home, taking screenshots with the case, and the multi-device and control function. Among the graphic labels of the multi-device and control function, the digital labels on the devices can be used to identify the distance between the device and the sound source target, that is, the user.
[0053] It should be noted that, as the policy execution device in the embodiment of the present application, the arbitration device 200 can have various data and signaling interaction methods with other devices (such as the electronic device 100 and the server 300), and there is no unique limitation here. For example, the arbitration device 200 can directly connect to the electronic device 100 to obtain corresponding information, and the arbitration device 200 can connect to the server 300 through a mobile communication network to achieve corresponding information interaction, etc.
[0054] Please refer to Figure 1d , Figure 1d which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device is applied to the above device control system 10. The electronic device includes an application processor 120, a memory 130, a communication module 140, and one or more programs 131. The application processor 120 is communicatively connected to the memory 130 and the communication module 140 through an internal communication bus.
[0055] Among them, the one or more programs 131 are stored in the above-mentioned memory 130 and are configured to be executed by the above-mentioned application processor 120. The one or more programs 131 include instructions for executing any step in the above method embodiment.
[0056] Among them, the application processor 120 can be, for example, a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, units, and circuits described in connection with the disclosure of the present application. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and so on. The communication unit can be the communication module 140, a transceiver, a transceiver circuit, etc., and the storage unit can be the memory 130.
[0057] The memory 130 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0058] In a specific implementation, the application processor 120 is configured to execute any step performed by the arbitration device or the target device or the model training device in the method embodiment of the present application.
[0059] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of a method for determining a distance relationship provided by an embodiment of the present application, and is applied to the arbitration device 200 in the above device control system 10. As shown in the figure, the device control method includes the following operations.
[0060] Step 201: Obtain a plurality of voice acquisition data corresponding to a plurality of devices one by one. Each voice acquisition data in the plurality of voice acquisition data includes reference voice data obtained by the corresponding device collecting the sound of the sound source target and the device identifier of the device.
[0061] Among them, the plurality of devices may be the electronic devices 100 in the above device control system 10, and the quantity and device type are not uniquely limited.
[0062] Among them, the sound source target includes a user or a sound-producing device, and there is no unique limitation here. The sound of the sound source target can be a wake-up voice, such as "Hello, Xiaoou", etc. In addition, in a heterogeneous distributed scenario, each device can wait for the wake-up voice simultaneously. When the user speaks the wake-up voice, the device obtains the monophonic speech segments of their respective wake-up voices.
[0063] Among them, the reference speech data can be first subjected to a 4KHz low-pass filter to suppress the non-human voice audio part.
[0064] In specific implementation, the reference speech data can be the speech data before frequency response ability alignment and / or feature extraction and feature fusion processing. In this case, the arbitration device uniformly performs relevant preprocessing on the reference speech data of multiple devices.
[0065] In addition, the reference speech data can also be the speech data after each device has independently performed frequency response ability alignment and / or feature extraction and feature fusion processing. In this case, the arbitration device no longer performs unified preprocessing, but after obtaining the reference speech data and device identifiers of multiple devices, it calls a pre-trained distance comparison model to predict at least one relative distance identifier of at least one device.
[0066] Step 202: According to the reference speech data, device identifiers in the multiple sound acquisition data, and the pre-trained distance comparison model, determine at least one relative distance identifier corresponding to at least one device among the multiple devices. Each relative distance identifier in the at least one relative distance identifier is used to indicate the position of the corresponding device in the device distance relationship sequence. The device distance relationship sequence is a sequence formed by sorting the multiple devices according to a preset sorting strategy for distances. The distance refers to the distance between the device and the sound source target.
[0067] Among them, the preset sorting strategy can include from large to small or from small to large, etc., and there is no unique limitation here.
[0068] Among them, the reference speech data includes monophonic speech data or multichannel speech data, that is, there is no high requirement for the speech acquisition ability of the device, and the application convenience is higher.
[0069] Among them, the relative distance identifier can be a number (such as 1 / 2 / 3 / 4), a graph (such as line segments with different lengths), etc., and there is no unique limitation here.
[0070] For example, assume that multiple devices include Device A, Device B, Device C, and Device D, and the device distance relationship sequence is Device A → Device D → Device B → Device C, and the current distance sorting relationship is from far to near, that is, Device A has the closest distance to the sound source target, and Device C has the farthest distance to the sound source target. Then the prediction result can be: the relative distance identifier of Device A is 1 (closest), the relative distance identifier of Device B is 3, the relative distance identifier of Device C is 4 (farthest), and the relative distance identifier of Device D is 2.
[0071] In specific implementation, the arbitration device can directly select the target device with the closest relative distance to the sound source target as the wake-up device and execute the user's voice command through this target device. In addition, for the special case where the prediction results of two devices are the same in the distance comparison result, the arbitration device can further query the preset device wake-up priority set to determine the device to be woken up first. This can improve the accuracy and success rate of device control.
[0072] In specific implementation, if the input of the distance comparison model only contains audio data information, the output of the distance comparison model is the relative distance identifier of each voice data. The arbitration device can further determine the relative distance identifier of the corresponding device according to the correspondence between the voice data and the device identifier. In this case, the device identifier only needs to be the device identifier corresponding to the voice data in the prediction result. That is, the voice data corresponds to the device identifier, the model input data does not include the device identifier, the prediction result corresponds to the voice data one by one, and the prediction result corresponds to the device identifier indirectly through the voice data.
[0073] If the input of the distance comparison model contains audio data information and device identifiers, such as Mobile Phone 1, Mobile Phone 2, Mobile Phone 3, the device types of Mobile Phone 1 and Mobile Phone 2 are Type 1, and the device type of Mobile Phone 3 is Type 2. Then the device identifier of Mobile Phone 1 can be Type 1 + Mobile Phone 1 name, the identifier of Mobile Phone 2 is Type 1 + Mobile Phone 2 name, and the device identifier of Mobile Phone 3 is Type 2 + Mobile Phone 3 name. Then the output of the distance comparison model is directly the relative distance identifier of each device identifier, that is, the prediction result can be directly corresponding to the device identifier.
[0074] In a possible example, determining at least one relative distance identifier corresponding to at least one device among the multiple devices according to the reference voice data and device identifier in the multiple voice acquisition data and a pre-trained distance comparison model includes: performing frequency response ability alignment on each of the multiple reference voice data in the multiple voice acquisition data to obtain multiple target voice data after frequency response ability alignment corresponding to the multiple voice acquisition data; determining at least one relative distance identifier corresponding to at least one device among the multiple devices according to the multiple target voice data, the device identifier in the multiple voice acquisition data, and the pre-trained distance comparison model.
[0075] It can be seen that in this example, due to the different performance of sensor devices in different types and models of devices, there are significant differences in the collected activation voice signals at the same distance. Therefore, through frequency response ability alignment processing, the voice signals obtained by heterogeneous devices are adapted to the same standard, providing a unified data basis for distance comparison. Compared with the existing simple normalization scheme only through the voice energy ratio, it is beneficial to improve the accuracy of the model prediction result.
[0076] In this possible example, performing frequency response ability alignment on the multiple reference voice data in the multiple voice acquisition data to obtain multiple target voice data after frequency response ability alignment includes: performing the following operations on each of the multiple reference voice data in the multiple voice acquisition data to obtain multiple target voice data after frequency response ability alignment: obtaining the preset frequency response unit impulse response of the current device relative to the reference device according to the device identifier of the currently processed reference voice data; performing convolution operation on the reference voice data and the frequency response unit impulse response to perform gain adjustment to obtain the target voice data after frequency response ability alignment.
[0077] In specific implementation, if the current device and the reference device are the same device, the voice after the reference voice data is adjusted remains unchanged. The reference device can be a specified device among the multiple devices.
[0078] Among them, the frequency response unit impulse response of the current acquisition device relative to the reference device is determined through the following steps a1 to a5:
[0079] Step a1, assume that the total number of device types for which perceptual ability alignment is performed is L, and L is a positive integer. Place multiple devices at the same distance from the sound source, play a 0 - 8KHz sweep signal, and record the frequency response curves of each device. Record N times at the same distance, and each device obtains N frequency response curves. Perform M batches of acquisitions under different sound source distance conditions (such as 0.5m, 0.8m, 1m, 1.2m, 1.5m, 1.8m, 2m, 2.2m, 2.5m, 2.8m, 3m). Denote is the frequency response curve recorded for the j-th time of the i-th batch of the l-th device, where l = 1, 2, …, L; i = 1, 2, …, M; j = 1, 2, …, N, and L, M, and N are positive integers.
[0080] Step a2, calculate the statistical average of all the frequency response curves of each device as the final frequency response curve of the device:
[0081]
[0082] where K is the number of points of the frequency response curve, which is a positive integer.
[0083] Step a3, select a device as the reference device, denoted as type b ∈ [1, 2, …, L]. The selection method can be the device that the user uses most frequently, etc. Calculate the ratio of the frequency response curves of each device to the reference device to obtain the frequency response transfer function between devices:
[0084]
[0085] If l = b, the reference device does not need to align its capabilities with itself.
[0086] Step a4, perform a discrete fast Fourier inverse transform (IDFT) on the frequency response transfer function to obtain the frequency response unit impulse response corresponding to the transfer function:
[0087]
[0088] In the formula j is the symbol of the imaginary part of the complex number.
[0089] Step a5, save the frequency response unit impulse response t , ,
[0088] , , , , ,
[0089] , l ,
[0087] , ,
[0090] ,
[0091] (n) in the corresponding device.
[0090] It can be seen that in this example, due to the different performances of the sensor devices of different types and models of devices, there are significant differences in the collected voice signals activated at the same distance. Therefore, the frequency response unit impulse response is obtained through the frequency response curve, and the signals obtained from heterogeneous devices are adapted to the same standard, providing a unified data basis for distance comparison. Compared with the existing simple normalization scheme only through the voice energy ratio, it is beneficial to improve the accuracy of the model prediction results.
[0091] In a possible example, determining at least one relative distance identifier corresponding to at least one device among the multiple devices according to the target voice data, the device identifier in the multiple voice acquisition data, and the pre-trained distance comparison model includes: performing multi-dimensional feature extraction on each of the multiple target voice data to obtain multiple voice feature sets corresponding to the multiple target voice data one by one, where each voice feature set includes multi-dimensional feature extraction results; performing feature fusion on each of the multiple voice feature sets to obtain multiple fused voice features; and determining at least one relative distance identifier corresponding to at least one device among the multiple devices according to the multiple fused voice features, the device identifier in the multiple voice acquisition data, and the pre-trained distance comparison model.
[0092] In this possible example, performing multi-dimensional feature extraction on each of the multiple target voice data to obtain multiple voice feature sets corresponding to the multiple target voice data one by one includes: performing the following operations on each of the multiple target voice data to obtain multiple voice feature sets: extracting the scalar voice feature and the vector voice feature of the currently processed target voice data; performing dimensionality reduction and secondary feature extraction on the vector voice feature to obtain vector-derived voice features.
[0093] In this possible example, extracting the scalar voice feature and the vector voice feature of the currently processed target voice data includes: preprocessing the currently processed target voice data to obtain the preprocessed target voice data; and extracting the scalar voice feature and the vector voice feature of the preprocessed target voice data; where the preprocessing includes at least one of the following: voice activity detection (VAD) processing, pre-emphasis processing through a high-frequency filter, frame segmentation processing, and windowing processing.
[0094] Among them, the purpose of the voice activity detection (VAD) processing is to identify and eliminate long silent periods therefrom so as to intercept the most effective voice segments for use in subsequent steps. For example, the WebRTC VAD algorithm can be used. The purpose of the pre-emphasis processing is to compensate for the high-frequency components lost by the voice signal due to the influence of the pronunciation system and highlight the high-frequency formants. Pre-emphasis is achieved through a high-frequency filter, and its transfer function is:
[0095] H(z) = 1 - μz -1 ,
[0096] Among them, z is the Z-transform of the voice signal, and μ is the coefficient of pre-emphasis, generally between 0.9 and 1.0. In this application, for example, it can be 0.97.
[0097] In addition, for ease of processing, based on the short-term stationarity of the speech signal, it is necessary to perform frame segmentation on the speech signal. In this solution, the frame length is set to 25 ms and the frame shift is set to 10 ms, that is, there is an overlapping area of 15 ms between two adjacent frames, so as to avoid the influence caused by excessive changes in the speech signals of two adjacent frames. For the convenience of expression, let s i [n] represent the data of the i-th frame.
[0098] After frame segmentation, in order to eliminate the possible signal discontinuity at both ends of each frame and prevent spectral leakage, windowing processing needs to be performed to obtain s i,w [n] = s i [n] × w[n], where w[n] represents the window function. In the embodiment of the present application, the Hamming window is adopted, and its formula is as follows:
[0099]
[0100] where N is the length of s i [n].
[0101] In specific implementation, the above scalar speech features can extract 5 scalar speech features, as shown in Table 1. The scalar speech features are extracted according to their definitions, and the specific calculation process will not be elaborated.
[0102] Table 1 Scalar Speech Features
[0103] Feature Type Chinese Explanation English Explanation LP Linear Prediction Linear Prediction LPRR LP Residual Peak-to-Root Mean Square Ratio LP Residual Ratio LPRK LP Residual Kurtosis LP Residual Kurtosis LPRHP LP Residual Histogram Peak LP Residual Histogram Peak SPSK Spectrogram Skewness Spectrogram Skewness SHPP Spectrogram Histogram Peak Position Spectrogram Histogram Peak Position
[0104] The above scalar speech features can all be eigenvalue in the form of (1, ), mainly depicting the characteristics of the user's current sound field environment, such as reverberation, reflection, transmission gain, noise, etc. Here, the form of (1, ) represents the data dimension, (1, ) represents a 1×1 dimensional vector, (245, ) represents a 245×1 dimensional vector, where the 245×1 vector is composed of several previous features spliced together. Example: several features: (a, ), (b, ),...., finally spliced into (a + b +...,, ) = (245, ). This representation is for 1 piece of speech data, and obtaining (245, ) means a 245×1 feature vector. In the actual calculation process, multiple pieces of speech will be processed in batches, denoted as N pieces, where N is a positive integer. Then each feature correspondingly becomes each batch of features, and the dimension changes: (1, ) → (1, N), that is, 1×N dimension, and similarly (245, ) → (245, N), that is, 245×N dimension.
[0105] In specific implementation, the above vector speech features can extract 8 vector speech features, as shown in Table 2. The vector speech features are extracted respectively according to the existing methods, and the specific calculation process will not be elaborated.
[0106] Table 2 Vector Speech Features
[0107] Feature Type Chinese Explanation English Explanation MFCC Mel-Frequency Cepstral Coefficients Mel-Frequency Cepstral Coefficients LPCC Linear Predictive Cepstral Coefficients Linear Predictive Cepstral Coefficients MHEC Mean Hilbert Envelope Coefficients Mean Hilbert Envelope Coefficients BFCC Bark-Frequency Cepstral Coefficients Bark-Frequency Cepstral Coefficients LFCC Linear-Frequency Cepstral Coefficients Linear-Frequency Cepstral Coefficients GFCC Gammatone-Frequency Cepstral Coefficients Gammatone-Frequency Cepstral Coefficients NGCC Normalized Gammachirp-Frequency Cepstral Coefficients NorGammachirp-Frequency Cepstral Coefficients MSRCC Magnitude-based Spectral Root Cepstral Coefficients Magnitude-based Spectral Root Cepstral Coefficients
[0108] The above vector voice features are all feature vectors in the form of (12, N f ), where 12 is the dimension of the feature, indicating that the vector voice features all have 12-dimensional feature components, and N f represents the number of frames of the voice. The vector voice features can depict the voice pickup characteristics of different voice pickup devices, such as spectrum, pitch, formant, etc.
[0109] In addition, the above 8 kinds of vector voice features have a large number of parameters and are not convenient to be directly combined with the scalar features for use. They can be further processed through vector derivative feature extraction. Vector derivative feature extraction includes three functional modules: feature component screening, differential feature calculation, and vector feature scalar quantization, which will be introduced separately below.
[0110] (1) Feature component screening module
[0111] For the above 8 kinds of vector voice features, they are all feature vectors in the form of (12, N f ), with a large number of parameters and not convenient to be directly used. In addition, the contribution degrees of each dimensional feature component to model training are different, and some contain less information, while some may contain redundant information. Therefore, it is necessary to screen the feature components of the vector voice features. This solution uses the Fisher criterion to evaluate the discrimination ability of each dimensional feature component. The Fisher discrimination criterion is as follows:
[0112]
[0113] where r Fisher is the Fisher ratio of the feature component. The larger this value is, the stronger the discrimination ability of this dimensional feature component; σ b represents the between-class variance of the feature component, that is, the variance of the means of the voice feature components of different distance types; σ w represents the within-class variance of the feature component, that is, the variance of the means of the voice feature components of the same distance type. The calculation formulas of σ b and σ w are as follows:
[0114]
[0115]
[0116] In the above two formulas: M represents the number of samples, represents the mean value of the k-th dimensional component of a certain feature vector of the i-th voice sample, and m k represents the mean value of the k-th dimensional component of a certain feature vector of all samples, and n iDenote the number of frames of a certain speech sample. Denote the feature value of the k-th dimension and the c-th frame of the i-th speech for a certain feature.
[0117] As described above, the feature component screening module uses the Fisher criterion to separately select 3 feature components with the largest Fisher ratio from the 8 types of vector speech features. In this way, each type of vector speech feature is transformed from (12, N f ) to (3, N f ), greatly reducing the number of feature parameters.
[0118] (2) Differential feature calculation module
[0119] The above 8 types of vector speech features can only reflect the static characteristics of speech, and their dynamic characteristics can be described by the differential parameters of vector speech features. The formula for calculating the differential parameters is as follows:
[0120]
[0121] Among them, c l Denote the l-th dimensional feature component of a certain vector speech feature. Denote the first-order difference value of the l-th dimensional feature component of the vector speech feature at the t-th frame. Θ is a constant, and this value can be used to represent that the size of the differential window is 2Θ + 1. In this scheme, Θ = 2. Through this formula, the first-order differences of the above 8 types of vector speech features can be obtained, which are respectively denoted as ΔMFCC, ΔLPCC, ΔMHEC, ΔBFCC, ΔLFCC, ΔGFCC, ΔNGCC, ΔMSRCC. These vector differential features are feature vectors in the form of (3, N f ).
[0122] (3) Vector feature quantization module
[0123] For the above 8 types of vector speech features such as MFCC and 8 types of vector differential features such as ΔMFCC, the number of frames N of the speech with a long duration f is often very large, that is, the dimension of the feature in the "frame" direction is still very high; in addition, the number of frames N of speeches with different durations f is different, which is not convenient for training machine learning models. To solve these two problems, this scheme uses the vector feature quantization module to perform quantization processing on the above vector features. Taking MFCC as an example below, the specific process of vector feature quantization is described in detail:
[0124] Step a: Use GMM and its EM algorithm to cluster each dimensional feature component of MFCC. The number of clusters is 4, and 4 clustering center values of each dimensional feature component are obtained, resulting in a feature in the form of (4, ), denoted as F1 l , where l ∈ (1, 2, 3) represents the label of the feature component;
[0125] Step b, calculate the maximum value, minimum value, and end value of each dimensional feature component of MFCC, obtaining a feature in the form of (3,), denoted as
[0126] Step c, calculate the maximum value, minimum value, and sum of squares of each dimensional feature component of the vector difference feature ΔMFCC of the MFCC, obtaining a feature in the form of (3,), denoted as F3 l ;
[0127] Step d, splice the three feature vectors F1 l , F2 l , F3 l obtained in steps a, b, and c, obtaining a feature in the form of (10,), denoted as F l ;
[0128] Step e, splice the respective features F l of the feature components again, obtaining a new feature in the form of (30,), denoted as F MFCC , which is used to represent the original MFCC feature;
[0129] Step f, perform operations on the other 7 types of vector speech features using the above steps a - step e., and the respective corresponding new features can be obtained, denoted as F LPCC , F MHEC , F BFCC , F BFCC , F LFCC , F GFCC , F NGCC , F MSRCC ;
[0130] Step g, splice and fuse the above 8 new features in the form of (30,), and a vector-derived speech feature in the form of (240,) can be obtained, denoted as F D , F D Each eigenvalue in has a specific physical meaning.
[0131] In specific implementation, the above feature fusion steps are used to fuse scalar speech features and vector-derived speech features. First, splice the 5 scalar speech features, obtaining a feature vector in the form of (5,), denoted as F S ; Then splice and fuse F S with the vector-derived feature F D , and finally obtain a fused feature in the form of (245,), denoted as F Fusion , that is, each speech is finally extracted with a feature vector in the form of (245,), which is used for the training of the speaker and heterogeneous device distance relationship model.
[0132] It can be seen that in this example, the fused speech feature is less affected by the characteristics of the environmental sound field and random noise, is applicable to various scenarios and heterogeneous distributed devices, is more robust than single features such as energy and signal-to-noise ratio, and has stronger generalization ability. This fused speech feature can be used for training the distance relationship model between the speaker and the distributed device, so as to judge the distance between different intelligent devices and the user, making the "human-machine distance" an important decision dimension in multi-device wake-up and improving the user experience. Specifically, in the same space, the user has multiple distributed devices that support the same wake-up word. After the user says the wake-up word, the device closest to the user responds to achieve proximity wake-up; or the human-machine distance is combined with other dimensions such as device status, device service ability, and user intention for comprehensive determination to select the most suitable device to respond to the user's request.
[0133] In a possible example, the determining of at least one relative distance identifier corresponding to at least one device among the multiple devices according to the reference speech data and device identifier in the multiple sound acquisition data and the pre-trained distance comparison model includes: performing multi-dimensional feature extraction on each sound acquisition data in the multiple sound acquisition data to obtain multiple speech feature sets corresponding to the multiple sound acquisition data one by one, where each speech feature set in the multiple speech feature sets includes the multi-dimensional feature extraction results; performing feature fusion on each speech feature set in the multiple speech feature sets to obtain multiple fused speech features; and determining at least one relative distance identifier corresponding to at least one device among the multiple devices according to the multiple fused speech features, the device identifiers in the multiple sound acquisition data, and the pre-trained distance comparison model.
[0134] Among them, feature extraction can also be used for other distance relationship determination methods, and the calculation results can be used for other application scenarios, etc.
[0135] In this possible example, the performing of multi-dimensional feature extraction on each sound acquisition data in the multiple sound acquisition data to obtain multiple speech feature sets corresponding to the multiple sound acquisition data one by one includes: performing frequency response ability alignment on each reference speech data in the multiple reference speech data in the multiple sound acquisition data to obtain multiple target speech data after frequency response ability alignment corresponding to the multiple sound acquisition data one by one; and performing multi-dimensional feature extraction on each target speech data in the multiple target speech data to obtain multiple speech feature sets corresponding to the multiple target speech data one by one.
[0136] In this possible example, the frequency response capability alignment for each of the multiple reference voice data in the multiple voice acquisition data to obtain multiple target voice data after frequency response capability alignment corresponding one-to-one to the multiple voice acquisition data includes: performing the following operations for each of the multiple reference voice data in the multiple voice acquisition data to obtain multiple target voice data after frequency response capability alignment corresponding one-to-one to the multiple voice acquisition data: obtaining a preset frequency response unit impulse response of the current device relative to the reference device according to the device identifier associated with the currently processed reference voice data; performing a convolution operation on the reference voice data and the frequency response unit impulse response and performing gain adjustment to obtain the target voice data after frequency response capability alignment.
[0137] In this possible example, the multi-dimensional feature extraction for each of the multiple target voice data to obtain multiple voice feature sets corresponding one-to-one to the multiple target voice data includes: performing the following operations for each of the multiple target voice data to obtain multiple voice feature sets: extracting the scalar voice feature and the vector voice feature of the currently processed target voice data; performing dimensionality reduction and secondary feature extraction on the vector voice feature to obtain a vector-derived voice feature.
[0138] In this possible example, the extraction of the scalar voice feature and the vector voice feature of the currently processed target voice data includes: preprocessing the currently processed target voice data to obtain the preprocessed target voice data; extracting the scalar voice feature and the vector voice feature of the preprocessed target voice data; where the preprocessing includes at least one of the following: silence suppression processing, pre-emphasis processing through a high-frequency filter, frame segmentation processing, and windowing processing.
[0139] It should be noted that the implementation principles of feature extraction, feature fusion, and frequency response capability alignment involved in this branch embodiment are similar to the corresponding contents in the foregoing embodiments, and will not be elaborated here.
[0140] It can be seen that in this example, the arbitration device can first perform feature extraction and feature fusion on the voice data, and optionally further perform frequency response capability alignment on the processed voice data, improving the flexibility of voice data preprocessing.
[0141] In a possible example, after determining at least one relative distance identifier corresponding to at least one device among the multiple devices according to the reference voice data and device identifier in the multiple voice acquisition data and a pre-trained distance comparison model, the method further includes: determining a target device for executing the voice instruction associated with the sound source target among the multiple devices according to the at least one relative distance identifier; if it is detected that the target device is a device other than the arbitration device, sending an indication message to the target device, where the indication message is used to instruct the target device to execute the operation indicated by the voice instruction; if it is detected that the target device is the arbitration device, executing the operation indicated by the voice instruction.
[0142] Among them, the voice instruction associated with the sound source target can be various user instructions such as "play music", and there is no unique limitation here.
[0143] In specific implementation, in the near wake-up scheme, the arbitration device preferentially selects the device with the shortest distance to the sound source target as the target device. That is to say, in this application scenario, the at least one relative distance identifier should at least include the relative distance identifier of the device with the shortest distance to the sound source target. In addition, combined with different application scenarios, the specific manifestation form of the at least one relative distance identifier can be various, and there is no unique limitation here.
[0144] It can be seen that in this example, the arbitration device can intelligently determine the target device for executing the voice instruction associated with the sound source target according to at least one relative distance identifier, improving the convenience and intelligence of device control.
[0145] It can be seen that in the embodiments of the present application, the arbitration device first obtains multiple voice collection data of multiple devices. Secondly, according to the reference voice data and device identifiers in the multiple voice collection data and a pre-trained distance comparison model, at least one relative distance identifier corresponding to at least one of the multiple devices is determined. Since each relative distance identifier is used to indicate the position of the corresponding device in the device distance relationship sequence, and the device distance relationship sequence is a sequence formed by sorting multiple devices according to a preset sorting strategy of distances, and the distance refers to the distance between the device and the sound source target, it can be seen that the model can predict the relative position of the device in the device distance relationship sequence. For example, if the device distance relationship sequence formed by devices 1, 2, and 3 in the order of the distance from the sound source target from near to far is device 3 → device 2 → device 1, the prediction result can be indicated by the relative distance identifier of device 3 being 1 to indicate that the distance between device 3 and the sound source target is the closest. Compared with the existing solution of the model isolatedly predicting the absolute distance, the present application realizes predicting the relative distance relationship containing global information in the multi-device voice interaction scenario through the distance comparison model. At the same time, since each voice collection data includes the reference voice data obtained by the corresponding device collecting the sound of the sound source target and the device identifier of the device, it can be seen that there is no restriction on the number of channels for collecting the reference voice data, so the hardware requirements and algorithm complexity problems of the conventional microphone array sound source localization algorithm can be overcome, which is beneficial to improving the efficiency and applicability of the relative distance relationship prediction.
[0146] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of a device control method provided by an embodiment of the present application, and is applied to a target device in a device control system 10. As shown in the figure, the present device control method includes the following operations.
[0147] Step 301: Obtain the indication information of the arbitration device. The indication information is generated by the arbitration device when determining the voice command for the target device among the multiple devices to perform voice association of the sound source target according to at least one relative distance identifier corresponding to at least one of the multiple devices. The at least one relative distance identifier is obtained by the arbitration device performing the following operations: obtaining multiple sound collection data corresponding to the multiple devices, where each sound collection data in the multiple sound collection data includes reference voice data obtained by the corresponding device collecting the sound of the sound source target and the device identifier of the device; and determining the at least one relative distance identifier according to the reference voice data and device identifier in the multiple sound collection data and a pre-trained distance comparison model. Each relative distance identifier in the at least one relative distance identifier is used to indicate the position of the corresponding device in the device distance relationship sequence, and the device distance relationship sequence is a sequence formed by sorting the multiple devices according to a preset sorting strategy of distances, and the distance refers to the distance between the device and the sound source target.
[0148] Step 302: Execute the operation indicated by the voice command for voice association of the sound source target according to the indication information.
[0149] Among them, the target device is the device closest to the sound source target among the multiple devices. This situation applies to the near wake-up product solution.
[0150] Among them, the target device is the arbitration device; or, the target device is a device other than the arbitration device among the multiple devices.
[0151] In a specific implementation, if the target device is the arbitration device, the arbitration device directly generates indication information and executes the operation indicated by the voice command for voice association of the sound source target according to the indication information.
[0152] If the target device is a device other than the arbitration device among the multiple devices, the arbitration device generates indication information, sends the indication information to the target device, and the target device executes the operation indicated by the voice command for voice association of the sound source target according to the indication information.
[0153] In addition, the method further includes: if it is detected that the distance between the target device and the sound source target is greater than a preset distance, output a prompt message to prompt the user to approach the target device; and / or, if it is detected that the distance between the target device and the sound source target is greater than the preset distance, increase the output volume of the target device.
[0154] If it is detected that the distance between the target device and the sound source target is less than or equal to the preset distance, the output volume of the target device is turned down. This can improve the intelligence of device control and enhance the user experience.
[0155] Among them, the preset distance can be, for example, 5 meters, 10 meters, etc.
[0156] It can be seen that in the embodiment of the present application, the target device first obtains the indication information of the arbitration device. Secondly, it performs the operation indicated by the voice command associated with the sound source target according to the indication information. Since the indication information is generated by the arbitration device when determining the target device among multiple devices for executing the voice command associated with the sound source target according to at least one relative distance identifier corresponding to at least one device among the multiple devices, and the relative distance identifier is used to indicate the position of the corresponding device from the sound source target in the distance sequence, and the distance sequence is used to indicate a sequence formed by sorting multiple distances according to a preset sorting strategy, and the multiple distances include the distances between each device among the multiple devices and the sound source target. Compared with the existing solution of determining the nearby wake-up device based on the absolute distance between the device and the sound source target, the present application realizes predicting the relative distance relationship including global information in the multi-device voice interaction scenario through a distance comparison model, and determining the target device to be woken up according to this relative distance relationship. At the same time, since each sound collection data includes the reference voice data obtained by the corresponding device collecting the sound of the sound source target and the device identifier of the device, it can be seen that there is no limitation on the number of channels for collecting the reference voice data. Therefore, it is possible to overcome the problems of high hardware requirements and complex algorithms of the conventional microphone array sound source localization algorithm, which is beneficial to improving the efficiency and applicability of predicting the relative distance relationship.
[0157] Please refer to Figure 4 , Figure 4 is a schematic flowchart of a method for training a distance comparison model provided by an embodiment of the present application, which is applied to a model training device, such as Figure 4 shown, the device control method includes the following operations.
[0158] Step 401, obtain training data, where the training data includes multiple voice data sets, each voice data set in the multiple voice data sets contains multiple reference voice data corresponding to multiple devices, each reference voice data in the multiple reference voice data is voice data obtained by the corresponding device collecting the sound of the sound source target, and the multiple voice data sets correspond to voice data sets collected in different sound collection environments, and the sound collection environment at least includes the position where the sound source target is located.
[0159] The sound collection environment refers to the acoustic environment for collecting sound data. The sound collection environment can be diversified. In addition to the location of the sound source target, a differentiated sound collection environment can be further constructed by the difference of at least one of the following characteristics: room area, noise level, etc.
[0160] For example, the area is 10m 2 ,20m 2 ,30m 2 ,40m 2 The system collects wake-up speech in both quiet and noisy environments. Noise can be added manually, and options include Gaussian white noise from the noise database, electrical noise (such as fans and air conditioners), and traffic noise. Depending on the sound pressure level of the pure wake-up speech, the signal-to-noise ratio can be set to -15dB, -10dB, -5dB, 0dB, 5dB, 10dB, or 15dB.
[0161] Step 402: Train a preset distance comparison model based on the reference speech data of the multiple speech data sets and a preset loss function to obtain a trained distance comparison model. The loss function is used to characterize the loss of the distance comparison model from the dimension of the prediction accuracy of the relative distance relationship between two devices in a device pairing group and the sound source target under the same sound collection environment. The device pairing group is composed of any two devices from the multiple devices.
[0162] Among them, the distance comparison model can specifically be a deep neural network, and the deep neural network can be, for example, a convolutional neural network or a deep residual network, etc., which is not limited here.
[0163] In one possible example, the relative distance relationship is characterized by defining a score for an event in which the first device of the two devices is closer to the sound source target than the second device, and the value of the score is associated with a distance difference, where the distance difference is the difference between a first distance and a second distance, the first distance being the distance between the first device and the sound source target, and the second distance being the distance between the second device and the sound source target.
[0164] The score can be calculated and expressed in a probability-like data format, and the value range of the score falls within the interval (0, 1).
[0165] For example, if the first distance is 50 cm, the second distance is 80 cm, and the third distance (the distance between the third device and the sound source target) is 90 cm, then the score for the event that the first device is closer to the sound source target than the second device can be 0.8, the score for the event that the first device is closer to the sound source target than the third device can be 0.9, the score for the event that the second device is closer to the sound source target than the third device can be 0.7, the score for the event that the second device is closer to the sound source target than the first device can be 0.08, the score for the event that the third device is closer to the sound source target than the first device can be 0.05, the score for the event that the third device is closer to the sound source target than the second device can be 0.09, and so on.
[0166] It can be seen that in this example, due to the relative distance relationship between the devices and the sound source, and this relationship is strong or weak, by associating the value of the score for the event that the first device is closer to the sound source target than the second device with the distance difference between the first distance and the second distance, the distance comparison model can learn this difference rather than just learning the coarse-grained relationship of far / near.
[0167] In a possible example, the score is calculated from at least one score of at least one group of adjacent devices that form the direct or indirect adjacent relationship between the two devices; the score of the adjacent devices is calculated from two relative distance identifiers of two devices in the adjacent devices, the relative distance identifier corresponds to the prediction result of the distance comparison model, the relative distance identifier is used to indicate the position of the corresponding device in the device distance relationship sequence, and the device distance relationship sequence is a sequence formed by sorting the multiple devices according to a preset sorting strategy of distances, and the distance refers to the distance between the device and the sound source target.
[0168] It can be seen that in this example, since the relative distance identifier can indicate the position of the corresponding device in the device distance relationship sequence, the device distance relationship sequence is a sequence formed by sorting the multiple devices according to a preset sorting strategy of distances, and the distance refers to the distance between the device and the sound source target, so the relative distance identifier can represent the model prediction result containing global information and improve the accuracy of the model prediction result representation.
[0169] In this possible example, training the preset distance comparison model according to the reference speech data of the multiple speech data sets and a preset loss function to obtain a trained distance comparison model includes: dividing the training data into a training set and a test set, where the training set includes some of the multiple speech data sets; using the training set to train the preset distance comparison model at least once until the accuracy of the distance comparison result of the test set predicted by the trained distance comparison model is greater than the preset accuracy.
[0170] Among them, the preset accuracy can be, for example, 98%, 99%, etc., and there is no unique limitation here.
[0171] It can be seen that in this example, by training the distance comparison model, the model is enabled to have the prediction ability that meets the requirements of the preset accuracy.
[0172] In this possible example, the training includes forward propagation and backpropagation optimization; in the forward propagation, the predicted relative distance identifier is calculated using the speech features of the speech data set; in the backpropagation optimization, the predicted score and the true score are calculated using the predicted relative distance identifier and the true relative distance identifier, and the loss of the distance comparison model is calculated using the loss function, the predicted score and the true score, and the parameters of the distance comparison model are adjusted according to the loss of the distance comparison model.
[0173] In a specific implementation, the design of this loss function is realized through the following steps a to e:
[0174] Step a, for the data collected for the same set of wake-up actions in the training data, any two devices can form a pair. Without loss of generality, assume that the current set contains the data of 5 devices, and record their types and relative distance relationships as: Among them Indicates that device a is closer to the sound source target than device b, then the relative distance identifiers of the devices are L a =1, L b =2, L c =3, L d =4, L e =5.
[0175] Step b, record the feature vectors extracted by each device as x a , x b , x c , x d , x e , and record the mapping of the feed-forward process of the deep neural network as f. Then, the output layer result corresponding to the feature vector (taking x a as an example) is o a =f(x a ).
[0176] Step c, a score can be obtained for the pairing of two devices from the reference speech data:
[0177]
[0178] Step d: For the true labels, the scores between any two paired devices are still calculated. Since there are relative relationships among the distances between multiple devices, first, the scores between adjacent paired devices are calculated, and based on this, the scores between non - adjacent devices are calculated.
[0179] For adjacent devices (taking b, c as an example), the score is:
[0180]
[0181] For non - adjacent devices (taking a, c as an example), b is the common adjacent device between them, and the score is:
[0182]
[0183] Furthermore, for the device pair (a, e):
[0184]
[0185] Step e: After the input data undergoes forward propagation (feedforward) through the deep neural network, backpropagation is performed according to the loss function between the actual output and the true labels, and the network parameters are iteratively adjusted to improve the network performance.
[0186] Taking any device pair (i, j) as an example, the loss function is calculated as follows:
[0187]
[0188] It can be seen that in this example, the loss function can quantitatively measure the difference between the estimated label and the true label of the voice data of any two devices in the current device set after passing through the distance comparison model, and adjust the parameters of the distance comparison model through this difference until the model prediction accuracy meets the requirements.
[0189] It can be seen that in the embodiments of the present application, the device first obtains training data. Secondly, a preset distance comparison model is trained according to the reference voice data of the multiple voice data sets and a preset loss function, and a trained distance comparison model is obtained. Since the training data includes multiple voice data sets, each voice data set contains multiple reference voice data corresponding one-to-one to multiple devices, each reference voice data is voice data obtained by the corresponding device collecting the sound of the sound source target, and the multiple voice data sets correspond to voice data sets collected in different sound collection environments. At the same time, the loss function is used to characterize the loss of the distance comparison model from the dimension of the prediction accuracy of the relative distance relationship between two devices in the device pairing group and the sound source target in the same sound collection environment. The device pairing group is composed of any two devices among the multiple devices. Therefore, the distance prediction model has the ability to predict the relative distance relationship between the device and the sound source target, and there is no limit on the number of channels of the reference voice data. Therefore, it is possible to overcome the problems of high hardware requirements and complex algorithms of conventional microphone array sound source localization algorithms, which is beneficial to improving the efficiency and applicability of relative distance relationship prediction.
[0190] The embodiments of the present application provide a distance relationship determination device, and the distance relationship determination device can be an arbitration device. Specifically, the distance relationship determination device is used to execute the steps performed by the arbitration device in the above distance relationship determination method. The distance relationship determination device provided by the embodiments of the present application may include modules corresponding to the respective steps.
[0191] The embodiments of the present application can divide the function modules of the distance relationship determination device according to the above method examples. For example, each function can be corresponding to each function module, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software function modules. The division of modules in the embodiments of the present application is illustrative, and is only a logical function division. There may be other division methods in actual implementation.
[0192] In the case of dividing each function module corresponding to each function, Figure 5 A possible structural schematic diagram of the distance relationship determination device involved in the above embodiments is shown. As Figure 5 shown, the distance relationship determination device 5 is applied to the arbitration device 200 in the device control system 10; the device includes:
[0193] An acquisition unit 50, configured to obtain multiple sound collection data corresponding one-to-one to multiple devices, and each sound collection data in the multiple sound collection data includes reference voice data obtained by the corresponding device collecting the sound of the sound source target and the device identifier of the device;
[0194] A determining unit 51, configured to determine at least one relative distance identifier corresponding to at least one device among the multiple devices according to reference voice data and device identifiers in the multiple voice acquisition data and a pre-trained distance comparison model, where each relative distance identifier in the at least one relative distance identifier is used to indicate the position of the corresponding device in a device distance relationship sequence, and the device distance relationship sequence is a sequence formed by sorting the multiple devices according to a preset sorting strategy of distances, and the distance refers to the distance between the device and the sound source target.
[0195] In a possible example, in terms of determining at least one relative distance identifier corresponding to at least one device among the multiple devices according to reference voice data and device identifiers in the multiple voice acquisition data and a pre-trained distance comparison model, the determining unit 51 is specifically configured to perform frequency response ability alignment on each reference voice data in the multiple reference voice data in the multiple voice acquisition data to obtain multiple target voice data after frequency response ability alignment corresponding to the multiple voice acquisition data; and determine at least one relative distance identifier corresponding to at least one device among the multiple devices according to the multiple target voice data, the device identifiers in the multiple voice acquisition data, and the pre-trained distance comparison model.
[0196] In a possible example, in terms of performing frequency response ability alignment on the multiple reference voice data in the multiple voice acquisition data to obtain multiple target voice data after frequency response ability alignment, the determining unit 51 is specifically configured to: perform the following operations on each reference voice data in the multiple voice acquisition data to obtain multiple target voice data after frequency response ability alignment: obtain a preset current frequency response unit impulse response of the current device relative to a reference device according to the device identifier of the currently processed reference voice data; perform a convolution operation on the reference voice data and the frequency response unit impulse response to perform gain adjustment to obtain the target voice data after frequency response ability alignment.
[0197] In a possible example, when determining, according to the target voice data, the device identifier in the multiple voice acquisition data, and the pre-trained distance comparison model, at least one relative distance identifier corresponding to at least one device among the multiple devices, the determining unit 51 is specifically configured to: perform multi-dimensional feature extraction on each of the multiple target voice data to obtain multiple voice feature sets corresponding to the multiple target voice data one by one, where each voice feature set includes multi-dimensional feature extraction results; perform feature fusion on each of the multiple voice feature sets to obtain multiple fused voice features; and determine, according to the multiple fused voice features, the device identifier in the multiple voice acquisition data, and the pre-trained distance comparison model, at least one relative distance identifier corresponding to at least one device among the multiple devices.
[0198] In a possible example, when determining, according to the reference voice data and device identifier in the multiple voice acquisition data, and the pre-trained distance comparison model, at least one relative distance identifier corresponding to at least one device among the multiple devices, the determining unit 51 is specifically configured to: perform multi-dimensional feature extraction on each of the multiple voice acquisition data to obtain multiple voice feature sets corresponding to the multiple voice acquisition data one by one, where each voice feature set in the multiple voice feature sets includes multi-dimensional feature extraction results; perform feature fusion on each of the multiple voice feature sets to obtain multiple fused voice features; and determine, according to the multiple fused voice features, the device identifier in the multiple voice acquisition data, and the pre-trained distance comparison model, at least one relative distance identifier corresponding to at least one device among the multiple devices.
[0199] In a possible example, when performing multi-dimensional feature extraction on each of the multiple voice acquisition data to obtain multiple voice feature sets corresponding to the multiple voice acquisition data one by one, the determining unit 51 is specifically configured to: perform frequency response ability alignment on each of the multiple reference voice data in the multiple voice acquisition data to obtain multiple target voice data after frequency response ability alignment corresponding to the multiple voice acquisition data one by one; and perform multi-dimensional feature extraction on each of the multiple target voice data to obtain multiple voice feature sets corresponding to the multiple target voice data one by one.
[0200] In a possible example, in terms of aligning the frequency response capabilities of each of the multiple reference speech data among the multiple voice acquisition data to obtain multiple target voice data after frequency response capability alignment corresponding one by one to the multiple voice acquisition data, the determining unit 51 is specifically configured to: perform the following operations for each of the multiple reference speech data among the multiple voice acquisition data to obtain multiple target voice data after frequency response capability alignment corresponding one by one to the multiple voice acquisition data: and obtain the preset current frequency response unit impulse response of the current device relative to the reference device according to the device identifier associated with the currently processed reference speech data; and perform a convolution operation on the reference speech data and the frequency response unit impulse response to perform gain adjustment to obtain the target voice data after frequency response capability alignment.
[0201] In a possible example, in terms of extracting multi-dimensional features of each of the multiple target voice data to obtain multiple voice feature sets corresponding one by one to the multiple target voice data, the determining unit 51 is specifically configured to: perform the following operations for each of the multiple target voice data among the multiple target voice data to obtain multiple voice feature sets: extract the scalar voice feature and the vector voice feature of the currently processed target voice data; and perform dimensionality reduction and secondary feature extraction on the vector voice feature to obtain a vector-derived voice feature.
[0202] In a possible example, in terms of extracting the scalar voice feature and the vector voice feature of the currently processed target voice data, the determining unit 51 is specifically configured to: preprocess the currently processed target voice data to obtain the preprocessed target voice data; and extract the scalar voice feature and the vector voice feature of the preprocessed target voice data; where the preprocessing includes at least one of the following: silence suppression processing, pre-emphasis processing through a high-frequency filter, frame segmentation processing, and windowing processing.
[0203] In a possible example, after the determining unit determines at least one relative distance identifier corresponding one by one to at least one of the multiple devices according to the reference speech data, the device identifier, and the pre-trained distance comparison model among the multiple voice acquisition data, it is further configured to: determine the target device for executing the voice command associated with the sound source target among the multiple devices according to the at least one relative distance identifier; if it is detected that the target device is a device other than the arbitration device, send an indication message to the target device, where the indication message is used to instruct the target device to perform the operation indicated by the voice command; if it is detected that the target device is the arbitration device, perform the operation indicated by the voice command.
[0204] In the case of adopting an integrated unit, the structural schematic diagram of another distance relationship determination device provided by an embodiment of the present application is as follows Figure 6 shown. In Figure 6 , the distance relationship determination device 6 includes: a processing module 60 and a communication module 61. The processing module 60 is used to control and manage the actions of the device control device. For example, it obtains the steps executed by the acquisition unit 50 and the determination unit 51, and / or is used to execute other processes of the technologies described herein. The communication module 61 is used to support the interaction between the device control device and other devices. As Figure 6 shown, the distance relationship determination device may further include a storage module 62, and the storage module 62 is used to store the program code and data of the distance relationship determination device.
[0205] Among them, the processing module 60 may be a processor or a controller. For example, it may be a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an ASIC, an FPGA or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logic blocks, modules and circuits described in combination with the disclosure of the present application. The processor may also be a combination that realizes computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, and so on. The communication module 61 may be a transceiver, an RF circuit or a communication interface, etc. The storage module 62 may be a memory.
[0206] Among them, all relevant contents of each scenario involved in the above method embodiment can be cited in the function description of the corresponding functional module, and will not be elaborated here. Both the above distance relationship determination device 5 and the distance relationship determination device 6 can execute the steps executed by the arbitration device in the above Figure 2 shown distance relationship determination method.
[0207] An embodiment of the present application provides a device control device, and this device control device may be an arbitration device. Specifically, the device control device is used to execute the steps executed by the target device in the above device control method. The device control device provided by the embodiment of the present application may include modules corresponding to the corresponding steps.
[0208] The embodiment of the present application can divide the functional modules of the device control device according to the above method example. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. The division of modules in the embodiment of the present application is illustrative, and is only a logical function division. There may be other division methods in actual implementation.
[0209] In the case of dividing each functional module corresponding to each function, Figure 7 A possible structural schematic diagram of the device control device involved in the above embodiment is shown. As Figure 7 shown, the device control device 7 is applied to the target device; the device includes:
[0210] An obtaining unit 70, configured to obtain indication information of an arbitration device, where the indication information is generated by the arbitration device when determining a voice command for sound association of the target device among the multiple devices to execute a sound source target according to at least one relative distance identifier corresponding to at least one device among the multiple devices, and the at least one relative distance identifier is obtained by the arbitration device performing the following operations: obtaining multiple sound collection data corresponding to the multiple devices, where each sound collection data in the multiple sound collection data includes reference voice data obtained by the corresponding device collecting the sound of the sound source target and the device identifier of the device; and determining the at least one relative distance identifier according to the reference voice data and device identifier in the multiple sound collection data and a pre-trained distance comparison model, where each relative distance identifier in the at least one relative distance identifier is used to indicate the position of the corresponding device in the device distance relationship sequence, and the device distance relationship sequence is a sequence formed by sorting the multiple devices according to a preset sorting strategy of distances, and the distance refers to the distance between the device and the sound source target;
[0211] An execution unit 71, configured to execute an operation indicated by the voice command for sound association of the sound source target according to the indication information.
[0212] In a possible example, the target device is the device among the multiple devices that is closest to the sound source target.
[0213] In a possible example, the target device is the arbitration device; or, the target device is a device among the multiple devices other than the arbitration device.
[0214] In the case of adopting an integrated unit, a structural schematic diagram of another device control device provided in an embodiment of the present application is as Figure 8 shown. In Figure 8 , the device control device 8 includes: a processing module 80 and a communication module 81. The processing module 80 is used to control and manage the actions of the device control device. For example, the steps executed by the obtaining unit 70 and the execution unit 71, and / or other processes for executing the technologies described herein. The communication module 81 is used to support the interaction between the device control device and other devices. As Figure 8As shown, the device control device may further include a storage module 82, which is used to store the program code and data of the device control device.
[0215] Among them, the processing module 80 may be a processor or a controller. For example, it may be a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an ASIC, an FPGA or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logic blocks, modules and circuits described in connection with the disclosure of the present application. The processor may also be a combination that realizes computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and so on. The communication module 81 may be a transceiver, an RF circuit or a communication interface, etc. The storage module 82 may be a memory.
[0216] Among them, all the relevant contents of each scenario involved in the above method embodiments can be cited in the function descriptions of the corresponding functional modules, and will not be elaborated here. Both the above device control device 7 and the device control device 8 can execute the Figure 3 steps performed by the target device in the device control method shown above.
[0217] An embodiment of the present application provides a training device for a distance comparison model. The training device for the distance comparison model may be a model training device for training a model. Specifically, the training device for the distance comparison model is used to execute the steps performed by the model training device in the above distance comparison model training method. The training device for the distance comparison model provided by the embodiment of the present application may include modules corresponding to the respective steps.
[0218] Embodiments of the present application can divide the functional modules of the training device for the distance comparison model according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. The division of modules in the embodiments of the present application is illustrative, and is only a logical function division. In actual implementation, there may be other division methods.
[0219] In the case of dividing each functional module corresponding to each function, Figure 9 shows a possible structural schematic diagram of the training device for the distance comparison model involved in the above embodiments. As Figure 9 shown, the training device 9 for the distance comparison model is applied to a model training device; the device includes:
[0220] An acquisition unit 90 is configured to acquire training data, where the training data includes a plurality of speech data sets. Each speech data set in the plurality of speech data sets contains a plurality of reference speech data corresponding to a plurality of devices one by one. Each reference speech data in the plurality of reference speech data is speech data obtained by a corresponding device collecting the sound of a sound source target, and the plurality of speech data sets correspond to speech data sets collected in different sound collection environments. The sound collection environment at least includes the location where the sound source target is located.
[0221] A training unit 91 is configured to train a preset distance comparison model according to the reference speech data of the plurality of speech data sets and a preset loss function, so as to obtain a trained distance comparison model. The loss function is used to characterize the loss of the distance comparison model from the dimension of the prediction accuracy of the relative distance relationship between two devices in the device pairing group and the sound source target in the same sound collection environment. The device pairing group is composed of any two devices in the plurality of devices.
[0222] In a possible example, the relative distance relationship is characterized by defining a score for an event that the first device in the two devices is closer to the sound source target than the second device, and the value of the score is associated with the distance difference. The distance difference is the difference between the first distance and the second distance. The first distance is the distance between the first device and the sound source target, and the second distance is the distance between the second device and the sound source target.
[0223] In a possible example, the score is calculated by at least one score of at least one group of adjacent devices constituting the direct or indirect adjacent relationship between the two devices.
[0224] The score of the adjacent devices is calculated by two relative distance identifiers of two devices in the adjacent devices. The relative distance identifier corresponds to the prediction result of the distance comparison model, and the relative distance identifier is used to indicate the position of the corresponding device in the device distance relationship sequence. The device distance relationship sequence is a sequence formed by sorting the plurality of devices according to a preset sorting strategy of distances. The distance refers to the distance between the device and the sound source target.
[0225] In a possible example, in terms of training the preset distance comparison model according to the reference speech data of the plurality of speech data sets and the preset loss function to obtain a trained distance comparison model, the training unit 91 is specifically configured to: divide the training data into a training set and a test set, where the training set includes some speech data sets in the plurality of speech data sets; and use the training set to train the preset distance comparison model at least once until the accuracy of the distance comparison result of the test set predicted by the trained distance comparison model is greater than a preset accuracy.
[0226] In a possible example, the training includes forward propagation and backpropagation optimization;
[0227] In the forward propagation, the predicted relative distance identifier is calculated using the speech features of the speech data set;
[0228] In the backpropagation optimization, the predicted score and the true score are calculated using the predicted relative distance identifier and the true relative distance identifier, and the loss of the distance comparison model is calculated using the loss function, the predicted score, and the true score. The parameters of the distance comparison model are adjusted according to the loss of the distance comparison model.
[0229] In the case of adopting an integrated unit, the structural schematic diagram of another distance comparison model training device provided by the embodiment of the present application is as Figure 10 shown. In Figure 10 , the distance comparison model training device 10 includes: a processing module 100 and a communication module 101. The processing module 100 is used to control and manage the actions of the distance comparison model training device. For example, it obtains the steps executed by the acquisition unit 90 and the training unit 91, and / or is used to execute other processes of the technologies described herein. The communication module 101 is used to support the interaction between the distance comparison model training device and other devices. As Figure 10 shown, the distance comparison model training device may further include a storage module 102, and the storage module 102 is used to store the program code and data of the distance comparison model training device.
[0230] Among them, the processing module 100 may be a processor or a controller. For example, it may be a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an ASIC, an FPGA or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in combination with the disclosure of the present application. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and so on. The communication module 101 may be a transceiver, an RF circuit, or a communication interface, etc. The storage module 102 may be a memory.
[0231] Among them, all relevant contents of the various scenarios involved in the above method embodiments can be cited in the function descriptions of the corresponding functional modules, and will not be repeated here. Both the above distance comparison model training device 9 and the distance comparison model training device 10 can execute the above Figure 4Steps performed by a model training device in the training method of the distance comparison model shown above.
[0232] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more collections of available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0233] The embodiments of the present application also provide a computer storage medium. The computer storage medium stores a computer program for electronic data exchange, and the computer program causes a computer to execute some or all of the steps of any of the methods described in the above method embodiments. The above computer includes an electronic device.
[0234] The embodiments of the present application also provide a computer program product. The computer program product includes a non-transitory computer-readable storage medium storing a computer program. The computer program is operable to cause a computer to execute some or all of the steps of any of the methods described in the above method embodiments. The computer program product can be a software installation package. The above computer includes an electronic device.
[0235] It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0236] In several embodiments provided by this application, it should be understood that the disclosed methods, devices, and systems can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for example, the division of the units is only a logical function division, and there can be other division methods in actual implementation; for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0237] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0238] In addition, the functional units in each embodiment of the present invention can be integrated in a processing unit, or each unit can be physically included separately, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of a combination of hardware and software functional units.
[0239] The above integrated units implemented in the form of software functional units can be stored in a computer-readable storage medium. The above software functional units stored in a storage medium include several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute some steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0240] Although the present invention is disclosed as above, the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions without departing from the spirit and scope of the present invention, and can make various changes and modifications, including the combination of the above different functions and implementation steps, including software and hardware implementation methods, all within the protection scope of the present invention.
Claims
1. A training method for a distance comparison model, characterized in that: include: Acquiring training data, the training data comprising a plurality of speech data sets, each of the plurality of speech data sets comprising a plurality of reference speech data corresponding one-to-one to a plurality of devices, each of the plurality of reference speech data being speech data obtained by collecting a sound of a sound source target by a corresponding device, and the plurality of speech data sets corresponding to speech data sets collected in different sound collection environments, the sound collection environment at least including a location of the sound source target; A preset distance comparison model is trained based on the reference speech data of the multiple speech data sets and a preset loss function to obtain a trained distance comparison model, wherein the loss function is used to characterize the loss of the distance comparison model from the dimension of prediction accuracy of the relative distance relationship between two devices in a device pairing group and the sound source target under the same sound collection environment, wherein the device pairing group is composed of any two devices from the multiple devices, and the relative distance relationship is characterized by defining a score of an event in which the first device of the two devices is closer to the sound source target than the second device, and the value of the score is associated with a distance difference, wherein the distance difference is the difference between the first distance and the second distance, and the first distance is The distance between the first device and the sound source target, the second distance is the distance between the second device and the sound source target, the score is calculated by at least one score of at least one group of adjacent devices that constitute a direct or indirect adjacent relationship between the two devices; the score of the adjacent device is calculated by two relative distance identifiers of two devices in the adjacent devices, the relative distance identifier corresponds to the prediction result of the distance comparison model, and the relative distance identifier is used to indicate the position of the corresponding device in the device distance relationship sequence. The device distance relationship sequence is a sequence formed by sorting the multiple devices according to a preset sorting strategy of distance. The distance refers to the distance between the device and the sound source target.
2. The method according to claim 1, characterized in that The step of training a preset distance comparison model based on the reference speech data of the plurality of speech data sets and a preset loss function to obtain a trained distance comparison model comprises: Dividing the training data into a training set and a test set, wherein the training set includes a portion of the plurality of speech data sets; The preset distance comparison model is trained at least once using the training set until the accuracy of the distance comparison result of the test set predicted by the trained distance comparison model is greater than a preset accuracy.
3. The method according to claim 2, characterized in that The training includes forward propagation and back propagation optimization; In the forward propagation, the speech features of the speech data set are used to calculate the predicted relative distance identifier; The predicted relative distance identifier and the true relative distance identifier are used in the back-propagation optimization to calculate the predicted score and the true score, and the loss of the distance comparison model is calculated using the loss function, the predicted score and the true score, and the parameters of the distance comparison model are adjusted according to the loss of the distance comparison model.
4. A training device for a distance comparison model, characterized in that: include: an acquisition unit, configured to acquire training data, the training data comprising a plurality of speech data sets, each of the plurality of speech data sets comprising a plurality of reference speech data corresponding one-to-one to a plurality of devices, each of the plurality of reference speech data being speech data obtained by collecting a sound of a sound source target by a corresponding device, and the plurality of speech data sets corresponding to speech data sets collected in different sound collection environments, the sound collection environment at least including a location of the sound source target; A training unit is used to train a preset distance comparison model based on the reference voice data of the multiple voice data sets and a preset loss function to obtain a trained distance comparison model, wherein the loss function is used to characterize the loss of the distance comparison model from the dimension of prediction accuracy of the relative distance relationship between two devices in a device pairing group and the sound source target under the same sound collection environment, wherein the device pairing group is composed of any two devices from the multiple devices, and the relative distance relationship is characterized by defining a score of an event in which the first device of the two devices is closer to the sound source target than the second device, and the value of the score is associated with a distance difference, wherein the distance difference is the difference between the first distance and the second distance, and the first The distance is the distance between the first device and the sound source target, the second distance is the distance between the second device and the sound source target, the score is calculated by at least one score of at least one group of adjacent devices that constitute a direct or indirect adjacent relationship between the two devices; the score of the adjacent device is calculated by two relative distance identifiers of two devices in the adjacent devices, the relative distance identifier corresponds to the prediction result of the distance comparison model, and the relative distance identifier is used to indicate the position of the corresponding device in the device distance relationship sequence. The device distance relationship sequence is a sequence formed by sorting the multiple devices according to a preset sorting strategy of distance. The distance refers to the distance between the device and the sound source target.
5. An electronic device, characterized in that: The electronic device comprises: one or more processors; One or more memories for storing programs, The one or more memories and the program are configured so that the one or more processors control the device to execute the steps of the method according to any one of claims 1 to 3.
6. A computer-readable storage medium, characterized in that A computer program for electronic data exchange is stored, wherein the computer program enables a computer to execute the method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Speaker confirmation method and speaker confirmation device used in short voice condition
CN105845140A
Method and apparatus for calibrating microphones of electronic device, and electronic device
CN106658329A
Method and apparatus for determining sound source distance
CN107507625A
Microphone array correction system and method
CN109951766A
Pickup volume control method and device and storage medium
CN111417053A