Device management method and electronic device

By acquiring and analyzing the voice feature data of electronic devices, the system can identify and control the device closest to the speaker, thus solving the problem of echo signals in remote calls and improving call clarity and quality.

CN119520689BActive Publication Date: 2026-01-16HONOR DEVICE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510031398.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2026-01-16
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

During remote calls, the difference in distance between different electronic devices and the speaker can cause inconsistent sound signal pickup times, resulting in echo signals and affecting call clarity.

Method used

By acquiring voice feature data from multiple electronic devices, the device closest to the speaker is identified as the first electronic device, and it is controlled to send the picked-up sound signal to other devices. Other devices do not send the picked-up sound signal, and the audio output device is controlled not to output the picked-up sound signal to reduce echo.

Benefits of technology

It improves the clarity of remote calls, reduces echo signals, and enhances call quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119520689B_ABST
    Figure CN119520689B_ABST
Patent Text Reader

Abstract

The application discloses a device management method and an electronic device. The method comprises the following steps: acquiring voice feature data of at least two electronic devices in a first space; determining a first electronic device based on the voice feature data of the at least two electronic devices, that is, determining a device used by a speaker; controlling the first electronic device to send a picked-up sound signal to other electronic devices; controlling a second electronic device other than the device used by the speaker to not send a picked-up sound signal to other electronic devices; and controlling an audio output device of the second electronic device to not output a sound signal picked up in a current time window in the first space. In this way, echo signals in a remote call process can be reduced. By using the application, echo signals in a remote call process can be reduced, the echo signals include echo signals in a local space and echo signals in a remote space, and the intelligibility of the remote call is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of terminal, in particular to a device management method and an electronic device. BACKGROUND

[0002] Remote call is widely used in the scene of multi-person voice chat, video chat and the like. In the case that there are multiple electronic devices participating in remote call in the local space, when the user of one of the electronic devices is speaking, the sound will be picked up by the electronic device and also picked up by other electronic devices. Since the distances between the electronic devices and the speaker are different, the time of picking up the sound signal by each electronic device is inconsistent, and the time of receiving the sound signal picked up by each electronic device in the remote space is also different, resulting in echo signals in the remote space. Moreover, the electronic device used by the speaker transmits the picked-up sound signal to other electronic devices in the local space, and the other electronic devices play the sound signal of the speaker. However, since the sound of the speaker will also propagate in the local space through the air, the users in the local space hear multiple same sound signals with time difference, resulting in echo signals in the local space. Therefore, how to solve the echo signals in the process of remote call is a problem to be solved. SUMMARY

[0003] The present application provides a device management method and an electronic device, which can reduce echo signals in the process of remote call, including echo signals in the local space and echo signals in the remote space, thereby improving the clarity of remote call.

[0004] In a first aspect, an embodiment of the present application provides a device management method, which comprises: obtaining voice feature data of at least two electronic devices in a first space; the voice feature data of each electronic device is obtained by performing feature extraction on the picked-up sound signal in a current time window by each electronic device; determining a first electronic device based on the voice feature data of the at least two electronic devices, and controlling the first electronic device to send the picked-up sound signal to other electronic devices, and controlling a second electronic device not to send the picked-up sound signal to other electronic devices; the second electronic device includes electronic devices other than the first electronic device in the first space, and the other electronic devices include electronic devices in the first space and electronic devices in a second space; controlling an audio output device of the second electronic device not to output the sound signal of the first space picked up in the current time window.

[0005] It can be seen that by determining the first electronic device in the first space, the first electronic device can be controlled to send the picked-up sound signal to other electronic devices, so that the sound signal output by the electronic device in the second space is clearer. By controlling the second electronic device not to send the picked-up sound signal to other electronic devices, the problem of echo signal in the second space caused by sending multiple sound signals with time difference to the electronic device in the second space can be avoided. Moreover, by controlling the audio output device of the second electronic device not to output the sound signal picked up in the current time window in the first space, no echo signal will be generated in the first space, thereby solving the echo problem in the first space during remote communication and improving the clarity of remote communication.

[0006] In combination with the first aspect, in a possible implementation, determining the first electronic device based on the voice feature data of the at least two electronic devices includes: performing feature comparison on audio feature information included in the voice feature data of the at least two electronic devices, to determine a first similarity between sound signals picked up by the at least two electronic devices; and in a case where the first similarity is greater than or equal to a first similarity threshold, performing comparison on time indication information included in the voice feature data of the at least two electronic devices, to determine the first electronic device that picks up the earliest sound signal among the at least two electronic devices.

[0007] It can be seen that by comparing the first similarity between sound signals picked up by the at least two electronic devices and comparing the time indication information of the at least two electronic devices, it can be determined whether the sound signals picked up by the at least two electronic devices are the speech content of the same speaker, and it can be determined which electronic device among the multiple electronic devices picks up the earliest sound signal, so as to take the electronic device as the first electronic device. Since the sound signal picked up by the first electronic device is the earliest and is closest to the speaker, the picked-up sound signal is the clearest and has the lowest delay, and after being sent to other electronic devices, the sound signal output by the electronic device in the second space is also clearer, which can improve the clarity of remote communication.

[0008] In combination with the first aspect, in a possible implementation, performing feature comparison on audio feature information included in the voice feature data of the at least two electronic devices to determine a first similarity between sound signals picked up by the at least two electronic devices includes: inputting audio feature information included in the voice feature data of the at least two electronic devices into a voice processing model for feature comparison, to determine the first similarity between sound signals picked up by the at least two electronic devices; wherein the voice processing model is obtained by training an initial model based on a similarity label between at least two sample sound signals and a sample similarity between the at least two sample sound signals, and the sample similarity is obtained by performing feature comparison on sample audio feature information of the at least two sample sound signals based on the initial model.

[0009] It can be seen that, by using the voice processing model to perform feature comparison on the audio feature information included in the voice feature data of the at least two electronic devices, since the voice processing model is obtained by training the initial model, the accuracy of the feature comparison can be improved, and using the voice processing model to perform feature comparison can also improve the efficiency of the feature comparison.

[0010] In combination with the first aspect, in a possible implementation, the method further includes: obtaining audio feature information of the sound signal received by the second electronic device; performing feature comparison on the audio feature information of the sound signal received by the second electronic device and the audio feature information included in the voice feature data of the first electronic device, to determine a second similarity between the sound signal received by the second electronic device and the sound signal picked up by the first electronic device; and in a case where the second similarity is greater than or equal to a second similarity threshold, determining that the sound signal received by the second electronic device is the sound signal of the first space picked up in the current time window, and controlling the audio output apparatus of the second electronic device to not output the received sound signal.

[0011] It can be seen that, by performing feature comparison on the audio feature information of the sound signal received by the second electronic device and the audio feature information included in the voice feature data of the first electronic device to determine the second similarity therebetween, it can be determined whether the sound signal received by the second electronic device is the sound signal of the first space picked up in the current time window, and if so, the audio output apparatus of the second electronic device is controlled to not output the received sound signal, so that only the speech of the speaker that propagates through the air in the first space, and the sound signal that propagates through the network is not output, which can avoid the generation of echo signals in the first space.

[0012] In combination with the first aspect, in a possible implementation, the method further includes: in a case where the second similarity is less than the second similarity threshold, determining that the sound signal received by the second electronic device is the sound signal of the second space picked up in the current time window, and controlling the audio output apparatus of the second electronic device to output the received sound signal.

[0013] It can be seen that, in a case where it is determined that the sound signal received by the second electronic device is the sound signal of the second space picked up in the current time window, the audio output apparatus of the second electronic device is controlled to output the received sound signal, so that the sound signal picked up by the electronic device in the second space is output in the first space, so as to avoid the problem of missing the speech of the second space, such as a remote space, in a remote call.

[0014] In a possible implementation of the first aspect, the method further includes: sending a first control instruction to the first electronic device; the first control instruction is used to control the first electronic device to send the picked-up sound signal to the other electronic device; sending a second control instruction to the second electronic device; the second control instruction is used to control the second electronic device not to send the picked-up sound signal to the other electronic device.

[0015] It can be seen that, by sending the first control instruction to the first electronic device, the first electronic device can send the picked-up sound signal to the other electronic device. By sending the second control instruction to the second electronic device, the second electronic device can not send the picked-up sound signal to the other electronic device. By sending different control instructions to the first electronic device and the second electronic device respectively, the accuracy of controlling each electronic device can be improved, and the situation that the first electronic device and the second electronic device both send the picked-up sound signal to the other electronic device to cause the echo signal in the second space can be avoided.

[0016] In a possible implementation of the first aspect, the method further includes: obtaining an initial working state of the first electronic device in a current time window; the initial working state is used to indicate whether to send the picked-up sound signal to the other electronic device; and in a case where the initial working state of the first electronic device indicates not to send the picked-up sound signal to the other electronic device, sending the first control instruction to the first electronic device.

[0017] It can be seen that, by obtaining the initial working state of the first electronic device in the current time window, it can be determined which control instruction to send to the first electronic device, the control accuracy for the first electronic device is improved, and thus the echo problem in the remote call is reduced.

[0018] In a possible implementation of the first aspect, the method further includes: receiving a third electronic device picked-up sound signal sent by a third electronic device; and in a case where the third electronic device is an electronic device in the second space, controlling the audio output apparatus of the second electronic device to output the third electronic device picked-up sound signal.

[0019] It can be seen that, in a case where the electronic device in the second space picks up the sound signal, the third electronic device picked-up sound signal is controlled to be output by the audio output apparatus of the second electronic device, so as to avoid the problem of missing broadcast of the sound in the remote space in the remote call.

[0020] In a possible implementation of the first aspect, the obtaining of the voice feature data of the at least two electronic devices in the first space includes: establishing a communication connection with the at least two electronic devices in the first space; and receiving the voice feature data transmitted by the at least two electronic devices based on the communication connection.

[0021] It can be seen that by establishing the communication connection with the plurality of electronic devices, the voice feature data transmitted by each electronic device can be received based on the communication connection, and then the first electronic device can be determined based on the voice feature data of each electronic device, thereby improving the accuracy of the determination of the first electronic device.

[0022] In a second aspect, the present application provides an electronic device, which includes one or more processors and one or more memories; the one or more memories are coupled with the one or more processors, and the one or more memories are configured to store computer program codes, the computer program codes including computer instructions, which, when executed by the one or more processors, cause the electronic device to perform the device management method in any possible implementation manner of the first aspect.

[0023] In a third aspect, the present application provides a chip system, which includes a processor and an interface, and the processor and the interface are coupled; the interface is configured to receive or output signals, and the processor is configured to execute code instructions to perform the device management method in any possible implementation manner of the first aspect.

[0024] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the device management method in any possible implementation manner of the first aspect.

[0025] In a fifth aspect, the present application provides a computer program product, which, when running on a computer, causes the computer to perform the device management method in any possible implementation manner of the first aspect. BRIEF DESCRIPTION OF DRAWINGS

[0026] Figure 1a An application scenario schematic diagram provided by an embodiment of the present application is shown in FIG. 1;

[0027] Figure 1b An application scenario schematic diagram provided by an embodiment of the present application is shown in FIG. 2; Figure Two ;

[0028] Figure 2 A flowchart of a device management method provided by an embodiment of the present application is shown in FIG. 3;

[0029] Figure 3 An application scenario schematic diagram provided by an embodiment of the present application is shown in FIG. 4; Figure Three;

[0030] Figure 4 An application scenario schematic diagram for processing echo of a remote space is provided for an embodiment of the present application.

[0031] Figure 5 A flowchart of a device management method is provided for an embodiment of the present application Figure Two ;

[0032] Figure 6 An application scenario schematic diagram for processing echo of a local space is provided for an embodiment of the present application.

[0033] Figure 7 A module interaction flowchart of an electronic device is provided for an embodiment of the present application.

[0034] Figure 8 A hardware structure schematic diagram of an electronic device is provided for an embodiment of the present application.

[0035] Figure 9 A software structure schematic diagram of an electronic device is provided for an embodiment of the present application. DETAILED DESCRIPTION

[0036] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. In the description of the embodiments of the present application, unless otherwise specified, “ / ” represents or, for example, A / B can represent A or B; the “and / or” in the text only represents a description of the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which means that there are three cases of A alone, A and B together, and B alone. In addition, in the description of the embodiments of the present application, “multiple” means two or more than two.

[0037] It should be understood that the terms “first”, “second” and the like in the specification and claims of the present application and the drawings are used to distinguish different objects, and are not used to describe a specific order. In addition, the terms “include” and “have” and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but can optionally include steps or units that are not listed, or can optionally include other steps or units inherent to the process, method, product or device.

[0038] Reference to an "embodiment" in this application means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the application. The appearances of the phrase in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily mutually exclusive of one another. It is expressly understood that the embodiments described in this application can be combined with each other in their various permutations and combinations.

[0039] To facilitate understanding of the schemes provided by the embodiments of the present application, the related concepts involved in the embodiments of the present application are introduced as follows:

[0040] I. Remote call

[0041] The remote call refers to a scenario in which electronic devices located in multiple different spaces participate in the same call, for example, can refer to participating in the same voice chat, video chat, conference, live room, game, and the like. The multiple different spaces of electronic devices participating in the same remote call can realize the interaction between the multiple different spaces of electronic devices. The two devices on the electronic device participating in the remote call are mainly used to realize the pickup and output of the sound signal in the space during the remote call, that is, the audio input device and the audio output device.

[0042] II. Audio input device

[0043] The audio input device is a device for picking up and transmitting the sound signal in the space of the electronic device. The audio input device may, for example, refer to a microphone (Microphone, Mic), and the microphone is also called a sound receiver or a microphone.

[0044] III. Audio output device

[0045] The audio output device is a device for outputting the sound signal received by the electronic device. The audio output device may, for example, refer to a loudspeaker (Loud speakers), and the loudspeaker can also be called a loudspeaker.

[0046] IV. Echo signal of remote space

[0047] In the case that multiple electronic devices exist in the same space, such as the local space, to participate in the remote call, when the user of one of the electronic devices is speaking, the sound will be picked up by the electronic device and also picked up by each of the other electronic devices in the local space. Since the distance between each electronic device in the local space and the speaker is different, the time at which each electronic device picks up the sound signal is inconsistent. And since each electronic device in the local space sends the sound signal picked up to all electronic devices participating in the remote call, including the electronic devices in the remote space and all electronic devices in the local space, through the network, the time at which the electronic devices in the remote space receive the sound signals picked up by each electronic device in the local space is also inconsistent. The electronic devices in the remote space output the received multiple sound signals with time difference, which will generate echo signals in the remote space.

[0048] Fifth, echo signals in the local space

[0049] The electronic device used by the speaker transmits the sound signal picked up to each electronic device in the local space through the network, and the audio output device of each electronic device in the local space outputs the received sound signal of the speaker. Since the sound of the speaker will also propagate in the local space through the air, and the speed of air propagation is greater than the speed of network propagation, multiple same sound signals with time difference will exist in the local space, thus causing echo signals in the local space.

[0050] The application scenarios of the embodiments of the present application are introduced as follows:

[0051] Please refer to Figure 1a , Figure 1a An application scenario schematic diagram one provided by the embodiments of the present application is shown in FIG. 1. Figure 1a As shown in FIG. 1, Figure 1a For example, it can refer to the scenario in which multiple participants use their respective office devices to access the same conference in a collective office scenario, or the scenario in which multiple game players use their respective game devices to join the same game in the same room, the multiple game players in the room access the call system, or the scenario in which the host and multiple audiences join the same live streaming room for voice interaction in the same room in a live streaming scenario, etc. Figure 1aThe remote call process is illustrated by taking a collective office scenario as an example. Spaces where multiple electronic devices participating in the remote call are located can be divided into a remote space and a local space. Electronic devices in the remote space and the local space both participate in the same remote call, such as conference A. When two electronic devices participating in the remote call share the same space, such as space 1, for example, multiple electronic devices exist in the local space, a user of an electronic device 1 in the multiple electronic devices can be defined as a participant 1, and a user of an electronic device 2 can be defined as a participant 2. The remote space is illustrated by taking an electronic device participating in the remote conference as an example. A user of an electronic device in the remote space is defined as a remote participant. Alternatively, multiple electronic devices in the remote space can also participate in conference A. The number of electronic devices in the remote space is not limited in this application. The remote space and the local space can transmit sound signals through a network.

[0052] Please refer to Figure 1b , Figure 1b An application scenario provided for an embodiment of the application is shown in Figure Two . As shown in Figure 1b , in the remote call process, for example, when the participant 2 speaks, as shown in ① in Figure 1b , an audio input device (such as a microphone) on the electronic device 2 used by the participant 2 picks up a sound signal of the participant 2, and as shown in ② in Figure 1b , an audio input device (such as a microphone) on the electronic device 1 of the participant 1 can pick up a sound signal of the local space, so the sound signal of the participant 2 is also picked up. Because the electronic device 1 and the electronic device 2 are different distances from the participant 2, the time when the electronic device 2 picks up the sound signal of the participant 2 is different from the time when the electronic device 1 picks up the sound signal of the participant 2. When the electronic device 1 and the electronic device 2 respectively send the picked sound signals to the electronic device in the remote space through the network, the remote electronic device outputs two sound signals with a time difference through an audio output device (such as a loudspeaker), thereby generating an echo signal in the remote space, which reduces the clarity and quality of the call in the remote space.

[0053] Therefore, in order to solve the problem that the echo in the remote space affects the clarity of the remote call, by determining the device where the speaker in the local space is located in the current time window, for example, when the participant 2 speaks, the electronic device 2 used by the participant 2 is controlled to send the picked sound signal to other electronic devices, and the electronic device 1 is controlled not to send the picked sound signal to other electronic devices. Then the remote electronic device only receives the sound signal sent by the electronic device 2, thereby solving the problem of the echo signal in the remote space and improving the clarity and quality of the call in the remote space in the remote call process. The other electronic devices at least include the electronic device in the remote space, such as the remote electronic device in Figure 1b .

[0054] For the echo problem of the local space in the remote call process, as shown in ③ of Figure 1b , the voice signal of the participant 2 is transmitted through the air to the ears of each participant in the local space, such as the ears of the participant 1 and the participant 2. As shown in ④ of Figure 1b , the voice signal of the participant 2 is picked up by the electronic device 2 used by the participant 2 in the manner of ① and sent to the electronic device 1 through the network. The audio output device (such as a loudspeaker) of the electronic device 1 outputs the voice signal of the participant 2 received by the electronic device 2, and the received voice signal of the participant 2 is played through the loudspeaker in the local space. Since the voice signals output through ③ and ④ have a time difference, each participant in the local space will hear two echo signals with a time difference. Figure 1b Figure 1b Therefore, in order to solve the echo problem of the local space, in the case where the speaker in the current time window in the local space is determined, such as the participant 2 speaking, the electronic device 2 picks up the voice signal and sends it to the electronic device 1, and by controlling the audio output device of the electronic device 1 not to output the received voice signal of the participant 2 picked up by the electronic device 2, only the speaking voice of the participant 2 transmitted through the air is played in the local space, so that the echo problem of the local space can be solved, the call clarity in the local space in the remote call process is improved, and the remote call experience is improved. Figure 1b The following will be described in detail in combination with

[0055] the specific flow of the device management method provided by the embodiments of the present application. Please refer to ,

[0056] the flowchart of the device management method provided by the embodiments of the present application, Figure 2 the main problem to be solved is to solve the echo problem of the remote space in the remote call, as shown in Figure 2 , the method can include but is not limited to the following steps: Figure 2 Figure 2 S101, obtaining voice feature data of at least two electronic devices in a first space. Figure 2

[0057] S101, obtaining voice feature data of at least two electronic devices in a first space.

[0058] ​​The device management method in the embodiments of the present application can be executed by any one of electronic devices in any one space, or executed by an audio management module on the any one of electronic devices, or executed by a chip or a chip system with the function of the audio management module. There are multiple electronic devices in the any one space, which share the any one space. The electronic device can be a mobile phone, a tablet computer, a notebook computer, a palm computer, a mobile internet device (MID), a wearable device, a virtual reality (VR) device, an augmented reality (AR) device, a wireless terminal in industrial control, a wireless terminal in self driving, a wireless terminal in remote medical surgery, a wireless terminal in smart grid, a wireless terminal in transportation safety, a wireless terminal in smart city, a wireless terminal in smart home, a personal digital assistant (PDA), etc., which are not limited in the embodiments of the present application.

[0059] The first space can be a local space, and multiple electronic devices in the first space share the first space, which means that the multiple electronic devices participate in the same remote call in the first space, and each electronic device can pick up the sound signal in the first space and transmit it to other electronic devices participating in the remote call.

[0060] Optionally, the embodiments of the present application can include two spaces, i.e. the first space and the second space. The second space can be a space other than the first space. For example, the first space is a local space, and the second space is a remote space. The first space and the second space are different spaces, and the first space and the second space are relative, i.e. when the first space is a remote space, the second space is a local space. The first space and the second space are located in different positions, and the electronic devices in the first space and the electronic devices in the second space can transmit the sound signals picked up in their respective spaces through a network. For example, the electronic devices in the first space receive the sound signals transmitted by the electronic devices in the second space through the network, and the sound signals received can be output through each electronic device in the first space. The electronic devices in the second space receive the sound signals transmitted by the electronic devices in the first space through the network, and the sound signals received can be output through each electronic device in the second space, so that the remote call between the electronic devices in the two different spaces can be realized.

[0061] It can be understood that, taking the first space as an example, the number of the second spaces can be multiple. When there are multiple different spaces, any one of the spaces can be selected as the first space, and the remaining spaces can be referred to as the second spaces.

[0062] In the embodiments of the present application, an electronic device is selected from the first space with multiple electronic devices sharing the space, and the device management method is executed by the selected electronic device. For ease of description, the selected electronic device in the first space is referred to as the designated device of the first space. By selecting the designated device from the first space, other electronic devices in the first space can send their respective data to the designated device for analysis and processing, thereby determining the device used by the speaker, and then performing different device management on the device used by the speaker and other devices to solve the echo signal problem in the remote call.

[0063] Alternatively, the designated device of the first space can be selected by one or a combination of the following methods: determining an electronic device in the first space that has been in the remote call for more than a time threshold as the designated device of the first space; or receiving a selection instruction for any electronic device in the first space, and determining the selected electronic device as the designated device of the first space; or obtaining performance level indication information for each electronic device in the first space, and determining an electronic device with a performance level indicated by the performance level indication information greater than a level threshold as the designated device in the first space; or receiving voting information for each electronic device in the first space, and determining an electronic device with a number of votes indicated by the voting information greater than a number threshold as the designated device in the first space. Wherein, the higher the performance level of the electronic device is, the higher the efficiency of the electronic device in processing subsequent data is.

[0064] In the embodiments of the present application, the designated device can be selected from the multiple electronic devices in the first space by one or a combination of the above methods, which can improve the accuracy of the selection of the designated device. Furthermore, each electronic device in the first space can determine to send its own state information such as the state information of the audio input device and the audio output device, voice feature data and other information to the designated device, and the designated device can analyze each electronic device in the first space, thereby performing different types of control on each electronic device, improving the accuracy of the control on each electronic device, and further reducing the echo signal in the remote call.

[0065] The designated device of the first space can obtain the voice feature data of the designated device, and determine the voice feature data of at least two electronic devices in the first space according to the voice feature data of the designated device and the voice feature data received from each electronic device in the first space. The voice feature data of at least two electronic devices in the first space is determined, so as to determine the device used by the speaker in the first space based on the voice feature data of at least two electronic devices in the first space, and then different control is performed on the device used by the speaker and other devices. The device used by the speaker can be the electronic device closest to the speaker, for example, the electronic device held by the speaker, or the device participating in the conference using the account information of the speaker. When the speaker switches, the device used by the speaker also switches. For example, when participant 1 speaks, the electronic device used by the speaker is electronic device 1, and when participant 2 speaks, the electronic device used by the speaker is electronic device 2.

[0066] Optionally, before the designated device of the first space obtains the voice feature data of at least two electronic devices in the first space, the designated device of the first space can establish a communication connection with at least two electronic devices in the first space respectively, and receive the voice feature data transmitted by at least two electronic devices based on the communication connection respectively.

[0067] The communication connection can be established in a manner including but not limited to Bluetooth (BT), wireless local area networks (WLAN) (such as Wi-Fi network) or other local area shared device interconnection, and the application does not limit the communication connection between electronic devices. By establishing a communication connection with at least two electronic devices in the first space, the efficiency of sending voice feature data from each electronic device to the designated device of the first space can be improved, and the network transmission channel in the remote call can not be occupied, and the stability of the remote call can be improved.

[0068] The voice feature data of each electronic device is obtained by extracting the feature of the sound signal picked up in each time window. In the embodiment of the application, the sound signal picked up in each time window is taken as an example to determine the device used by the speaker in each time window, so as to manage each electronic device in the first space in different types according to the device used by the speaker in the time window.

[0069] Optionally, each electronic device can perform feature extraction on the sound signal picked up in the current time window to obtain the voice feature data before sending the voice feature data to the designated device in the first space. After each electronic device performs feature extraction on the sound signal, the voice feature data of at least two electronic devices can be directly analyzed by the designated device in the first space to determine the electronic device used by the speaker in the first space. Instead of transmitting the sound signal itself by each electronic device and then performing feature extraction by the designated device in the first space, the efficiency of voice feature data transmission can be improved.

[0070] The voice feature data may, for example, include but is not limited to one or more of Mel-scale frequency cepstral coefficients (MFCC) features, energy, chroma, phase, etc. of the sound signal. The energy is related to the amplitude of the sound signal, and the greater the amplitude, the higher the energy of the sound signal, which is usually represented as the loudness of the sound. Chroma is used to describe timbre, which is determined by the frequency spectrum structure of the sound. Different people will produce different timbres when speaking. The phase represents the position of the waveform on the time axis, and the change of the phase will affect the waveform of the sound, thereby affecting the characteristics of the sound. Optionally, the voice feature data of the electronic device can be in the form of a feature vector or a feature matrix. The feature vector or feature matrix represented by the voice feature data of each electronic device is different, and the similarity between the voice feature data of two electronic devices can be determined by calculating the similarity distance between the feature vectors or feature matrices of the two electronic devices, thereby determining the similarity between the sound signals picked up by the two electronic devices, and determining whether the sound content of the sound signals picked up by the two electronic devices is the same. The similarity between two voice feature data can be determined by a similarity distance calculation formula, which may, for example, include but is not limited to cosine distance, Euclidean distance, Manhattan distance, Hamming distance, Chebyshev distance, etc.

[0071] Optionally, each electronic device can perform feature extraction on the sound signal picked up in the current time window to obtain audio feature information such as MFCC features in the following manner: sampling the sound signal picked up by each electronic device at a preset sampling period, converting the continuous sound signal into a discrete sound signal, and the sampling period can be determined according to the Nyquist sampling theorem; then filtering the discrete sound signal through a digital filter with a transfer function of H(Z) = 1 - aZ-1, where a is a pre-emphasis coefficient, a > 0.9 and a < 1, to increase the high-frequency resolution of the sound signal; then, the discrete sound signal can be processed by a window function to obtain a plurality of speech frames, that is, N speech frames are obtained, where the window function can be any one of a rectangular window, a Hamming window, or a Hanning window. Further, the energy distribution of each speech frame in the frequency spectrum is obtained by performing fast Fourier transform on each speech frame, that is, the frequency spectrum of each speech frame is obtained by performing fast Fourier transform on each speech frame after windowing, and the energy spectrum of the sound signal is obtained by taking the square of the frequency spectrum of the sound signal. Then, the energy spectrum is filtered through a set of Mel scale triangular filter banks, and finally the log energy output by each triangular filter bank is calculated, and the log energy output by each triangular filter bank is input into a discrete cosine transform (DCT) to obtain the MFCC features of the sound signal.

[0072] Optionally, each electronic device in the first space can also send the sound signal picked up in the current time window to a designated device in the first space, and the selected designated device can perform feature extraction on the sound signal sent by each electronic device to obtain the speech feature data of each electronic device. Since each electronic device does not need to extract the speech feature data of the sound signal picked up in the current time window, the computing power of each electronic device can be reduced.

[0073] Please refer to Figure 3 , Figure 3 An application scenario provided by the embodiments of the present application Figure Three , for example Figure 3As shown, by adding the audio management module on each electronic device in the local space, for example, the audio management module can be added in the audio path on each electronic device, such as the audio processing object (APO), and a communication connection is established between the audio management modules on each electronic device to realize the transmission of voice feature data in the audio management module. By transmitting the voice feature data of each electronic device to the audio management module of the designated device in the first space, the audio management module of the designated device in the first space can analyze and determine the similarity between the sound signals picked up by at least two electronic devices in the first space, and perform time delay synchronization analysis on the sound signals picked up by at least two electronic devices to determine the first electronic device that picks up the sound signal earliest.

[0074] Optionally, if there are multiple electronic devices sharing the space in the remote space, the audio path on each electronic device in the remote space can also add an audio management module, and a communication connection is established between the audio management modules on each electronic device to realize the transmission of voice feature data in the audio management module on each electronic device in the remote space, the similarity analysis between the sound signals, and the determination of the first electronic device that picks up the sound signal earliest in the remote space.

[0075] S102, the audio feature information included in the voice feature data of at least two electronic devices is compared, and the first similarity between the sound signals picked up by at least two electronic devices and the first electronic device that picks up the sound signal earliest among at least two electronic devices are determined.

[0076] Here, by selecting the designated device in the first space from the multiple electronic devices in the first space, the designated device in the first space can analyze based on the voice feature data of at least two electronic devices to determine the electronic device used by the speaker in the first space, i.e. the first electronic device. Since the first electronic device is closest to the speaker, the sound signal picked up by the first electronic device is the clearest, and the speed of picking up the sound signal is the fastest, i.e. the delay in the sound signal is the lowest, so the first electronic device can be controlled to send the picked-up sound signal to other electronic devices, and the second electronic device other than the first electronic device can be controlled not to send the picked-up sound signal to other electronic devices. It can be avoided that the first electronic device and the second electronic device both send sound signals with time difference and different clarity to other electronic devices, resulting in echo signals.

[0077] Optionally, the electronic device in the first space can determine the first electronic device by analyzing the voice feature data of the at least two electronic devices in the first space. For example, the designated device of the first space can perform feature comparison on the voice feature data of the at least two electronic devices, and determine the first similarity between the sound signals picked up by the at least two electronic devices and the first electronic device that picks up the sound signal earliest among the at least two electronic devices.

[0078] The voice feature data can include audio feature information and time indication information. The audio feature information can reflect the sound characteristics in the sound signal, for example, the audio feature information can include one or more of the energy, chroma, loudness, etc. of the sound signal. By comparing the similarity between the audio feature information of two sound signals, the similarity between the sound contents of the two sound signals can be determined, so as to determine whether the contents of the two sound signals are the same, that is, whether the two sound signals are picked up by different electronic devices for the same speaker.

[0079] The time indication information between two sound signals can reflect the time delay between the two sound signals. The lower the time delay between the two sound signals, the closer the electronic device that picks up the two sound signals to the speaker. Since the multiple electronic devices in the first space are different distances from the speaker, the sound picked up by each electronic device has a time delay difference, that is, the sound signal picked up by the electronic device farther away is slower than the sound signal picked up by the electronic device closer to the speaker, so that the sound signals picked up by different electronic devices have a time delay difference. By comparing the time delay difference between any two sound signals, the first electronic device that picks up the sound signal earliest, that is, the first electronic device closest to the speaker, can be determined. For example, the time indication information can be the phase in the sound signal or other information that can be used to indicate the time delay between the sound signals.

[0080] Optionally, the time indication information can also be the time when each electronic device picks up the sound signal of the speaker in the first space. Before obtaining the time indication information in the voice feature data of each electronic device, the device time on each electronic device can be calibrated to make the time on each electronic device consistent, thereby improving the accuracy of subsequent comparison of the time when each electronic device picks up the sound signal of the speaker in the first space.

[0081] In the embodiments of the present application, by comparing the audio feature information in the voice feature data of the at least two electronic devices, it can be determined whether the contents of the sound signals picked up by the at least two electronic devices are the same or similar. If so, it is considered that the sound signals picked up by the at least two electronic devices are the speech sounds picked up by different electronic devices for the same speaker in the current time window. Further, by comparing the time indication information included in the voice feature data of the at least two electronic devices, the time delay difference between the sound signals picked up by the at least two electronic devices can be determined, so as to determine the first electronic device that picks up the earliest sound signal.

[0082] Optionally, the designated device in the first space can use a voice processing model to perform feature comparison on the voice feature data of the at least two electronic devices, to determine the first similarity between the sound signals picked up by the at least two electronic devices and the first electronic device that picks up the earliest sound signal among the at least two electronic devices. The voice processing model may, for example, include but is not limited to a Deep Neural Networks (DNN), a Recurrent Neural Network (RNN), a Residual Neural Network (ResNet), a Convolutional Neural Network (CNN), etc.

[0083] Optionally, before using the voice processing model, a large number of sample sound signals can be obtained to train an initial model to obtain the voice processing model, so that the voice processing model trained can have the ability to perform feature comparison on the voice feature data of the at least two electronic devices, to determine the first similarity between the sound signals picked up by the at least two electronic devices, and to determine the first electronic device that picks up the earliest sound signal among the at least two electronic devices, so as to improve the feature comparison efficiency and accuracy when using the voice processing model for feature comparison subsequently. The process of training the initial model by the designated device in the first space is described below:

[0084] The voice feature data corresponding to the at least two sample sound signals and the initial model are obtained from the training set, the voice feature data corresponding to the at least two sample sound signals is input into the initial model for feature comparison, to determine the sample similarity between the at least two sample sound signals and to determine the sample electronic device that picks up the earliest sample sound signal among the at least two sample electronic devices; the similarity labels between the voice feature data corresponding to the at least two sample sound signals are obtained, and the initial model is trained based on the similarity labels, the sample similarity and the sample electronic device, to obtain the voice processing model.

[0085] The at least two sample sound signals can be obtained by the at least two sample electronic devices picking up sample sound signals of the speaker at different distances from the speaker, and the voice feature data corresponding to the at least two sample sound signals can be obtained by each sample electronic device performing feature extraction on a sample sound signal picked up in a sample time window. The similarity label between the voice feature data corresponding to the at least two sample sound signals can include a sound similarity label and a sample electronic device label. The sound similarity label can be used to reflect whether the at least two sample sound signals are similar, and the sample electronic device label can be used to indicate an electronic device that picks up a sample sound signal earliest among the at least two sample electronic devices, that is, a sample real data. The sample similarity between the at least two sample sound signals and the sample electronic device that picks up the sample sound signal earliest among the at least two sample electronic devices are model prediction data. By comparing whether the sample real data and the model prediction data are consistent, it can be determined whether the initial model converges. When the sample real data and the model prediction data are inconsistent, the model parameters in the initial model can be adjusted so that the sample real data and the model prediction data of the initial model are as consistent as possible. When the sample real data and the model prediction data of the initial model are consistent, it is determined that the initial model converges, and the initial model at this time is saved as a voice processing model.

[0086] Optionally, the process of training the initial model based on the similarity label, the sample similarity, and the sample electronic device to obtain the voice processing model can be as follows: determining a first deviation feature for the initial model based on the sound similarity label and the sample similarity, determining a second deviation feature for the initial model based on the sample electronic device and the sample electronic device label, and training the initial model based on the first deviation feature and the second deviation feature to obtain the voice processing model.

[0087] The process of training the initial model can be a process of adjusting the model parameters in the initial model. The first deviation feature and the second deviation feature can be a loss function of the initial model. The greater the value of the loss function of the initial model, the lower the prediction accuracy of the initial model. Therefore, the model parameters in the initial model can be adjusted to reduce the value of the loss function in the initial model. When the value of the loss function in the initial model is less than a loss threshold, for example, the first deviation feature is less than a first deviation threshold and the second deviation feature is less than a second deviation threshold, or when the number of iterations of the initial model is greater than or equal to a number threshold, or when the training duration of the initial model satisfies a preset duration, it is considered that the initial model at this time converges, and the initial model at this time is saved as a voice processing model. Since the initial model is trained by selecting a large amount of voice feature data corresponding to sample sound signals from the training set, the accuracy and stability of the voice processing model can be improved.

[0088] That is, in this scenario, the voice processing model can directly compare the voice feature data of the at least two electronic devices, output two results, one result is used to indicate whether the at least two sample sound signals are similar, that is, the sample similarity result, and the other result is used to indicate the sample electronic device that picks up the sample sound signal earliest among the at least two sample electronic devices. By combining the two results output by the voice processing model, the first electronic device can be determined, and subsequent different controls can be performed on the first electronic device and other electronic devices, thereby improving the accuracy of the control of the electronic device.

[0089] Optionally, in the case of comparing the voice feature data of the at least two electronic devices in the first space and determining that the voice feature data of the at least two electronic devices in the first space are similar, if the voice feature data of the at least two electronic devices in the first space both satisfy the target transmission condition, any one of the electronic devices can be selected as the first electronic device. The voice feature data of the electronic device satisfying the target transmission condition can mean that the clarity of the sound signal of the electronic device is greater than the clarity threshold, and the sound signal delay is less than the delay threshold, that is, when the sound signal of the electronic device is transmitted to the electronic device in the second space through the network, the sound signal output by the electronic device in the second space is also clear. For example, if the voice feature data of the at least two electronic devices in the first space both satisfy the target transmission condition, any one of the electronic devices with a network speed greater than a network speed threshold, or a network remaining traffic greater than a traffic threshold, or a device performance greater than a performance threshold can be selected as the first electronic device.

[0090] In the embodiments of the present application, for example, the area of the first space is small, and the sound of the same speaker can be picked up by each electronic device in the first space, and the picked-up sound signal is clear and the sound signal delay is low, then the optimal first electronic device can be selected by further combining the network speed, remaining traffic, device performance and the like, so that the sound in the first space can be picked up by the first electronic device subsequently, and the sound signal transmission efficiency in the first space is improved.

[0091] S103, control the first electronic device to send the picked-up sound signal to the other electronic devices, and control the second electronic device not to send the picked-up sound signal to the other electronic devices.

[0092] The first electronic device refers to an electronic device used by a speaker in the first space within a current time window, the second electronic device includes electronic devices other than the first electronic device in the first space, and the other electronic device includes electronic devices in the first space and electronic devices in the second space. That is, controlling the first electronic device to send the picked-up sound signal to the other electronic device means controlling the first electronic device to send the picked-up sound signal to the electronic devices other than the first electronic device in the first space and the electronic devices in the second space. Controlling the second electronic device not to send the picked-up sound signal to the other electronic device means controlling the second electronic device not to send the picked-up sound signal to the electronic devices other than the first electronic device in the first space and the electronic devices in the second space.

[0093] It can be understood that the first electronic device may, for example, refer to a main electronic device in the first space, and the second electronic device may, for example, refer to a secondary electronic device in the first space. When there is only one speaker in the first space, the number of main electronic devices in the first space is one, and the number of secondary electronic devices in the first space can be the total number of devices other than the main electronic device. When there are multiple speakers in the first space, the number of main electronic devices in the first space is multiple, and the number of secondary electronic devices in the first space can be the total number of devices other than the multiple main electronic devices. In the case where multiple speakers in the first space speak in the current time window, the number of first electronic devices can be multiple, and the processing of the specified device of the first space for multiple first electronic devices can be the same.

[0094] Optionally, the electronic device in the first space can determine the first electronic device by: performing feature comparison on audio feature information included in voice feature data of at least two electronic devices to determine a first similarity between sound signals picked up by the at least two electronic devices; and in a case where the first similarity is greater than or equal to a first similarity threshold, comparing time indication information included in the voice feature data of the at least two electronic devices to determine a first electronic device that picks up the earliest sound signal among the at least two electronic devices.

[0095] The feature comparison on the audio feature information included in the voice feature data of the at least two electronic devices to determine the first similarity between the sound signals picked up by the at least two electronic devices can refer to calculating a similarity distance between the audio feature information included in the voice feature data of the at least two electronic devices. The similarity distance calculation method can refer to the description in the foregoing, which will not be repeated here.

[0096] It can be seen that, by performing feature comparison on the audio feature information included in the speech feature data of the at least two electronic devices, it can be determined whether the sound signals picked up by the at least two electronic devices are the same or similar. If the sound signals picked up by the at least two electronic devices are not the same or similar, it can be indicated that the sound signals picked up by the two electronic devices are not the sound signals of the same speaker, for example, there are multiple speakers in the first space speaking at the same time, and the electronic devices used by each speaker pick up the speech of the respective speaker. If the sound signals picked up by the at least two electronic devices are the same or similar, it can be indicated that the sound signals picked up by the at least two electronic devices are the sound of the same speaker, for example, the sound signals picked up by the two electronic devices are the same, which can indicate that the distances between the two electronic devices and the speaker are the same, and the sound signals picked up by the two electronic devices are similar, which can indicate that the distances between the two electronic devices and the speaker are different, but both pick up the speech of the same speaker. Then, by further determining the time indication information included in the speech feature data of the at least two electronic devices, the first electronic device that picks up the earliest sound signal can be determined.

[0097] That is, in the case where the first similarity is less than the first similarity threshold, there is no need to compare the time indication information included in the speech feature data of the at least two electronic devices, which can improve the determination efficiency of the first electronic device. In the case where it is determined that the first similarity is greater than or equal to the first similarity threshold, the time indication information included in the speech feature data of the at least two electronic devices is compared to determine the first electronic device that picks up the earliest sound signal among the at least two electronic devices, which can improve the accuracy of the determination of the electronic device.

[0098] Optionally, the process of performing feature comparison on the audio feature information included in the speech feature data of the at least two electronic devices to determine the first similarity between the sound signals picked up by the at least two electronic devices can be implemented by a speech processing model. Specifically, the specified device in the first space can input the audio feature information included in the speech feature data of the at least two electronic devices into the speech processing model for feature comparison to determine the first similarity between the sound signals picked up by the at least two electronic devices.

[0099] The voice processing model is obtained by training an initial model based on a similarity label between the at least two sample sound signals and a sample similarity between the at least two sample sound signals, and the sample similarity is obtained by comparing sample audio feature information of the at least two sample sound signals based on the initial model. The first similarity between the sound signals picked up by the at least two electronic devices refers to a model output result, and the similarity label between the at least two sample sound signals is a sample real result. By comparing the difference between the model output result and the sample real result, the initial model can be adjusted based on the difference. When the difference between the model output result and the sample real result is less than a difference threshold, or the number of iterations of the initial model is greater than or equal to a number threshold, or the training time length of the initial model satisfies a preset time length, it can be determined that the initial model converges, and the initial model at this time is saved as the voice processing model.

[0100] That is, when training the initial model, the initial model can only be trained for the similarity calculation capability between the sound signals picked up by the at least two electronic devices, and when performing feature comparison using the voice processing model, the first similarity between the sound signals picked up by the at least two electronic devices can be output. For the capability of determining the first electronic device that picks up the sound signal earliest among the at least two electronic devices, the initial model can be used or not used. When the initial model is used to determine the capability of the first electronic device that picks up the sound signal earliest among the at least two electronic devices, the training process of the initial model can refer to the foregoing description, which will not be described here.

[0101] In the embodiment of the application, after the specified device in the first space determines the first electronic device, the first electronic device sends the picked-up sound signal to the second electronic device and the electronic devices in the second space. After the electronic devices in the second space receive the sound signal picked up by the first electronic device, the electronic devices in the second space can output the sound signal picked up by the first electronic device in the second space, thereby realizing transmission of the sound signal in the first space to the ears of the participants in the second space.

[0102] Please refer to Figure 4 , Figure 4 An application scenario diagram for processing echo in a remote space is provided for the embodiment of the application, as shown in Figure 4As shown, there are two electronic devices in the local space, which are the electronic device 41 of the participant 1 and the electronic device 42 of the participant 2. The current state is that the participant 2 is speaking, and the electronic device 41 of the participant 1 does not actively mute the audio input device such as a microphone. The voice signal of the participant 2 is picked up by the electronic device 42, and the voice signal energy is strong and the voice is clear. The electronic device 41 of the participant 1 also picks up the voice signal of the participant 2, but the voice signal energy is weak and there is a delay. The voice signals picked up by the electronic device 41 and the electronic device 42 can be transmitted to each electronic device in the local space and the electronic device in the remote space through the cloud. When the voice signals picked up by the electronic device 41 and the electronic device 42 are transmitted to the electronic device in the remote space through the cloud, the problem of echo occurs when the participant 2 speaks in the remote space.

[0103] Therefore, by adding an audio management module in the electronic device 41 and the electronic device 42 in the local space, the audio management module of each electronic device extracts the voice feature data of the voice signal picked up by itself, and transmits the voice feature data of each electronic device to the audio management module of the designated device (such as the electronic device 42) in the first space based on the communication connection. The audio management module of the designated device in the first space compares the voice feature data of at least two electronic devices to determine the first electronic device that picks up the earliest voice signal among the at least two electronic devices, i.e., the electronic device closest to the speaker. Moreover, the audio management module of the designated device in the first space controls the first electronic device to send the picked-up voice signal to other electronic devices, and controls the second electronic device not to send the picked-up voice signal to other electronic devices. Then, only the first electronic device (for example, the first electronic device is the designated device in the first space, i.e., the electronic device 42) in the first space sends the picked-up voice signal to the cloud. As shown in ② in FIG. 1B, the second electronic device (such as the electronic device 41) does not send the voice signal with a delay to the cloud, i.e., does not send the picked-up voice signal to the cloud. The remote space receives the picked-up voice signal of the first electronic device and outputs it, so there is no echo signal in the remote space. Figure 4

[0104] Optionally, after the second electronic device in the first space receives the picked-up voice signal of the first electronic device, the second electronic device can determine whether to output the picked-up voice signal of the first electronic device in the first space according to its own situation.

[0105] ​In one case, for example, the area of the first space is greater than the area threshold, then for the second electronic device which is greater than the distance threshold from the speaker, the sound signal picked up by the first electronic device can be controlled to be output; for the second electronic device which is less than or equal to the distance threshold from the speaker, the sound signal picked up by the first electronic device can be controlled not to be output. Since when the area of the first space is large, the user of the electronic device which is greater than the distance threshold from the speaker can hardly hear the speech content of the speaker clearly, the sound signal picked up by the first electronic device is output through the second electronic device, which can ensure the quality and accuracy of the remote call. The user of the electronic device which is less than or equal to the distance threshold from the speaker can hear the speech content of the speaker clearly, so there is no need to output the sound signal, which can reduce the power consumption of the second electronic device while ensuring the quality of the remote call.

[0106] In another case, for example, when the area of the first space is greater than the area threshold, the audio output device of the first electronic device is controlled to output the picked-up sound signal, and the second electronic device is controlled not to output the sound signal picked up by the first electronic device. Wherein, when the audio output device of the first electronic device outputs the picked-up sound signal, the sound signal can be amplified, so that when the audio output device of the first electronic device outputs the picked-up sound signal, each participant in the first space can clearly hear the sound signal, ensuring the accuracy of the remote call.

[0107] It can be understood that if multiple electronic devices share the second space in the second space, any electronic device in the second space can be selected as the designated device of the second space by referring to the way of selecting the designated device in the first space, so that the designated device of the second space can receive the voice feature data sent by each electronic device in the second space to determine the first electronic device in the second space, i.e. the device used by the speaker in the second space, so as to realize different control operations for different electronic devices in the second space. The specific control method can refer to the control operation of the electronic device in the first space, which will not be described in detail here.

[0108] In the embodiments of the present application, since the first electronic device is determined to be the device used by the speaker in the current time window, i.e., the device closest to the speaker, the sound signal picked up by the first electronic device in the current time window is clearer and has the lowest delay compared with the second electronic device in the first space. By controlling the first electronic device to send the picked-up sound signal to other electronic devices and controlling the second electronic device not to send the picked-up sound signal to other electronic devices, the sound signal received by the electronic device receiving the sound signal is the clearest and has the lowest delay, i.e., the sound effect is the best, thereby solving the echo problem caused by receiving multiple sound signals with time difference by the electronic device receiving the sound signal and improving the call quality in remote calls.

[0109] Optionally, in the case where the specified device in the first space is the first electronic device, the control of the first electronic device to send the picked-up sound signal to other electronic devices means that the first electronic device sends the picked-up sound signal to other electronic devices. In the case where the specified device in the first space is not the first electronic device, the control of the first electronic device to send the picked-up sound signal to other electronic devices means that the specified device in the first space controls the first electronic device to send the picked-up sound signal to other electronic devices.

[0110] Optionally, the manner in which the selected electronic device in the first space controls the first electronic device to send the picked-up sound signal to other electronic devices can be as follows: sending a first control instruction to the first electronic device; the first control instruction is used to control the first electronic device to send the picked-up sound signal to other electronic devices.

[0111] By sending the first control instruction to the first electronic device, the first electronic device can control the first electronic device to send the picked-up sound signal to other electronic devices based on the first control instruction.

[0112] Optionally, the manner in which the selected electronic device in the first space controls the second electronic device not to send the picked-up sound signal to other electronic devices can be as follows: sending a second control instruction to the second electronic device; the second control instruction is used to control the second electronic device not to send the picked-up sound signal to other electronic devices.

[0113] By sending the second control instruction to the second electronic device, the second electronic device can control the second electronic device not to send the picked-up sound signal to other electronic devices based on the second control instruction. By sending different control instructions to the first electronic device and the second electronic device respectively, the first electronic device and the second electronic device can know whether to send the picked-up sound signal to other electronic devices, thereby avoiding the first electronic device and the second electronic device both sending the picked-up sound signal to other electronic devices to cause echo.

[0114] Optionally, when the first control instruction is sent to the first electronic device, the initial working state of the first electronic device in the current time window can be determined first, so that the control instruction sent to the first electronic device is determined in combination with the initial working state of the first electronic device in the current time window. For example, the selected electronic device in the first space can obtain the initial working state of the first electronic device in the current time window; the initial working state is used to indicate whether to send the picked-up sound signal to other electronic devices; in the case where the initial working state of the first electronic device indicates not to send the picked-up sound signal to other electronic devices, the first control instruction is sent to the first electronic device.

[0115] The initial working state of each electronic device can include the initial working state of the audio input device such as the microphone and the initial working state of the audio output device such as the loudspeaker of each electronic device. The initial working state of each electronic device can be obtained by each electronic device reading the flag bit controlled by the audio input device such as the microphone and the flag bit controlled by the audio output device such as the loudspeaker of itself. By reading the flag bit controlled by the microphone and the flag bit controlled by the loudspeaker of itself, the initial working state of the microphone and the loudspeaker of itself can be determined, so that the initial working state of itself can be transmitted to the designated device of the first space through the communication connection.

[0116] By determining the control instruction for the first electronic device in combination with the initial working state of the first electronic device in the current time window, the control accuracy for the first electronic device can be improved.

[0117] Optionally, when the second control instruction is sent to the second electronic device, the initial working state of the second electronic device in the current time window can be determined first, so that the control instruction sent to the second electronic device is determined in combination with the initial working state of the second electronic device in the current time window. For example, the selected electronic device in the first space can obtain the initial working state of the second electronic device in the current time window; the initial working state is used to indicate whether to send the picked-up sound signal to other electronic devices; in the case where the initial working state of the second electronic device indicates to send the picked-up sound signal to other electronic devices, the second control instruction is sent to the second electronic device.

[0118] By determining the control instruction for the second electronic device in combination with the initial working state of the second electronic device in the current time window, the control accuracy for the second electronic device can be improved.

[0119] In the embodiment of the present application, by determining the first electronic device in the first space, the first electronic device can be controlled to send the picked-up sound signal to other electronic devices, so that the sound signal output by the electronic device in the second space is clearer. By controlling the second electronic device not to send the picked-up sound signal to other electronic devices, the problem that the second space, such as a remote space, has an echo signal can be avoided when multiple sound signals with time difference are sent to the electronic device in the second space.

[0120] In the following, the embodiment of the present application is combined Figure 5 The specific process of the device management method provided by the embodiment of the present application is described in detail, please refer to Figure 5 , Figure 5 The flowchart of the device management method provided by the embodiment of the present application Figure Two , Figure 5 The main problem to be solved in the present application is the echo signal problem in the local space in remote communication, as shown in Figure 5 The method can include but is not limited to the following steps:

[0121] S201, obtaining voice feature data of at least two electronic devices in a first space.

[0122] S202, determining a first electronic device based on the voice feature data of the at least two electronic devices.

[0123] The specific implementation of steps S201-S202 can refer to the implementation of steps S101-S102 described above, which will not be repeated here.

[0124] Optionally, after determining the first electronic device in the first space, the designated device of the first space can control the audio output device of the second electronic device not to output the sound signal picked up in the current time window in the first space.

[0125] In the embodiment of the present application, since the designated device of the first space determines the first electronic device used by the speaker in the first space, and the first electronic device sends the picked-up sound signal to other electronic devices, when the second electronic device in the first space receives the picked-up sound signal sent by the first electronic device, the second electronic device can output the picked-up sound signal sent by the first electronic device in the first space. However, since the speaker speaks in the first space, all users of electronic devices in the first space, i.e. all participants in the first space, can hear the speaker's voice, so there will be air propagation of the speaker's voice in the first space, and there will also be the sound signal picked up by the first electronic device received by the second electronic device through the network. Since the speed of sound signal propagation through air is greater than the speed of network transmission, an echo signal with time difference will be generated in the first space, all participants in the first space will hear the echo signal with time difference, which affects the quality of remote communication.

[0126] Therefore, the designated device of the first space controls the audio output device of the second electronic device not to output the sound signal of the first space picked up in the current time window, the second electronic device in the first space does not output the sound signal of the first space picked up in the current time window, and only the user (i.e. the speaker) of the first electronic device in the first space outputs the sound signal through air propagation, so that the echo signal in the first space is not generated, and the clarity of the remote call in the first space is improved.

[0127] Optionally, the audio output device of the second electronic device can be muted, so that the audio output device of the second electronic device does not output the sound signal picked up in the current time window.

[0128] Optionally, the designated device of the first space can control the audio output device of the second electronic device to output the sound signal of the second space picked up by the electronic device in the second space in the current time window.

[0129] That is, in the case of the current time window, if there are two speakers in the first space and the second space speaking at the same time, each electronic device in the first space can output the sound signal of the speaker picked up in the second space, and each electronic device in the second space can output the sound signal of the speaker picked up in the first space, so as to avoid missing of the sound signal in the first space and the second space and improve the accuracy of the remote call.

[0130] Optionally, when controlling the audio output device of the second electronic device not to output the sound signal of the first space picked up in the current time window, the audio output device of the second electronic device can be directly controlled to be muted, so that the audio output device of the second electronic device does not output all the sound signals picked up in the current time window. However, since there may be a speaker in the second space speaking in the current time window, it is further determined whether only the speaker in the local space is speaking in the current time window, so as to avoid missing the sound signal of the remote space when there is a speaker in the remote space speaking.

[0131] S203, acquiring audio feature information of the sound signal received by the second electronic device.

[0132] S204, performing feature comparison on the audio feature information of the sound signal received by the second electronic device and the audio feature information included in the voice feature data of the first electronic device, to determine a second similarity between the sound signal received by the second electronic device and the sound signal picked up by the first electronic device.

[0133] S205, in a case where the second similarity is greater than or equal to the second similarity threshold, determining that the sound signal received by the second electronic device is the sound signal of the first space picked up in the current time window, and controlling the audio output apparatus of the second electronic device not to output the received sound signal.

[0134] The audio feature information of the sound signal received by the second electronic device is obtained by feature extraction on the sound signal received by the second electronic device. The sound signal received by the second electronic device can include the sound signal of the first space picked up in the current time window by the first electronic device and the sound signal of the second space picked up in the current time window by the electronic device in the second space and transmitted through the network. The audio feature information is obtained by feature extraction on the two different types of sound signals respectively. By comparing the audio feature information of the sound signal received by the second electronic device with the audio feature information included in the voice feature data of the first electronic device, it can be determined whether the sound content of the sound signal received by the second electronic device is the same as the sound content of the sound signal picked up by the first electronic device in the current time window, i.e., whether it is the speech content of the same speaker. In a case where the second similarity between the sound signal received by the second electronic device and the sound signal picked up by the first electronic device is greater than or equal to the second similarity threshold, it indicates that the sound signal received by the second electronic device is the sound signal picked up by the first electronic device and transmitted through the network, and there is no speaker in the second space in the current time window. Therefore, the audio output apparatus of the second electronic device can be controlled not to output the received sound signal, for example, directly mute the audio output apparatus of the second electronic device.

[0135] Optionally, the audio feature information of the sound signal received by the second electronic device and the audio feature information included in the voice feature data of the second electronic device can also be compared for feature comparison to determine the second similarity between the sound signal received by the second electronic device and the sound signal picked up by the second electronic device.

[0136] Since it is determined through the above steps that the sound signals of the first electronic device and the second electronic device are similar, i.e., the sound signals of the same speaker picked up by the two electronic devices, the results of the feature comparison are the same whether the audio feature information included in the voice feature data of the first electronic device or the voice feature data of the second electronic device is used for feature comparison with the audio feature information of the sound signal received by the second electronic device.

[0137] Optionally, since the first electronic device and the second electronic device are located in the first space, when the second electronic device receives the sound signal, the first electronic device will also receive the same sound signal. If it is determined that the sound signal received by the second electronic device is the sound signal picked up in the first space within the current time window, the audio output device of the first electronic device can also be controlled not to output the received sound signal, thereby avoiding the generation of an echo signal in the first space.

[0138] In other words, when each electronic device in the first space receives a sound signal sent by another electronic device in the first space, the audio output device of each electronic device in the first space does not output the sound signal. Thus, only the speaker's voice signal is transmitted through the air in the first space, so as to avoid generating an echo signal in the local space.

[0139] S206, if the second similarity is less than the second similarity threshold, determine that the sound signal received by the second electronic device is the second spatial sound signal picked up within the current time window, and control the audio output device of the second electronic device to output the received sound signal.

[0140] Since the second similarity between the sound signal received by the second electronic device and the sound signal picked up by the first electronic device is less than the second similarity threshold, it means that the sound signal received by the second electronic device is not sent through the network by the sound signal picked up by the first electronic device, but rather the sound signal of the speaker in the second space picked up by the electronic device in the second space within the current time window and sent through the network. Therefore, the audio output device of the second electronic device is controlled to output the received sound signal, thereby avoiding the omission of the call content in the second space and improving the accuracy of the remote call.

[0141] Please see Figure 6 , Figure 6 This application provides an illustration of an application scenario for processing echoes in the local space, as shown in the following diagram. Figure 6 As shown, there are two electronic devices in this space. Currently, participant 2 is speaking, and participant 1's electronic device 1 has not actively muted its audio output device (such as a speaker). Participant 2's voice signal is picked up by electronic device 62 and transmitted through the cloud. The audio output device (such as a speaker) of participant 1's electronic device 61 plays the signal, which participant 2 hears. Simultaneously, participant 1 can also hear participant 2's voice signal transmitted through the air. Due to the delay between the voice signal transmitted through the cloud and the voice signal transmitted through the air, all participants in this space, such as participant 1 and participant 2, will clearly hear an echo signal.

[0142] Therefore, by adding the audio management module in each electronic device in the local space, such as the electronic device 61 and the electronic device 62, the audio management module of each electronic device 61 can obtain the sound signal transmitted through the network, and perform feature extraction on the sound signal to obtain the speech feature data, transmit the respective speech feature data to the audio management module of the designated device 62 (such as the electronic device 62) in the first space based on the communication connection, and the audio management module of the designated device in the first space performs feature comparison on the speech feature data of at least two electronic devices, determines that the sound signal received by the second electronic device is the sound signal of the first space picked up in the current time window, that is, all the sound signals received by the second electronic device are the sound signals transmitted by the first electronic device through the network, and then mutes the audio output device of the second electronic device, such as controlling the audio output device of the second electronic device (such as the electronic device 61) not to output the received sound signal. Then, as shown in ③ in the figure, only the sound signal transmitted through the air exists in the local space, so that there is no echo signal in the local space. Figure 6

[0143] Alternatively, if it is detected that any electronic device in the first space picks up a sound signal in any time window, and it is determined that the electronic device in the second space does not pick up a sound signal in the any time window, it means that the sound signal received by each electronic device in the first space is picked up by the any electronic device in the first space, and the audio output device of all the electronic devices in the first space is controlled not to output the received sound signal, thereby reducing the echo signal in the local space. That is, when a participant in the first space speaks, if it is determined that no participant in the second space speaks at this time, all the loudspeakers in the first space can be muted to reduce the echo in the first space.

[0144] Alternatively, since the first electronic device and the second electronic device are located in the first space, when the second electronic device receives a sound signal, the first electronic device also receives the same sound signal. In the case where it is determined that the sound signal received by the second electronic device is the sound signal of the second space picked up in the current time window, the audio output device of the first electronic device can also be controlled to output the received sound signal to avoid the problem of missing the sound of the second space in the remote call.

[0145] That is, when each electronic device in the first space receives the sound signal sent by the electronic device in the second space, the audio output device of each electronic device in the first space outputs the sound signal to avoid missing the sound signal of the remote space.

[0146] ​Optionally, if a sound signal is received in any time window, a processing manner for the sound signal can be determined in combination with an electronic device to which the received sound signal belongs. The any time window can include a current time window or other time windows. For example, the specified device of the first space receives a sound signal picked up by the third electronic device sent by the third electronic device; in the case that the third electronic device is an electronic device in the second space, the audio output device of the second electronic device is controlled to output the sound signal picked up by the third electronic device.

[0147] For example, it can be determined that the third electronic device belongs to the first space or the second space by obtaining an identifier of the third electronic device. The identifier of the third electronic device may, for example, include but is not limited to participant identity information of the third electronic device, an Internet Protocol Address (IP address) of the third electronic device, an organizational location to which the third electronic device belongs, a factory number of the third electronic device, and the like. Since the sound signal picked up by the third electronic device in the remote space is received, the audio output devices of all the second electronic devices in the local space can be unmuted, so that the audio output devices of all the second electronic devices in the local space output the sound signal picked up by the third electronic device, reduce the omission of the sound signal of the remote space, and ensure the accuracy of the remote call.

[0148] Optionally, in the case that the third electronic device is an electronic device in the first space, the audio output device of the second electronic device is controlled not to output the sound signal picked up by the third electronic device. Since the received sound signal is a sound signal picked up by the local space, in order to reduce the echo signal generated by the local space, the audio output device of the second electronic device is controlled not to output the sound signal picked up by the third electronic device.

[0149] Optionally, in the case that it is determined that the second similarity between the sound signal picked up by the third electronic device sent by the third electronic device and the sound signal picked up by the first electronic device is small, it can be further determined whether the third electronic device is an electronic device in the second space to determine the processing manner for the audio output device of the third electronic device, so as to determine whether to output the sound signal sent by the third electronic device in the first space.

[0150] Specifically, the audio feature information of the sound signal sent by the third electronic device is compared with the audio feature information included in the voice feature data of the first electronic device to determine the second similarity between the sound signal sent by the third electronic device and the sound signal picked up by the first electronic device; in the case that the second similarity is less than a second similarity threshold and the third electronic device is an electronic device in the second space, the audio output device of the second electronic device is controlled to output the sound signal sent by the third electronic device.

[0151] The sound signal sent by the third electronic device refers to the sound signal picked up by the third electronic device and sent by the third electronic device. When receiving the sound signal sent by any electronic device, by comparing the sound signal sent by the electronic device with the sound signal picked up by the first electronic device in the local space, and further combining the space to which the electronic device belongs to determine whether to output the sound signal sent by the electronic device, the accuracy of the sound signal output can be improved. For example, when receiving the sound signal sent by the third electronic device, by further combining the space to which the third electronic device belongs to determine whether the sound signal sent by the third electronic device is the sound signal sent by the remote space or the sound signal sent by the local space, the accuracy of the sound signal judgment can be improved, so as to avoid the case that when there are multiple sound signals in the local space, the sound signal picked up by each electronic device is not accurate, which leads to inaccurate voice similarity comparison, thereby causing the remote signal to be missed in the remote call, and the accuracy of the remote call can be improved.

[0152] Optionally, in a case where the second similarity is less than the second similarity threshold and the second electronic device is an electronic device in the first space, the audio output apparatus of the second electronic device is controlled to not output the sound signal picked up by the second electronic device.

[0153] In the embodiment of the present application, in a case where the second similarity is less than the second similarity threshold and the second electronic device is an electronic device in the first space, it is indicated that the sound signal received by the second electronic device is the sound signal sent by the local space, and then the audio output apparatus of the second electronic device is controlled to not output the sound signal picked up by the second electronic device, so as to avoid the echo signal in the local space.

[0154] It can be seen that, by controlling the audio output apparatus of the second electronic device to not output the sound signal of the first space picked up in the current time window, the echo signal in the first space will not be generated, thereby solving the echo problem in the first space in the remote call process and improving the clarity of the remote call. Moreover, by further analyzing the second similarity between the sound signal received by the second electronic device and the sound signal picked up by the first electronic device, the accuracy of determining that the sound signal received by the second electronic device is the sound signal of the first space picked up in the current time window or the sound signal received by the second electronic device is the sound signal of the second space picked up in the current time window can be improved, thereby improving the accuracy of controlling the audio output apparatus of the second electronic device.

[0155] It can be understood that the solution to the echo signal in the first space in the embodiment of the present application and the solution to the echo signal in the first space in the embodiment of the present application Figure 2 are the same. Figure 5In the embodiments, the solutions for the echo signals in the second space can be executed synchronously, and the execution processes of the two solutions do not affect each other. Alternatively, the solutions for the echo signals in the first space can be executed first, followed by the solutions for the echo signals in the second space, or the solutions for the echo signals in the second space can be executed first, followed by the solutions for the echo signals in the first space. This application does not limit this.

[0156] The above Figure 2 , Figure 5 The embodiments illustrate the process of managing designated devices in the first space. The following describes the process in conjunction with... Figure 7 Regarding the above Figure 2 , Figure 5 The following embodiment illustrates the interaction flowchart of each module during device management by a designated device in the first space. For an example, please refer to... Figure 7 , Figure 7 The present application provides an embodiment of an electronic device with interactive flowcharts for each module. During device management, the interactive flow of each module within the electronic device is as follows:

[0157] S301, Identify the designated device from the first space.

[0158] Here, "designated device" refers to a designated device within the first space. For example, the designated device in the first space can be determined from within the first space, or it can be determined from any electronic device within the first space. After the designated device is determined, a communication connection is established between the audio management module of the designated device and the audio management modules of multiple electronic devices within the first space. The audio management module of the designated device can then receive and analyze the voice feature data sent by the audio management modules of other electronic devices to determine the first electronic device used by the speaker.

[0159] Optionally, Figure 7 The multiple electronic devices in the first space may include, for example, a designated device in the first space, a first electronic device, and a second electronic device. After the designated device is determined, the audio management module of the designated device can establish a connection with the audio management modules of each electronic device in the first space.

[0160] S302, a communication connection is established between the audio management module of the designated electronic device and the audio management module of the first electronic device.

[0161] S303, establish a communication connection between the audio management module of the designated electronic device and the audio management module of the second electronic device.

[0162] S304, the audio input device of the specified device picks up the sound signal in the first space and transmits to the audio management module of the specified device.

[0163] S305, the audio management module of the specified device extracts features from the sound signal to obtain speech feature data.

[0164] Among them, the audio input device of each electronic device picks up the sound signal in the first space, which can be transmitted to the audio management module of itself, and the audio management module of itself extracts features from the sound signal picked up by itself to obtain speech feature data.

[0165] S306, the audio input device of the first electronic device picks up the sound signal in the first space and transmits to the audio management module of the first electronic device.

[0166] S307, the audio management module of the first electronic device extracts features from the sound signal to obtain speech feature data, and transmits to the audio management module of the specified device.

[0167] S308, the audio input device of the second electronic device picks up the sound signal in the first space and transmits to the audio management module of the second electronic device.

[0168] S309, the audio management module of the second electronic device extracts features from the sound signal to obtain speech feature data, and transmits to the audio management module of the specified device.

[0169] Among them, each electronic device can transmit speech feature data to the audio management module of the specified device based on communication connection.

[0170] S310, the audio management module of the specified device determines the first electronic device based on the speech feature data of multiple electronic devices.

[0171] S311, the audio management module of the specified device sends the first control instruction to the audio management module of the first electronic device.

[0172] Among them, the audio management module of the specified device can send the first control instruction to the audio management module of the first electronic device, and the first control instruction is used to control the first electronic device to send the picked-up sound signal to other electronic devices, such as sending the picked-up sound signal to the second electronic device and the electronic device in the second space.

[0173] S312, the audio management module of the first electronic device sends the picked-up sound signal to the audio management module of the second electronic device.

[0174] S313, the audio output device of the second electronic device outputs the received sound signal.

[0175] The second electronic device can output the received sound signal through an audio output device such as a loudspeaker.

[0176] S314, the audio management module of the first electronic device sends the picked-up sound signal to the audio output device of the electronic device in the second space.

[0177] The audio management module of the first electronic device can send the picked-up sound signal to the electronic device in the second space. If the number of electronic devices in the second space is one, no audio management module can be added to the electronic device in the second space, and the audio management module of the first electronic device can send the picked-up sound signal to the electronic device in the second space or the audio output device of the electronic device in the second space. If the number of electronic devices in the second space is more than one, i.e., there are multiple electronic devices sharing the space in the second space, an audio management module can be added to each electronic device in the second space, and the audio management module of the first electronic device can send the picked-up sound signal to the audio management module of each electronic device in the second space, and the audio management module of each electronic device in the second space can output the received sound signal through an audio output device.

[0178] S315, the audio output device of the electronic device in the second space outputs the received sound signal.

[0179] For example, the electronic device in the second space can output the received sound signal through an audio output device such as a loudspeaker.

[0180] It can be understood that if the specified device in the first space is not the first electronic device, when the audio management module of the first electronic device sends the picked-up sound signal to the second electronic device, it also includes sending the picked-up sound signal to the audio output device of the specified device, and the audio output device of the specified device outputs the received sound signal.

[0181] It can be understood that if the specified device in the first space is the first electronic device, the audio output device of the specified device can output or not output the picked-up sound signal.

[0182] S316, the audio management module of the specified device sends a second control instruction to the audio management module of the second electronic device.

[0183] S317, the audio management module of the second electronic device does not send the picked-up sound signal to other electronic devices.

[0184] The audio management module of the specified device can send a second control instruction to the audio management module of the second electronic device, and the second control instruction is used to control the audio management module of the second electronic device to not send the picked-up sound signal to other electronic devices, that is, the second electronic device does not send the picked-up sound signal to the first electronic device, and the second electronic device does not send the picked-up sound signal to the electronic device in the second space.

[0185] Optionally, the manner in which the second electronic device is controlled to not send the picked-up sound signal to other electronic devices can include the following two manners:

[0186] The first manner: the OS system (Operating System) of the second electronic device receives the second control instruction sent by the specified device based on the communication connection, and the OS system in which the second electronic device is located closes the audio input device such as a microphone.

[0187] The second manner: the audio management module of the second electronic device intercepts the path of the second electronic device transmitting the sound signal to the network, so that the second electronic device does not transmit the sound signal to the network. It can be understood that the path of the second electronic device transmitting the sound signal to the network is always open during the entire remote call process, and when the second electronic device is controlled to not transmit the sound signal to the network, the path does not need to be closed, and only the sound signal needs to be intercepted.

[0188] It can be understood that when the second electronic device is controlled to not send the picked-up sound signal to other electronic devices in the embodiment of the application, the second electronic device can still pick up the sound signal in the first space, and the voice feature data can be obtained by performing feature extraction on the sound signal in each time window and transmitted to the specified device in the first space based on the communication connection, so as to determine whether the speaker changes, that is, whether the first electronic device changes, so as to control the changed speaker to use the first electronic device to send the picked-up sound signal to other electronic devices, and control the second electronic device to not send the picked-up sound signal to other electronic devices, so as to improve the clarity of the remote call.

[0189] S318, the audio management module of the specified device sends a third control instruction to the audio management module of the second electronic device.

[0190] S319, the audio output device of the second electronic device does not output the sound signal picked up in the first space in the current time window.

[0191] The third control instruction is used to control the audio output device of the second electronic device to not output the sound signal picked up in the first space in the current time window. For example, the audio management module of the specified device can control the audio output device of the second electronic device to not output the sound signal picked up in the first space in the current time window.

[0192] It can be understood that controlling the audio output device of the second electronic device not to output the first spatial sound signal picked up in the current time window can mean muting the audio output device of the second electronic device. After muting the audio output device of the second electronic device, the audio output device of the second electronic device can still receive the sound signal transmitted by the first electronic device and the electronic device in the second space through the network, but does not output the received sound signal, that is, does not play the received sound signal.

[0193] Optionally, the audio input device of the electronic device in the second space picks up the sound signal in the second space, which can be transmitted to the audio management module of the plurality of electronic devices in the first space, and the audio management module of the plurality of electronic devices in the first space controls the audio output device of itself to output the sound signal in the second space.

[0194] Optionally, Figure 7 The implementation manners not mentioned in steps S301-S319 in the corresponding embodiments can refer to the specific implementation of the foregoing steps S101-S103 and the foregoing steps S201-S206, and will not be described here.

[0195] Through the device management method in the embodiments of the present application, echo path elimination and echo suppression of remote call in a shared space can be realized. By establishing a communication connection between each electronic device in the same space, transmission of voice feature data can be realized, thereby realizing identification of the first electronic device (i.e., the main electronic device) and the second electronic device (i.e., the auxiliary electronic device), and controlling the main electronic device to send the picked-up sound signal to other devices, and controlling the auxiliary electronic device not to send the picked-up sound signal to other devices, thereby reducing the far-end echo signal in the remote call. In addition, when the user of the main electronic device is speaking, the speaker of the auxiliary electronic device is controlled not to output the first spatial sound signal picked up by the main electronic device, thereby reducing the near-end echo signal in the remote call.

[0196] The hardware structure of the electronic device 100 will be introduced as follows:

[0197] Please refer to Figure 8 , Figure 8 The hardware structure of the electronic device 100 provided in the embodiments of the present application is shown in the following figure.

[0198] The electronic device 100 can include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headset jack 170D, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 can include a pressure sensor 180A, a gyroscope sensor 180B, a barometric sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0199] It can be understood that the structure shown in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 can include more or fewer components than shown, or combine certain components, or split certain components, or different arrangement of components. The components shown can be implemented in hardware, software, or a combination of software and hardware.

[0200] The processor 110 can include one or more processing units, for example: the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units can be independent devices, or can be integrated in one or more processors.

[0201] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to instruction operation codes and timing signals, and complete the control of fetching and executing instructions.

[0202] The processor 110 can also have internal memory for storing instructions and data. In some embodiments, the internal memory of the processor 110 is a cache. The cache can hold instructions and data recently used by the processor 110 or in the process of being used by the processor 110. If the processor 110 needs to use the instructions or data again, it can be directly called from the cache. This avoids repeated access and reduces the waiting time of the processor 110, thereby improving the efficiency of the system.

[0203] In some embodiments, the processor 110 can include one or more interfaces. The interfaces can include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0204] It can be understood that the interface connection relationship between the modules shown in the embodiments of the present application is only illustrative and does not constitute a structural limitation of the electronic device 100. In some other embodiments of the present application, the electronic device 100 can also use different interface connection methods or a combination of multiple interface connection methods in the above embodiments.

[0205] The charging management module 140 is configured to receive a charging input from a charger. The charger can be a wireless charger or a wired charger.

[0206] The power management module 141 is configured to connect the battery 142 and the charging management module 140 to the processor 110. The power management module 141 receives the input of the battery 142 and / or the charging management module 140 to supply power to the processor 110, the internal memory 121, the external memory, the display screen 194, the camera 193, and the wireless communication module 160, etc. In some other embodiments, the power management module 141 can also be arranged in the processor 110.

[0207] The wireless communication function of the electronic device 100 can be implemented by the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor, and the baseband processor, etc.

[0208] The antenna 1 and the antenna 2 are used for transmitting and receiving electromagnetic wave signals. Each antenna in the electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization of the antennas. For example, the antenna 1 can be multiplexed as a diversity antenna of a wireless local area network. In some other embodiments, the antennas can be used in combination with a tuning switch.

[0209] The mobile communication module 150 can provide a solution including 2G / 3G / 4G / 5G wireless communication applied on the electronic device 100. The mobile communication module 150 can include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves by the antenna 1, and perform filtering, amplification, etc. on the received electromagnetic waves, and transmit the processed electromagnetic waves to the modem processor for demodulation. The mobile communication module 150 can also amplify the signals modulated by the modem processor, and convert the signals into electromagnetic waves radiated by the antenna 1. In some embodiments, at least part of the functional modules of the mobile communication module 150 can be arranged in the processor 110. In some embodiments, at least part of the functional modules of the mobile communication module 150 can be arranged in the same device as at least part of the modules of the processor 110.

[0210] The modem processor can include a modulator and a demodulator. The modulator is used to modulate a low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. The low-frequency baseband signal processed by the baseband processor is transmitted to the application processor.

[0211] The wireless communication module 160 can provide a solution for wireless communication including Wi-Fi network, Bluetooth, BLE broadcast, global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. applied on the electronic device 100. The wireless communication module 160 can be one or more devices integrated with at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and transmits the processed signals to the processor 110. The wireless communication module 160 can also receive signals to be transmitted from the processor 110, frequency modulate them, amplify them, and radiate them as electromagnetic waves via the antenna 2.

[0212] In some embodiments, the antenna 1 and the mobile communication module 150 of the electronic device 100 are coupled, and the antenna 2 and the wireless communication module 160 are coupled, so that the electronic device 100 can communicate with a network and other devices through wireless communication technology.

[0213] The electronic device 100 implements a display function through a GPU, a display screen 194, and an application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 can include one or more GPUs that execute program instructions to generate or change display information.

[0214] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. In some embodiments, the electronic device 100 can include 1 or N display screens 194, N being a positive integer greater than 1. Among them, the display screen 194 can include an OLED screen.

[0215] Optionally, the display screen 194 can further include an OLED glass layer, an OLED light-emitting unit, a fingerprint identification sensor, a microlens array, etc. The display screen 194 supports optical under-screen fingerprint identification.

[0216] The electronic device 100 can implement a photographing function through an ISP, a camera 193, a video codec, a GPU, a display 194, and an application processor, etc. The ISP is used to process data fed back by the camera 193. The camera 193 is used to capture a still image or a video. The camera 193 can include a front camera and a rear camera, the front camera is located in a display area of the screen, and the rear camera is located in a back area of the screen. The digital signal processor is used to process a digital signal, which can process not only a digital image signal but also other digital signals. The video codec is used to compress or decompress a digital video. The electronic device 100 can support one or more video codecs.

[0217] The NPU is a neural-network (NN) calculation processor, which can quickly process input information by referring to a biological neural network structure, for example, referring to a transmission mode between human brain neurons, and can also continuously self-learn. Through the NPU, intelligent cognition of the electronic device 100, such as image recognition, face recognition, voice recognition, text understanding, etc., can be realized.

[0218] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to realize a data storage function. For example, compressed driver files and other files are saved in the external memory card.

[0219] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes various function applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 can include a program storage area and a data storage area. The program storage area can store an operating system, at least one application program required by a function (such as an audio playing function), etc. The data storage area can store data (such as a video stream) created during use of the electronic device 100, etc. In addition, the internal memory 121 can include a high-speed random access memory, and can also include a non-volatile memory, such as a flash memory device, etc. Optionally, the code of the audio management module can be read into the internal memory 121, and the processor 110 executes operations such as determining a first electronic device in a first space, sending a first control instruction to the first electronic device, sending a second control instruction to a second electronic device, etc. by running the instructions stored in the internal memory 121.

[0220] The electronic device 100 can implement audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the earphone interface 170D, and the application processor, etc. For example, outputting the picked-up voice signal of the speaker of the first space, outputting the voice signal transmitted from the remote space, etc. through the speaker 170A; collecting the voice signal of the speaker through the microphone 170C, etc.

[0221] The audio module 170 is configured to convert digital audio information into an analog audio signal output, and to convert an analog audio input into a digital audio signal. The audio module 170 can also be configured to encode and decode audio signals. In some embodiments, the audio module 170 can be disposed in the processor 110, or some functional modules of the audio module 170 can be disposed in the processor 110.

[0222] The speaker 170A, also referred to as a "loudspeaker", is configured to convert an audio electrical signal into a sound signal. The receiver 170B, also referred to as a "earpiece", is configured to convert an audio electrical signal into a sound signal. The microphone 170C, also referred to as a "microphone", "sound transducer", is configured to convert a sound signal into an electrical signal. The earphone interface 170D is configured to connect a wired earphone. The pressure sensor 180A is configured to sense a pressure signal, and can convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor 180A can be disposed in the display screen 194. The gyroscope sensor 180B can be configured to determine the motion posture of the electronic device 100. The barometric pressure sensor 180C is configured to measure the air pressure. The magnetic sensor 180D includes a Hall sensor. The acceleration sensor 180E can detect the acceleration of the electronic device 100 in each direction (generally three axes). The distance sensor 180F is configured to measure the distance. The proximity light sensor 180G can include, for example, a light emitting diode (LED) and a light detector. The ambient light sensor 180L is configured to sense the ambient light brightness. The fingerprint sensor 180H is configured to collect a fingerprint. The temperature sensor 180J is configured to detect the temperature. The touch sensor 180K, also referred to as a "touch panel". The touch sensor 180K can be disposed in the display screen 194, and the touch sensor 180K and the display screen 194 together form a touch screen, also referred to as a "touch panel". The touch sensor 180K is configured to detect a touch operation acting on or near the touch sensor 180K. The bone conduction sensor 180M can acquire a vibration signal. The keys 190 include a power key, a volume key, etc. The motor 191 can generate a vibration prompt. The indicator 192 can be an indicator light, and can be configured to indicate a charging state, a power change, and can also be configured to indicate a message, a missed call, a notification, etc. The SIM card interface 195 is configured to connect a SIM card.

[0223] The software structure of the electronic device is described as follows:

[0224] Figure 9FIG. 1 is a schematic diagram of a software structure of an electronic device 100 according to an embodiment of the present application. The software structure adopts a layered architecture, which divides the software into a plurality of layers, each of which has a clear role and division of labor. The layers communicate with each other through a software interface. In the embodiment of the present application, the electronic device is taken as an example of an electronic device installed with Windows (an operating system), and the electronic device can include an application layer (application, APP), a framework layer, a kernel layer (Kernal), and a hardware layer (Hardware).

[0225] The application layer can include a series of application packages. As shown in FIG. 2, the application packages can include applications such as a conference, a live broadcast, a camera, a call, and a video. Figure 9

[0226] The framework layer refers to a layer for audio management, for example, a layer for communication and interaction between audio management modules, for example, a layer that can control whether to transmit or not transmit a sound signal picked up by the electronic device to the network, and control transmission of voice feature data between different electronic devices, and the like. For example, the framework layer can include an audio management module and a local communication module, and the like.

[0227] The audio management module is used to extract voice feature data from a sound signal, analyze voice feature data of a plurality of electronic devices to determine a first electronic device, control the first electronic device to send a picked-up sound signal to other electronic devices, and control a second electronic device not to send a picked-up sound signal to other electronic devices, and the like.

[0228] The local communication module is used to realize data transmission between a plurality of electronic devices in the same space, for example, transmission of voice feature data, initial working state transmission of a loudspeaker and a microphone, and the like. The local communication module can include, for example, a WiFi module and a Bluetooth module, and the like. The WiFi module and the Bluetooth module are both used to transmit voice feature data and initial working states of a loudspeaker and a microphone.

[0229] The kernel layer is responsible for managing hardware resources of the system and providing necessary services to the application program. The kernel layer includes drivers of various hardware devices, which are responsible for controlling and operating the hardware devices. The kernel layer can include an audio input device driver and an audio output device driver.

[0230] The audio input device driver is used to drive an audio input device to pick up a sound signal and upload it to the network or not. For example, the audio input device driver can be a microphone driver.

[0231] The audio output device driver is used to drive an audio output device to output or not output a received sound signal. For example, the audio output device driver can be a loudspeaker driver.​

[0232] The hardware layer includes a plurality of hardware, such as an audio input device and an audio output device. The audio input device may, for example, refer to a microphone, and the audio output device may, for example, refer to a loudspeaker.

[0233] The audio input device, such as a microphone, is used to collect a sound signal in a space in which the electronic device is located, and the audio output device, such as a loudspeaker, is used to play a received sound signal, or the loudspeaker can also output a sound signal picked up by the microphone of the electronic device after amplification processing.

[0234] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another, for example, the computer instructions can be transferred from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server, data center, etc. that includes one or more available media sets. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk), etc.

[0235] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by a computer program to instruct the relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiments. The aforementioned storage medium includes ROM or random access memory (RAM), magnetic disk or optical disk, and various media that can store program codes.

[0236] In summary, the above only describes the embodiments of the technical scheme of the present application, and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made according to the disclosure of the present application shall be included in the protection scope of the present application.

Claims

1. A device management method characterized by, The method is applied to a specified device in a first space, the specified device being determined from at least two electronic devices in the first space, and the method comprises: The specified device acquires voice feature data sent by at least two electronic devices in the first space; the voice feature data of each electronic device is obtained by feature extraction on a sound signal picked up by the electronic device in a current time window; The specified device determines a first electronic device based on the voice feature data of the at least two electronic devices, and controls the first electronic device to send the picked-up sound signal to other electronic devices, and controls a second electronic device not to send the picked-up sound signal to the outside; the second electronic device is different from the first electronic device, and the first electronic device and the second electronic device both belong to the electronic devices in the first space; the sound signal pickup function of each electronic device in the first space is kept open, and in the process that the second electronic device does not send the picked-up sound signal to the outside, the second electronic device continues to pick up the sound signal of the first space in the current time window to obtain voice feature data and transmit the voice feature data to the specified device; After receiving the voice feature data transmitted by the second electronic device, the specified device determines a first electronic device in a next time window of the current time window based on the voice feature data; The specified device controls the audio output device of the second electronic device not to play the sound signal of the first space picked up in the current time window, and controls the audio output device of the second electronic device to play the sound signal of a second space picked up in the current time window; The method comprises: The specified device performs feature comparison on the voice feature data of the at least two electronic devices in the first space; In a case that the voice feature data of the at least two electronic devices in the first space are similar, the specified device determines an electronic device whose voice feature data meet a target transmission condition; the target transmission condition includes that the clarity of the sound signal of the electronic device is greater than a clarity threshold and the sound signal delay is less than a delay threshold; In a case that the voice feature data of at least two electronic devices meet the target transmission condition, the specified device determines the first electronic device from the electronic devices whose voice feature data meet the target transmission condition based on a specified factor; the specified factor includes one or more of network speed, remaining traffic, and device performance.

2. The method of claim 1, wherein, The method comprises: The specified device performs feature comparison on the audio feature information included in the voice feature data of the at least two electronic devices to determine a first similarity between the sound signals picked up by the at least two electronic devices; In a case that the first similarity is greater than or equal to a first similarity threshold, the specified device performs comparison on time indication information included in the voice feature data of the at least two electronic devices to determine a first electronic device that picks up the sound signal earliest among the at least two electronic devices.

3. The method of claim 2, wherein, The feature comparison is performed on audio feature information included in the voice feature data of the at least two electronic devices, and a first similarity between sound signals picked up by the at least two electronic devices is determined. The feature comparison is performed on audio feature information included in the voice feature data of the at least two electronic devices, and a first similarity between sound signals picked up by the at least two electronic devices is determined. The voice processing model is obtained by training an initial model based on similarity labels between at least two sample sound signals and sample similarities between the at least two sample sound signals, and the sample similarities are obtained by performing the feature comparison on sample audio feature information of the at least two sample sound signals based on the initial model.

4. The method of claim 1, wherein, The control of the audio output apparatus of the second electronic device not to play the sound signal of the first space picked up in the current time window includes: obtaining audio feature information of a sound signal received by the second electronic device; performing the feature comparison on the audio feature information of the sound signal received by the second electronic device and the audio feature information included in the voice feature data of the first electronic device, and determining a second similarity between the sound signal received by the second electronic device and the sound signal picked up by the first electronic device; in a case where the second similarity is greater than or equal to a second similarity threshold, determining that the sound signal received by the second electronic device is the sound signal of the first space picked up in the current time window, and controlling the audio output apparatus of the second electronic device not to play the received sound signal.

5. The method of claim 4, wherein, The method further includes: in a case where the second similarity is less than the second similarity threshold, determining that the sound signal received by the second electronic device is the sound signal of the second space picked up in the current time window, and controlling the audio output apparatus of the second electronic device to play the received sound signal.

6. The method of claim 1, wherein, The control of the first electronic device to send the picked-up sound signal to other electronic devices and the control of the second electronic device not to send the picked-up sound signal to the other electronic devices include: sending a first control instruction to the first electronic device; the first control instruction is used to control the first electronic device to send the picked-up sound signal to the other electronic devices; sending a second control instruction to the second electronic device; the second control instruction is used to control the second electronic device not to send the picked-up sound signal to the other electronic devices.

7. The method of claim 6, wherein, The sending of the first control instruction to the first electronic device includes: obtaining an initial working state of the first electronic device in the current time window; the initial working state is used to indicate whether to send the picked-up sound signal to the other electronic devices; in a case where the initial working state of the first electronic device indicates not to send the picked-up sound signal to the other electronic devices, sending the first control instruction to the first electronic device.

8. The method according to any one of claims 1 to 7, characterized in that, The method further includes: receiving a sound signal picked up by a third electronic device and sent by the third electronic device; In a case that the third electronic device is an electronic device in the second space, controlling an audio output device of the second electronic device to play the sound signal picked up by the third electronic device.

9. The method according to any one of claims 1 to 7, characterized in that, The receiving the voice feature data transmitted by the at least two electronic devices in the first space comprises: establishing a communication connection with the at least two electronic devices in the first space; obtaining the voice feature data respectively transmitted by the at least two electronic devices based on the communication connection.

10. An electronic device, comprising: The electronic device comprises one or more processors, a memory and a touch screen; the memory is configured to store program code; the processor is configured to execute the program code, so that the electronic device implements the method according to any one of claims 1-9.

11. A chip system applied to an electronic device, characterized by comprising: The chip system comprises at least one processor and an interface configured to receive instructions and transmit the instructions to the at least one processor; the at least one processor executes the instructions so that the electronic device executes the method according to any one of claims 1-9.

12. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, which, when executed by a processor, implements the method according to any one of claims 1-9.

13. A computer program product comprising computer programs / instructions, characterized in that, The computer program / instructions, when executed by a processor, implement the steps of the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Multi-video conference collaborative meeting method, terminal and storage medium

    CN115209083A