Information processing device, method and equipment

By setting up a plurality of audio acquisition units and target signals in the information processing device, the distance information between the sound source and the device is determined, and the problem of only obtaining angle information under the far-field model is solved, and the accuracy and user experience of audio processing are improved.

CN120103319APending Publication Date: 2025-06-06LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510237771.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art can only obtain the angle information of the sound source under the far-field model, but cannot obtain the distance information, which affects the processing of sound signals and user experience.

Method used

By setting a plurality of audio acquisition units in the information processing device, combining the target signal, the distance information between the target object and the target device is determined, thereby realizing audio processing based on the distance information.

Benefits of technology

The distance information between the sound source and the device is obtained, which improves the accuracy and user experience of audio processing, and enhances the audio quality received by the counterpart devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120103319A_ABST
    Figure CN120103319A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an information processing device, method and equipment, and the device comprises target equipment; the first component has a first position relation with the target equipment, and the first component comprises a plurality of audio acquisition units; under the first position relation, the distance between the first component and the target equipment is smaller than or equal to a first target threshold value; the second part has a second position relation with the target equipment, and the second part is used for collecting audio; the distance between the second component in the second position relation and the target equipment is greater than the distance between the first component in the first position relation and the target equipment; and the third part can transmit and collect a target signal, and the target signal is used for determining first distance information between the target object and the target equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to but is not limited to the field of information processing technology, and in particular to an information processing device, method and equipment. Background Art

[0002] According to the distance between the sound source and the microphone array, the sound field model can be divided into a near-field model and a far-field model, where the far-field model regards the sound wave as a plane wave. In the related art, under the far-field model, a sound source localization algorithm based on the far-field model can be used to localize the sound source. However, the sound source localization algorithm based on the far-field model can only obtain the angle information of the sound source but not the distance information of the sound source. Therefore, it is impossible to process the sound signal based on the distance information, which affects the user's satisfaction. Summary of the invention

[0003] In view of this, the embodiments of the present application at least provide an information processing device, method and equipment.

[0004] The technical solution of the embodiment of the present application is implemented as follows:

[0005] On the one hand, an embodiment of the present application provides an information processing device, the information processing device comprising:

[0006] Target device;

[0007] A first component, the first component has a first position relationship with the target device, and the first component includes a plurality of audio acquisition units; in the first position relationship, a distance between the first component and the target device is less than or equal to a first target threshold;

[0008] a second component, the second component has a second positional relationship with the target device, and the second component is used to collect audio; the distance between the second component and the target device in the second positional relationship is greater than the distance between the first component and the target device in the first positional relationship;

[0009] The third component is capable of transmitting and collecting a target signal, and the target signal is used to determine first distance information between the target object and the target device.

[0010] On the other hand, an embodiment of the present application provides an information processing method, the information processing method comprising:

[0011] The target device determines first distance information between the target object and the target device based on the target signal;

[0012] The target device executes the target strategy based on the first distance information;

[0013] The target device executes a target strategy based on the first distance information, including any of the following:

[0014] The target device determines an audio collection component from at least two audio collection components based on the first distance information, so as to obtain target audio based on the audio collection component; the target audio is used to output to the opposite end device;

[0015] When the target device determines that the target object is located in the target area based on the first distance information, the target device performs signal enhancement processing on the audio of the target object collected by the audio collection component;

[0016] The target device performs a first process on the audio of the target object collected by the audio collection component based on the first distance information to obtain the target audio, so that the opposite device can perform a second process on the target audio based on the type of the fifth component; the first process is different from the second process.

[0017] On the other hand, an embodiment of the present application provides an information processing device, including:

[0018] A determination module, used for the target device to determine first distance information between the target object and the target device based on the target signal;

[0019] An execution module, configured for the target device to execute a target strategy based on the first distance information;

[0020] Execution modules include any of the following:

[0021] A determination unit, configured for the target device to determine an audio collection component from at least two audio collection components based on the first distance information, so as to obtain target audio based on the audio collection component; the target audio is used to be output to the opposite end device;

[0022] An enhancement unit, configured to perform signal enhancement processing on the audio of the target object collected by the audio collection component when the target device determines that the target object is located in the target area based on the first distance information;

[0023] A processing unit is used for the target device to perform a first process on the audio of the target object collected by the audio collection component based on the first distance information to obtain the target audio, so that the opposite device can perform a second process on the target audio based on the type of the fifth component; the first process and the second process are different.

[0024] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The drawings herein are incorporated into the specification and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and are used together with the specification to illustrate the technical solution of the present application.

[0026] Figure 1A schematic diagram of an information processing device provided in an embodiment of the present application;

[0027] Figure 2 A schematic diagram of an angular spectrum provided in an embodiment of the present application;

[0028] Figure 3 A schematic diagram of an information processing device and a target object provided in an embodiment of the present application;

[0029] Figure 4 A schematic diagram of the locations of components of an information processing device provided in an embodiment of the present application;

[0030] Figure 5A A schematic diagram of an implementation flow of an information processing method provided in an embodiment of the present application;

[0031] Figure 5B A schematic diagram of an information processing device and a sound source object provided in an embodiment of the present application;

[0032] Fig. 6A A schematic diagram of an implementation of a conference system provided in an embodiment of the present application;

[0033] Figure 6B A schematic diagram of the implementation process of sound source localization and virtual space rendering combined with UWB positioning technology in a conference system provided in an embodiment of the present application;

[0034] Figure 7 A schematic diagram of the structure of an information processing device provided in an embodiment of the present application;

[0035] Figure 8 A hardware entity schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0036] In order to make the purpose, technical solutions and advantages of the present application clearer, the technical solutions of the present application are further elaborated in detail below in conjunction with the drawings and embodiments. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.

[0037] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0038] It should be pointed out that the terms "first\second\third" involved in the embodiments of the present application are merely to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.

[0039] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those skilled in the art in the field to which the embodiments of the present application belong. It should also be understood that terms such as those defined in general dictionaries should be understood to have meanings consistent with the meanings in the context of the prior art, and will not be interpreted with idealized or overly formal meanings unless specifically defined as here.

[0040] The present application embodiment provides an information processing device, such as Figure 1 As shown, including:

[0041] Target device 1;

[0042] A first component 2, the first component 2 and the target device 1 have a first position relationship, and the first component 2 includes a plurality of audio acquisition units; in the first position relationship, a distance between the first component 2 and the target device 1 is less than or equal to a first target threshold;

[0043] a second component 3, the second component 3 having a second positional relationship with the target device 1, and the second component 3 being used for collecting audio; a distance between the second component 3 and the target device 1 in the second positional relationship is greater than a distance between the first component 2 and the target device 1 in the first positional relationship;

[0044] The third component 4 can transmit and collect a target signal, and the target signal is used to determine the first distance information between the target object and the target device 1.

[0045] Here, the target device refers to a device with information processing functions. The audio acquisition unit refers to a component with audio acquisition and transmission functions. The first component refers to an array formed by a plurality of audio acquisition units located at different positions arranged in a certain shape rule, and this arrangement enables the array to sample sound signals propagating in space. The first target threshold is used to determine whether the distance between the first component and the target device can be ignored. It can be understood that the value of the first target threshold should be small. For example, the first target threshold can be 25 centimeters (cm). Multiple audio acquisition units refer to at least two audio acquisition units. The third component refers to a component used to obtain the distance between the second component and the target device. The target object can be an object that emits audio, for example, the target object can be a person or an animal.

[0046] In some embodiments, the first component has a first positional relationship with the target device: the first positional relationship may be that the first component is built into the target device; or the first component is located outside the target device, such as Figure 1 As shown, and connected to the target device wirelessly or wiredly.

[0047] In some implementations, multiple audio collection units may be arranged on the same straight line, thus forming a first component of a linear array, such as Figure 1 As shown; it can also be formed by arranging multiple audio collection units on the same plane, thus forming the first component of the planar array; the first component can also be formed by arranging multiple audio collection units in the same three-dimensional space, thus forming the first component of the three-dimensional array.

[0048] In some implementations, the multiple audio acquisition units in the first component may be implemented as devices or components with audio acquisition functions, such as microphones, recorders, and audio acquisition cards.

[0049] In some embodiments, the information processing device further includes a second component because when the place where the information processing device is located is large, if only the first component is provided in the information processing device, the quality of the audio of the target object far away from the first component collected by the first component will be poor. By adding the second component to the information processing device, not only the range of audio collection of the information processing device can be expanded, but also higher quality audio can be obtained from the audio collected by the first component and the audio collected by the second component for subsequent audio transmission or audio processing, thereby improving the call quality of the information processing device.

[0050] In some implementations, the second component has a second positional relationship with the target device: the second positional relationship may be that the second component is located outside the target device and is connected to the target device wirelessly or by wire.

[0051] In some implementations, the second component may be implemented as a device or component with an audio acquisition function, such as a microphone, a recorder, or an audio acquisition card.

[0052] In some implementations, the second component may further include multiple audio acquisition units, which may be implemented as devices or components with audio acquisition functions, such as microphones, recorders, and audio acquisition cards. In the case where the second component includes multiple audio acquisition units, when implementing the information processing method, the target device may obtain the audio acquisition unit corresponding to the middle position of the second component, and apply the audio acquisition unit to the information processing method related to the second component.

[0053] In some embodiments, the type of the second component may be the same as or different from the type of the first component, wherein the application place of the information processing device may be expanded by setting the type of the second component to be different from the type of the first component. For example, the second component is a satellite microphone and the first component is a host microphone. Since the satellite microphone is an audio acquisition device that transmits audio via a satellite and the host microphone is an audio acquisition device that transmits audio via electrical signals, the information processing device may be placed in a place that relies on satellite communication or electrical signal communication.

[0054] In some embodiments, by setting the distance between the second component and the target device to be greater than the distance between the first component and the target device, the audio collection range corresponding to the second component can be made farther than the audio collection range corresponding to the first component, thereby expanding the audio collection range of the information processing device.

[0055] In some embodiments, the target signal is used to determine the first distance information between the target object and the target device by: acquiring the distance between the second component and the target device based on the target signal, and determining the first distance information based on the distance between the second component and the target device.

[0056] In some embodiments, the third component may be provided on the target device, such as Figure 1 As shown; it can also be set in the second component; it can also be set in the target device and the second component, wherein the setting position of the third component is related to the type of the third component.

[0057] In some embodiments, the third component may be a position acquisition device, and the target signal is a signal for controlling the position acquisition device to acquire a position. In this way, the distance between the second component and the target device can be obtained based on the position data respectively acquired by the position acquisition devices disposed on the target device and the second component.

[0058] During implementation, in response to the triggering of the target signal, the first position data of the second component is acquired based on the position acquisition device on the second component; the second position data of the target device is acquired based on the position acquisition device on the target device; and the distance information is obtained based on the first position data and the second position data. The position acquisition device can be a Global Positioning System (GPS) receiver or a Beidou navigation satellite system receiver.

[0059] In some embodiments, the third component may also be an image acquisition device, in which case the target signal is a signal for controlling the image acquisition device to acquire an image. In this way, the distance between the second component and the target device can be obtained based on the image data of the image acquisition device set on the second component or the target device.

[0060] During implementation, the image acquisition device is provided on the second component for illustration, and in response to the triggering of the target signal, the image acquisition device is controlled to take a picture of the target device from the position of the second component, so that image data including the target device can be acquired, and then the distance between the second component and the target device is obtained based on the acquired image data and the device parameters of the image acquisition device. The image acquisition device can be a camera or a video camera.

[0061] In some embodiments, the third component may also be a first sensing device, in which case the target signal is a signal emitted by the sensing device. The receiving end and the transmitting end of the first sensor are located in the same device, so that the distance between the second component and the target device can be obtained based on the target signal emitted by the first sensing device disposed on the second component or the target device.

[0062] During implementation, the target signal emitted from the second component or the target device by the sensing device is reflected when encountering the target component or the second component, and a first time duration from the emission of the target signal to the reflection back to the target device or the second component can be obtained, thereby obtaining the distance between the second component and the target device through the first time duration. The first sensing device can be an ultrasonic device or a time of flight (TOF) device.

[0063] It can be understood that the distance between the first component and the target device is less than or equal to the first target threshold, that is, in the information processing device proposed in the present application, the distance between the first component and the target device is negligible, and the first distance information is used to characterize the distance between the target object and the target device, that is, the first distance information can also be used to characterize the distance between the first component and the target object.

[0064] In the embodiment of the present application, on the one hand, the information processing device includes a first component and a second component for collecting audio, which can expand the range of the information processing device for collecting audio, so that the quality of the collected audio can be higher by applying the information processing device in a larger place. On the other hand, compared with the sound source localization algorithm based on the far-field model in the related art, which can only obtain the angle information between the target object and the target device, the target signal emitted by the third component in the embodiment of the present application can also obtain the distance information between the target object and the target device, so that the collected audio can be processed based on the distance information, so that the quality of the audio finally output to the user is better, which improves the user's satisfaction.

[0065] In some embodiments, the third component includes a first subcomponent and a second subcomponent,

[0066] The first subcomponent is used to transmit a target signal to the second subcomponent, and the second subcomponent is used to receive the target signal;

[0067] The first subcomponent is arranged in the second component, and the second subcomponent is arranged in the target device; or, the first subcomponent is arranged in the target device, and the second subcomponent is arranged in the second component.

[0068] Here, the third component may be a sensor device; the target signal is a signal emitted by the sensor device.

[0069] In some embodiments, the third component may be a second sensing device, in which case the target signal is a signal transmitted by the sensing device, the first subcomponent is a transmitting end, and the second subcomponent is a receiving end. In this way, the distance between the second component and the target device can be obtained based on the target signal transmitted by the transmitting end. The second sensing device may be an ultra-wideband (UWB) device, an infrared sensing device, a radar device, or a through-beam sensor.

[0070] In the above embodiment, by installing the third component on the second component and the target device, the target signal emitted by the third component can be used to obtain the second distance information between the second component and the target device with higher accuracy.

[0071] In some embodiments, the target device is further configured to:

[0072] Determining second distance information between the second component and the target device based on the target signal;

[0073] Determine angle information between the target object and the target device;

[0074] Determine third distance information based on the first audio collected by the first component and the second audio collected by the second component; the third distance information is used to represent the distance between the target device and the target object, and the distance difference between the distance between the second component and the target object;

[0075] Based on the second distance information, the angle information and the third distance information, the first distance information and the fourth distance information are determined; the fourth distance information is used to characterize the distance between the second component and the target object.

[0076] In some embodiments, the target device is used to determine the second duration corresponding to the target signal in response to the second subcomponent receiving the target signal; the target device is used to determine the second distance information between the target device and the second component based on the second duration. After receiving the target signal, the second subcomponent determines the second duration corresponding to the target signal and transmits the second duration to the target device. Further, when the second subcomponent is set on the target device, the transmission duration of the second duration is the shortest, because the distance between the target device and the second subcomponent is the shortest.

[0077] In some embodiments, a first audio collected by a first component is obtained; the first audio is subjected to Fourier processing to obtain a first processed audio; the first audio is used as an input of a sound source localization algorithm based on a far-field model to determine the angle information between the target object and the target device. The sound source localization algorithm based on a far-field model may be a Time Difference of Arrival (TDOA) algorithm, a Generalized Cross-Correlation Phase Transform (GCC-PHAT) algorithm, or a Steered Response Power Phase Transform (SRP-PHAT) algorithm.

[0078] The following is an explanation of the specific implementation process using the SRP-PHAT algorithm as an example:

[0079] Step S100: Acquire the first audio collected by the first component.

[0080] Here, the first audio is non-silent audio.

[0081] The first audio frequency is represented by the following formula (1):

[0082] x(l,n)=[x 1 (l,n),x 2 (l,n),…,x M (l,n)] (1);

[0083] Wherein, l represents a frame index, n represents a sampling point index, and M represents the number of audio acquisition units in the first component.

[0084] Step S101: Perform STFT (Short-Time Fourier Transform) processing on the above formula (1) to obtain X(l, k). The representation of X(l, k) is shown in the following formula (2):

[0085] X(l,k)=[X 1 (l,k),X 2 (l,k),…,X M (l,k)] (2);

[0086] Wherein, k represents the frequency index, k=1, 2, ..., K; K represents the number of points for Fourier transform, which is generally selected as 1024.

[0087] Step S103: uniformly select N direction vectors d within a preset angle range n , find the SRP-PHAT value corresponding to each direction vector.

[0088] Here, a frame of data X(l,k) is in d n′ The expression of the SRP-PHAT value in the direction is shown in the following formula (3):

[0089]

[0090] Among them, d n′ represents the direction vector, n′=1,2,…,N;R ij [τ ij (d n′ )] represents the generalized cross-correlation function of the sound signals collected by the i-th and j-th microphones based on the weighted phase transformation; τ ij (d n′ ) represents the direction vector d n′ The arrival time difference at the i-th and j-th microphones respectively.

[0091] R ij [τ ij (d n′ )], see the following formula (4):

[0092]

[0093] Among them, Ω represents the rotation factor, * represents conjugation, e represents a natural constant, and || represents the modulus.

[0094] The expression of Ω is shown in the following formula (5):

[0095]

[0096] Among them, F s Indicates the sampling frequency of the first component or the second component, which is generally 16 kilohertz (kHZ).

[0097] τ ij (d n′ ), see the following formula (6):

[0098]

[0099] Among them, r i represents the rectangular coordinate vector of the i-th host microphone, r j represents the rectangular coordinate vector of the jth host microphone, c represents the sound speed, and ‖‖ represents the 2-norm of the vector.

[0100] Step S104: Obtain the SRP-PHAT values ​​of the L frame data in N direction vectors respectively, and smooth the SRP-PHAT values ​​of the L frame data in the same direction vector to obtain the smoothed SRP-PHAT value of each direction vector, and perform maximum calculation on the smoothed SRP-PHAT values ​​of the N direction vectors to obtain d peak , and determine d peak The corresponding angle information.

[0101] In some implementations, the smoothing process of the SRP-PHAT value of the vector in the same direction of the L frame data may be implemented based on a sampling moving average method or a sampling exponential smoothing method.

[0102] In some embodiments, the direction vector d n′ In spatial coordinates, it can be decomposed into the pitch angle and azimuth angle θ, so we can get The relationship table of the three is laid out on a two-dimensional plane for visualization to obtain an angular spectrum diagram, thereby obtaining d peak .

[0103] The following is an explanation of the first component implemented by a linear array. Since the linear array cannot distinguish the pitch angle, the pitch angle is Fixed to 90°, such as Figure 2 As shown, the horizontal axis represents the azimuth, the range of the horizontal axis is (0°, 180°), and the vertical axis represents the SRP-PHAT value. It can be seen that Figure 2 There is a peak value d peak , that is to say, the orientation information of the target object is:

[0104] In some embodiments, based on the first audio and the second audio, the delay information between the first audio collected by the first component and the second audio collected by the second component is determined by a method of mutual interference delay, and the third distance information is determined based on the delay information.

[0105] In some embodiments, based on the positions of the target device, the second component, and the target object, the target device is connected to the second component, the second component is connected to the target object, and the target object is connected to the target device to form a positional relationship, thereby determining the first distance information and the fourth distance information based on the positional relationship, the second distance information, the angle information, and the third distance information.

[0106] The first component formed by a linear array is used for explanation. Since a linear array cannot distinguish between pitch angles, θ is used in the following calculations. peak Determined as the angle information between the target object and the target device, such as Figure 3 As shown, the target device 1, the second component 3 and the target object 5 form a triangle. Based on the characteristics of the triangle, the following set of equations can be constructed, see the following formula (7):

[0107]

[0108] Among them, D 0 Indicates the second distance information, D m Indicates the first distance information, D s Indicates the fourth distance information, D 2 Represents the third distance information, θ peak Indicates angle information.

[0109] Solving the above formula (7), we can get D m and D s , see the following formulas (8) and (9).

[0110]

[0111] In the above embodiment, first, the third distance information is obtained based on the first audio collected by the first component and the second audio collected by the second component; then, based on the angle information and the third distance information between the target object and the target device, the first distance information between the target object and the target device, and the fourth distance information between the second component and the target object are obtained. In this way, based on the first distance information and the fourth distance information, better quality audio can be obtained from the first audio and the second audio, and the better quality audio can be applied to improve call quality and thereby improve user satisfaction.

[0112] In some embodiments, the target device is further configured to:

[0113] Determine, from the first audio, a third audio collected by a target audio collection unit; the target audio collection unit is an audio collection unit located in the middle of the multiple audio collection units constituting the first component;

[0114] Based on the third audio and the second audio, determine first delay information; the first delay information is used to represent the delay between the target acquisition unit acquiring the third audio and the second component acquiring the second audio;

[0115] Based on the first time delay information, third distance information is determined.

[0116] In some embodiments, when the number of the multiple audio collection units is an even number, based on the physical layout formed by the multiple audio collection units, any audio collection unit that is closest to the center point of the physical layout may be used as the target audio collection unit; when the number of the multiple audio collection units is an odd number, the audio collection unit at the center point of the physical layout formed by the multiple audio collection units may be used as the target audio collection unit.

[0117] Exemplarily, taking the case of a linear array as an example, when the number of the multiple audio acquisition units is 7, the target audio acquisition unit is the 4th in the queue; when the number of the multiple audio acquisition units is 6, the target audio acquisition unit is the 3rd or 4th in the queue.

[0118] In some implementations, the first delay information is determined based on the third audio and the second audio by a mutual interference delay method.

[0119] The specific implementation process of obtaining the first delay information is introduced below:

[0120] Step S200: Since the target audio acquisition unit and the second component are located at different positions, when collecting audio of the same target object, the target audio acquisition unit and the second component have different durations for collecting audio, so it is necessary to perform time delay processing on the target audio acquisition unit.

[0121] First, the maximum number of delayed sampling points for delay processing of the target audio acquisition unit is obtained, see the following formula (10):

[0122]

[0123] Among them, T max represents the maximum number of delayed sampling points, c represents the speed of sound, which is generally 342 m / s, and D max Indicates the maximum positioning distance of the satellite microphone, which can be determined according to the actual use environment. For example, if the information processing device is set in a conference room, D max Generally it will not exceed 20 meters (m).

[0124] Then, based on the maximum number of delayed sampling points, the third audio x c (l,n) is subjected to time delay processing to obtain the second processed audio x c(l,nt), where t represents the delay, t = 1, 2, ...T max .

[0125] Step S201: Process the second audio x c (l,nt) and the second audio y(l,n) are processed by K-point STFT respectively, and we get and Y(l,k).

[0126] here, The expression of is shown in the following formula (11):

[0127]

[0128] Step S202: Determine the third distance information by using the mutual interference delay method.

[0129] here, The first mutual interference coefficient Coh between Y(l,k) 0 , see the following formula (12):

[0130]

[0131] in, For the calculation method of , see the following formula (13):

[0132]

[0133] Among them, E[] means to find the expected value.

[0134] From the above formula (13), multiple mutual interference coefficients can be obtained, and then the multiple mutual interference coefficients are subjected to maximum value processing to obtain the first target mutual interference coefficient, and the first delay information corresponding to the first target mutual interference coefficient is determined, see the following formula (14):

[0135]

[0136] Among them, max means maximum value processing.

[0137] In some implementations, the third distance information is determined based on the first delay information, as shown in the following formula (15):

[0138]

[0139] In the above embodiment, the first time delay information can be determined based on the third audio and the second audio collected by the target audio collection unit, so that the first time delay information can be applied to the process of determining the third distance information, so as to finally obtain the first distance information and the fourth distance information based on the third distance information.

[0140] In some embodiments, the information processing apparatus includes a fourth component, the fourth component includes a third subcomponent and a fourth subcomponent, the third subcomponent is used to output a fourth audio, the fourth subcomponent is used to output a fifth audio, the fourth audio and the fifth audio have different signal frequencies, and the target device is further used to:

[0141] Based on the fourth audio collected by the first component and the second component, determine second time delay information between the fourth audio collected by the second component and the fourth audio collected by the first component;

[0142] Based on the fifth audio collected by the first component and the second component, determine third time delay information between the fifth audio collected by the second component and the fifth audio collected by the first component;

[0143] When it is determined that the second component is located at the target position based on the second delay information, the third delay information and the second target threshold, it is verified whether the second distance information meets the target condition based on the second delay information, the third delay information, and the fifth distance information between the third subcomponent and the fourth subcomponent; the target position is related to the position of the target device.

[0144] In some embodiments, the fourth component may be an audio output device or component such as a speaker, a horn, a sound box, or a headset, and the third subcomponent and the fourth subcomponent may be audio output devices or components for outputting different channels. For example, when the fourth component is a horn, the third subcomponent may be a left channel speaker, and the fourth subcomponent may be a right channel speaker.

[0145] In some embodiments, when controlling the third subcomponent to output the fourth audio and simultaneously controlling the fourth subcomponent to output the fifth audio, the first component and the second component are controlled to simultaneously collect the fourth audio and the fifth audio, so that the audio collected by the first component obtained by the target host does not distinguish the corresponding first target audio source, that is, the fourth audio and the fifth audio are not distinguished.

[0146] In some embodiments, the fourth audio and the fifth audio have different signal frequencies, so that the first target sound source can be filtered out from the audio collected by the microphone through a filter. In implementation, the frequency of the fourth audio can be set at 1kHz to 4kHz, and the frequency of the fifth audio can be set at 4kHz to 7kHz.

[0147] Here, the first collected audio collected by the target audio unit is obtained from the fourth audio and the fifth audio output by the fourth component collected by the first component, and the first collected audio is the fourth audio and the fifth audio output by the fourth component.

[0148] The first collected audio is filtered through a bandpass filter to obtain a first sub-collected audio and a second sub-collected audio, see the following formula (16):

[0149]

[0150] Among them, x′ c (l,n) represents the first collected audio, x′ cL (l,n) represents the first sub-collected audio obtained by filtering, x′ cR (l,n) represents the second sub-collected audio obtained by filtering, l represents the frame index, and n represents the sampling point index.

[0151] The second collected audio collected by the second component is obtained, where the second collected audio is the fourth audio and the fifth audio output by the fourth component.

[0152] The second collected audio is filtered through a bandpass filter to obtain a third sub-collected audio and a fourth sub-collected audio, see the following formula (17):

[0153]

[0154] Among them, y′(l,n) represents the second collected audio, y′ L (l,n) represents the third sub-collected audio obtained by filtering, y′ R (l,n) represents the fourth sub-collected audio obtained by filtering.

[0155] The specific implementation process of the second delay information and the third delay information is introduced below:

[0156] Step S300: Obtain the maximum number of delayed sampling points for delay processing of the target audio acquisition unit, see the above formula (10).

[0157] Step S301: Delay processing is performed on the above formula (16) based on the maximum number of delayed sampling points, see the following formula (18):

[0158]

[0159] Where t = 1, 2, ..., T max .

[0160] Step S302: Perform K-point STFT processing on the first collected audio and the second collected audio respectively to obtain a third processed audio and the fourth processed audio Y′(l,k).

[0161] here, Refer to the following formula (19) for the expression of:

[0162]

[0163] Step S303: Obtaining a second mutual interference coefficient Coh′ between the third processed audio and the fourth processed audio 0.

[0164] Here, the second mutual interference coefficient Coh′ 0 , see the following formula (20):

[0165]

[0166] For the calculation method of , see the following formula (21):

[0167]

[0168] For the calculation method of , see the following formula (22):

[0169]

[0170] Step S304: Based on formula (20), multiple mutual interference coefficients can be obtained, and then the multiple mutual interference coefficients are subjected to maximum value processing to obtain the second target mutual interference coefficient and the third target mutual interference coefficient, and the second delay information corresponding to the second target mutual interference coefficient and the third delay information corresponding to the third target mutual interference coefficient are determined.

[0171] Here, the method for determining the second delay information is as follows:

[0172]

[0173] Among them, T L Indicates the second delay information.

[0174] The method for determining the third delay information is as follows:

[0175]

[0176] Among them, T R Indicates the third delay information.

[0177] In some implementations, the target position is related to the position of the target device, which may mean that the target position corresponds to a midpoint between the positions of the target device, such as Figure 3 As shown, it can be seen that the target position is located on the extension line of the middle point of the target device.

[0178] In some embodiments, a first difference between the second delay information and the third delay information is obtained; based on the comparison between the first difference and a preset second target threshold, when the absolute value of the first difference is less than the second target threshold, it is determined that the second component is located at the target position.

[0179] Based on the comparison between the first difference and the preset second target threshold, see the following formula (25):

[0180] △T=|T L -T R |<δ 0 (25);

[0181] Among them, △T represents the absolute value of the first difference, δ 0 Indicates the second target threshold.

[0182] It can be understood that the target position is any point on the extension line of the midpoint of the target device, and the fourth subcomponent and the third subcomponent are located on both sides of the target device. If the second component is located at the target position, and since the distance between the first component and the target device is negligible, the time corresponding to the first component collecting the fourth audio and the fifth audio can be ignored. Therefore, the second delay information corresponding to the fourth audio output by the third subcomponent received by the second component and the third delay information corresponding to the fifth audio output by the fourth subcomponent received by the second component should be approximately the same. That is, when the above formula (25) is satisfied, it is determined that the second component is located at the target position.

[0183] In the above embodiment, the second time delay information between the second component collecting the fourth audio and the first component collecting the fourth audio, and the third time delay information between the second component collecting the fifth audio and the first component collecting the fifth audio are determined. In this way, the third time delay information and the second time delay information can be used to verify whether the second component is located at the target position, and when the second component is located at the target position, the accuracy of the distance obtained based on the mutual interference delay method can be determined.

[0184] In some embodiments, the target device is further configured to:

[0185] When the second component is not located at the target position, a first prompt is output; the first prompt is used to prompt adjustment of the position of the second component.

[0186] In the above embodiment, when it is determined that the second device is not located at the target position, a prompt message is output to remind the target user that the current position of the second component is abnormal and the position of the second component needs to be adjusted. By adjusting the position of the second component to be normal, the accuracy of the distance obtained based on the mutual interference delay method can be determined.

[0187] In some embodiments, the target device is further configured to:

[0188] Determine sixth distance information between the third subcomponent and the target device based on the second delay information and the fifth distance information;

[0189] Determine seventh distance information between the fourth subcomponent and the target device based on the third delay information and the fifth distance information;

[0190] When it is determined that the second distance information meets the target condition based on the sixth distance information, the seventh distance information and the third target threshold, the third distance information is determined based on the first audio collected by the first component and the second audio collected by the second component.

[0191] like Figure 4 As shown in FIG. 1 , the positional relationship between the third subcomponent 61, the fourth subcomponent 62 and the second component 3 is shown. It can be seen that the distance between the second component 3 and the third subcomponent 61 is D L , the distance between the second component 3 and the fourth subcomponent 62 is D R , the distance between the fourth subcomponent 62 and the third subcomponent 61 is D 0 .

[0192] The method for determining the sixth distance information is as follows:

[0193]

[0194] Among them, D 1 Indicates the fifth distance information.

[0195] The seventh distance information is determined in the following formula (27):

[0196]

[0197] The following describes the specific implementation process of determining that the second distance information meets the target condition:

[0198] Step S401: Determine the tenth distance information D′ based on the fifth distance information, the sixth distance information and the seventh distance information 0 , wherein the tenth distance information is the distance between the second component and the target device determined by the mutual interference delay method.

[0199] When implemented, the tenth distance information D′ 0 Refer to the following formula (28) for the determination of

[0200]

[0201] Step S402: Through D′ 0 , and D measured by UWB 0 , determine that the second distance information meets the target condition.

[0202] During implementation, a second difference between the tenth distance information and the second distance information is obtained; based on the comparison between the second difference and a preset third target threshold, it is determined that the second distance information meets the target condition.

[0203] Based on the comparison between the second difference and the preset third target threshold, see the following formula (29):

[0204] △D=|D 0 -D′ 0 |<δ 1 (29);

[0205] Among them, △D represents the absolute value of the second difference, δ 1 Indicates the third target threshold.

[0206] According to formula (29), when the absolute value of the second difference is less than the third target threshold, it is determined that the second distance information meets the target condition.

[0207] In some embodiments, under the condition that the target condition is met, it is shown that the tenth distance information obtained based on the mutual interference delay method is correct, so that it can be determined that the distance information determined based on the mutual interference delay is credible, and the method can be applied to the determination process of the third distance information.

[0208] In the above embodiment, the sixth distance information between the third subcomponent and the target device, and the seventh distance information between the fourth subcomponent and the target device are first obtained, and then based on the sixth distance information and the seventh distance information, it is determined that the second distance information meets the target condition. This indicates that the distance information determined based on the mutual interference delay has a high accuracy, so that the method can be applied to the determination process of the third distance information, thereby improving the calculation accuracy of the third distance information.

[0209] The present application embodiment provides an information processing method, such as Figure 5A As shown, it may include step S510 and step S520:

[0210] Step S510: the target device determines first distance information between the target object and the target device based on the target signal;

[0211] Step S520: the target device executes the target strategy based on the first distance information;

[0212] Step S520 includes any of the following:

[0213] Step S521: the target device determines an audio collection component from at least two audio collection components based on the first distance information, and obtains target audio based on the audio collection component; the target audio is used to output to the opposite device;

[0214] Here, the peer device refers to a device with an information processing function, wherein the peer device may be of the same type as the target device or may be of a different type from the target device. The target policy refers to an application policy of the first distance information.

[0215] In some implementations, the first component and the second component are used as audio collection components for illustration, and the fourth distance information between the second component and the target object can be obtained; the audio corresponding to the smaller distance between the first distance information and the fourth distance information is used to determine the target audio. That is, the audio collected by the microphone close to the speaker is selected and sent to the opposite device (remote device).

[0216] In some embodiments, the first component and the second component are used as audio collection components for illustration, and fourth distance information between the second component and the target object can be obtained; when the first distance information and the fourth distance information meet the target distance condition, the first audio is determined as the target audio, otherwise the second audio is determined as the target audio.

[0217] Step S522: when the target device determines that the target object is located in the target area based on the first distance information, the target device performs signal enhancement processing on the audio of the target object collected by the audio collection component;

[0218] Here, the target area may be an area for performing signal enhancement, and the target area may be set with the location of the target device as a reference location.

[0219] In some embodiments, the target area can be implemented as a virtual sound pickup wall area. In this way, by setting a virtual sound pickup wall in the target place where the information processing device is located, the audio collected by the audio collection component can be divided, so that the audio located in the virtual sound pickup wall area, that is, the target area, is regarded as call audio, and the audio outside the virtual sound pickup wall area is regarded as noise audio. The audio collected by the audio collection component can be classified. By classifying the audio collected by the audio collection component, corresponding processing can be performed based on different types of audio. On the one hand, signal enhancement processing is performed on the call audio, which can effectively improve the audio quality of the call audio. On the other hand, signal elimination or reduction processing can be performed on the noise audio, which can reduce the interference of noise audio on the call audio and improve the call quality.

[0220] In some embodiments, signal enhancement processing of the audio of the target object can be achieved based on multi-level automatic gain control. During implementation, the sound intensity of the collected audio of the target object can be obtained, and the gain can be adjusted based on the sound intensity of the collected audio of the target object. This can make the collected audio of the target object at different sound intensities at a stable level, which can not only achieve signal enhancement processing of the collected audio of the target object, but also improve the clarity and recognizability of the collected audio of the target object.

[0221] In some embodiments, signal enhancement processing of the audio of the target object can be achieved based on echo cancellation processing. During implementation, when the audio collected from the target object includes the output audio of the fourth component, echo cancellation linear processing and / or echo cancellation nonlinear processing can be performed on the audio collected from the target object, which can not only achieve signal enhancement processing of the audio collected from the target object, but also improve the clarity of the audio collected from the target object.

[0222] In some embodiments, signal enhancement processing of the target object's audio can be achieved based on enhancing the audio's spectrum. During implementation, the spectrum of the collected target object's audio is acquired, and the spectrum of the collected target object's audio is enhanced by spectrum equalization, spectrum subtraction, or a spectrum enhancement algorithm based on machine learning, thereby achieving signal enhancement processing of the collected target object's audio.

[0223] In some implementations, eleventh distance information between the target area and the target device is obtained, wherein the eleventh distance information is the maximum distance between the target area and the target device, so that when the first distance information is less than the eleventh distance information, it is determined that the target object is located in the target area.

[0224] In some implementations, the target area may be preset by a user, or may be obtained based on preset angle information and distance information.

[0225] Step S523: the target device performs a first process on the audio of the target object collected by the audio collection component based on the first distance information to obtain the target audio, so that the opposite device can perform a second process on the target audio based on the type of the fifth component; the first process and the second process are different.

[0226] Here, the first processing may be a processing of performing spatial rendering on the audio, and the second processing may be a processing of performing anti-crosstalk on the audio.

[0227] In some implementations, first, the audio of the target object collected by the audio collection component is subjected to time delay processing; then, the audio after time delay processing is subjected to Fourier transform processing to obtain a first transformed audio. Finally, the first transformed audio is subjected to a third processing to obtain a fifth processed audio; and the fifth processed audio is subjected to the first processing to obtain the target audio.

[0228] During implementation, the method for determining the fifth processed audio Z(l, k) is as follows:

[0229]

[0230] Wherein, Z(l,k) represents the first transformed audio; a is an empirical value, which can be summarized through a large amount of measured data. When implemented, the perceptual evaluation of speech quality (PESQ) or other evaluation indicators of the audio collected by the first component and the second component can be obtained at a certain distance respectively, and the PESQ or other evaluation indicators output by the two components are balanced to obtain the value of a. Generally, a>1.

[0231] According to the above formula (30), for X(l, k), the third processing can be a spatial filter. When implemented, the spatial filter can be determined based on a beamforming algorithm, so that the sound signal collected by the host microphone array can be in the target direction θ peak The beamforming algorithm may be a minimum variance distortionless response (MVDR) beamforming algorithm.

[0232] For Y(l,k), the third processing is a denoising filter, which can be implemented based on a single-channel denoising algorithm, such as an optimally modified log-spectral amplitude spectrum estimation (OMLSA) algorithm, a log minimum mean square error (Log Minimum Mean Square Error, Log-MMSE) algorithm, and the like.

[0233] In some implementations, performing the first processing on the fifth processed audio to obtain the target audio may be: expanding the single-channel Z(l,k) into dual-channel data Z L (l,k) and Z R (l, k), see the following formula (31) and formula (32):

[0234] Z L (l,k)=Z(l,k)H L (θ peak ,k)b(D m ) (31);

[0235] Z R (l,k)=Z(l,k)H R (θ peak ,k)b(D m ) (32);

[0236] Among them, H L and H R is the virtual orientation filter coefficient, b(D m) is the attenuation coefficient associated with the first distance information.

[0237] The method for determining the virtual orientation filter coefficients is shown in the following formulas (33) and (34):

[0238]

[0239] Among them, h L It represents the lateral head-related impulse response of the third subcomponent. The lateral head-related impulse response of the third subcomponent can be realized by a head-related impulse response database.

[0240]

[0241] Among them, h R Represents the lateral head-related impulse response of the fourth subcomponent.

[0242] In the implementation of the present application, the audio collected by the audio collection component is processed based on the first distance information, so that the audio effect ultimately output to the user can be better, thereby optimizing the user experience.

[0243] In some embodiments, in the above step S500, the fifth component is used to output audio; and the peer device can perform a second process on the target audio based on the type of the fifth component, including:

[0244] In response to the fifth component being of the first type, the peer device acquires a sixth audio collected by an audio collection component connected to the peer device;

[0245] When determining that the sixth audio is a silent audio, the opposite end device performs a second process on the target audio based on the preset eighth distance information;

[0246] When determining that the sixth audio is non-silent audio, the opposite device determines ninth distance information between the opposite device and the sound source object corresponding to the sixth audio; and performs second processing on the target audio based on the ninth distance information.

[0247] Here, the first type is used to characterize that the fifth component outputs the target audio in the form of an external speaker.

[0248] In some embodiments, the fifth component also includes a fifth sub-component and a sixth sub-component. In this way, when the fifth component is of the first type, if the opposite device outputs the received target audio through the fifth component, the sound output by the fifth sub-component will not only be received by the user's left ear, but also by the right ear, resulting in crosstalk. Therefore, the opposite device needs to perform a second processing on the target audio so that the audio output by the fifth sub-component is transmitted to the left ear, and the audio output by the sixth sub-component is transmitted to the right ear, thereby achieving directional rendering of the target audio.

[0249] In some embodiments, the type of the fifth component also includes the second type, which is used to characterize that the fifth component outputs the target audio in a non-external form. In this way, when the fifth component is of the second type, if the opposite device outputs the received target audio through the fifth component, the sound output by the fifth subcomponent is received by the user's left ear, and the sound output by the sixth subcomponent is received by the user's right ear.

[0250] In some embodiments, determining whether the sixth audio is silent audio may be based on the energy of the sixth audio, may be based on the short-time zero-crossing rate of the sixth audio, or may be determined by VAD (Voice Activity Detection).

[0251] In some embodiments, the eighth distance information may be determined based on the distance between the sixth subcomponent and the seventh subcomponent, for example, the eighth distance information may be half of the distance between the sixth subcomponent and the seventh subcomponent. In other embodiments, the eighth distance information may be determined based on the occupied area of ​​the location where the opposite device is located.

[0252] In some implementations, the method for determining the ninth distance information may refer to the method for determining the first distance information described above.

[0253] In some implementations, the target audio is subjected to a second process based on the ninth distance information.

[0254] When implemented, Figure 5B As shown, the distance H from the vertical line of the sound source object 9 to the opposite device 7 can be obtained, see the following formula (35):

[0255] H=D′ m sin(θ′ peak ) (35);

[0256] Among them, D′ m Represents the ninth distance information, θ′ peak Indicates the angle information between the sound source object and the peer device.

[0257] The distance between the fifth subcomponent 81 and the foot of the perpendicular is as follows:

[0258] D′ L =0.5D′ 1 +D′ m cos(θ′ peak ) (36);

[0259] Among them, D′ 1 Indicates the distance between the fifth and sixth subcomponents.

[0260] The distance between the sixth subcomponent 82 and the vertical foot is as follows:

[0261] D′ R =0.5D′ 1 -D′ m cos(θ′ peak ) (37);

[0262] Based on the above formulas (36) and (37), the angle θ between the sound source object 9 and the fifth subcomponent 81 can be obtained as L , and the angle θ between the sound source object 9 and the sixth subcomponent 82 R , see the following formulas (38) and (39):

[0263]

[0264] Furthermore, the anti-crosstalk output of the fifth subcomponent is given by the following formula (40):

[0265]

[0266] The anti-crosstalk output of the sixth subcomponent is shown in the following formula (41):

[0267]

[0268] The dual-channel frequency domain signal O obtained by the above formula (40) and the above formula (41) is L (l,k) and O R (l,k) is overlapped and added to convert it into the time domain to obtain o L (l,k) and o R (l,k), thus outputting o through the fifth subcomponent L (l,k), output o through the sixth subcomponent R (l,k).

[0269] In some embodiments, when the sixth audio is non-silent audio, by obtaining the ninth distance information between the sound source object corresponding to the non-silent audio and the opposite device, and the ninth distance information is used to characterize the position of the sound source object, the sound source object is also the listener corresponding to the target audio, so that the target audio is processed for the second time based on the position of the listener, so that the sound source object has a better spatial experience when hearing the audio finally output by the fifth component, and the experience satisfaction of the sound source object with the target audio is improved. That is, in the scenario where someone speaks in the space where the opposite device is located, anti-crosstalk processing is performed at the detected speaking position to ensure that the speaking participants can experience a better audio space rendering effect. Of course, the essential purpose of the above embodiment is to ensure that each participant in the space can experience a better audio space rendering effect. Therefore, if all participants in the space can be identified by other methods, such as person detection, face tracking, etc., and then anti-crosstalk processing is performed at the position of each participant, it also belongs to the protection scope of this application.

[0270] In some embodiments, when the sixth audio is silent audio, the listener corresponding to the target audio cannot be obtained, and the second processing of the target audio cannot be performed. In this scenario, in order to achieve the second processing of the target audio, an anti-crosstalk position can be preset in the place where the opposite device is located, and the eighth distance information is used to characterize the preset anti-crosstalk position. In this way, the second processing of the target audio can be achieved in a scene where no one is speaking, so that the target listener in the place where the opposite device is located has a better spatial experience when receiving the audio output through the fifth component, thereby improving the target listener's experience satisfaction with the target audio, that is, in a scene where no one speaks in the space where the opposite device is located, anti-crosstalk processing can be performed at the default position of the space, such as the best seat in the conference room, because the best seat in the conference room has participants by default, and anti-crosstalk at this position can at least ensure that the participants in the best position can experience a better audio space rendering effect.

[0271] In the above embodiment, when the peer device is of the first type, if the peer device outputs the received target audio through the fifth component, the sound output by the fifth subcomponent will be received not only by the user's left ear, but also by the right ear, resulting in crosstalk. Therefore, after the peer device performs a second processing on the target audio based on the sound source object or the eighth distance information, the audio output by the fifth subcomponent is transmitted to the left ear, and the audio played by the sixth subcomponent is transmitted to the right ear, thereby achieving directional rendering of the target audio.

[0272] The following describes the application of the embodiments of the present application in actual scenarios.

[0273] Conference systems often focus only on the intelligibility and stability of speech. By applying dereverberation technology and automatic gain control (AGC) technology, as well as adopting a single-channel output method for speech signals, the sound received at the far end is relatively consistent regardless of the position of the near-end speaker, ignoring the sense of presence brought by the coordination of sound and picture display.

[0274] In addition, according to the distance between the sound source and the microphone array, the sound field model can be divided into a near-field model and a far-field model. The far-field model regards the sound wave as a plane wave. The microphone array used in the conference system has a small aperture. If the sound source is far away from the microphone array, the sound signal will propagate to the microphone array in the form of a plane wave. Therefore, the acoustic model corresponding to the conference system is a far-field model.

[0275] In the related art, under the far-field model, a sound source localization algorithm based on the far-field model can be used to locate the sound source. However, the sound source localization algorithm based on the far-field model can only obtain the angle information of the sound source but cannot obtain the distance information of the sound source. Therefore, it is impossible to accurately render the spatial sense of the sound signal, which affects the user's satisfaction with the conference system.

[0276] Based on this, the present application embodiment proposes a conference system, such as Fig. 6A As shown, it may include: a satellite microphone 10 (i.e. the above-mentioned second component), a host microphone array 11 (i.e. the above-mentioned first component), and a host 12 (i.e. the above-mentioned target device), wherein the satellite microphone is connected to the host via a wire, and the host microphone array is located outside the host and is connected to the host via a wire, so that the conference system can be applied in larger places, and the quality of the collected audio can be improved through the joint action of the host microphone array and the satellite microphone.

[0277] Based on the above technical problems, the embodiment of the present application provides a method for sound source localization and virtual space rendering in a conference system combined with UWB positioning technology, such as Figure 6B As shown, it may include steps S601 to S608:

[0278] Step S601: The user activates the call function.

[0279] Step S602: Obtain the distance D between the satellite microphone and the host through UWB positioning detection 0 .

[0280] Step S603: Check whether the placement of the satellite microphone is centered and the result D′ of the mutual interference method distance measurement 0 Is it credible?

[0281] Here, the conference system includes a dual-channel speaker (i.e., the third component) disposed on both sides of the host, and the dual-channel speaker can simultaneously emit regular sound signals, so that the placement of the satellite microphone can be checked to be centered by controlling the dual-channel speaker to emit sound signals of different frequency bands. The dual-channel speaker includes a left channel speaker and a right channel speaker.

[0282] Next, we will introduce the specific implementation process of verifying whether the satellite microphone is placed in the center:

[0283] Step S6031: Control the left channel speaker (i.e., the third sub-component) to emit a first sound signal (i.e., the third audio frequency) of 1kHz to 4kHz, and the right channel speaker to emit a second sound signal (i.e., the fourth audio frequency) of 4kHz to 7kHz (i.e., the fourth sub-component), and when the dual-channel speakers emit sound signals, control the host middle microphone (i.e., the target audio acquisition unit) and the satellite microphone (i.e., the second component) to simultaneously acquire the sound signals emitted by the dual-channel speakers. The host middle microphone refers to the microphone located in the middle position of the host microphone array (i.e., the first component).

[0284] Step S6032: After the host middle microphone and the satellite microphone complete the collection of the sound signals emitted by the dual-channel speakers, the sound signals collected by the host middle microphone and the satellite microphone are filtered by bandpass filters according to the different frequency bands corresponding to the first sound signal and the second sound signal, respectively, to obtain the first sound signal emitted by the left channel speaker and the second sound signal emitted by the right channel speaker collected by the host middle microphone and the satellite microphone, respectively, and the filtered first sound signal emitted by the left channel speaker and the second sound signal emitted by the right channel speaker are stored in a buffer.

[0285] It should be noted that the above processing process can be performed in real time or intermittently, and the duration of the sound signal emitted by the dual-channel speaker should be at least 2 seconds (s).

[0286] Here, the sound signal collected by the middle microphone of the host is filtered, that is, the bandpass filtered signal of the middle microphone of the host (that is, the first collected audio), see the above formula (16).

[0287] The sound signal collected by the satellite microphone is filtered to obtain the bandpass filtered signal of the satellite microphone (ie, the second collected audio), see the following formula (17).

[0288] Step S6033: Based on the first sound signal emitted by the left channel speaker collected by the filtered host middle microphone and the satellite microphone, determine the first time delay (i.e., the above-mentioned second time delay information) between the first sound signal collected by the host middle microphone and the first sound signal collected by the satellite microphone; and, based on the second sound signal emitted by the filtered host middle microphone and the satellite microphone and collected by the right channel speaker, determine the second time delay (i.e., the above-mentioned third time delay information) between the second sound signal collected by the host middle microphone and the second sound signal collected by the satellite microphone.

[0289] During implementation, when the satellite microphone and the host's middle microphone collect sound signals, the first collection time length used by the host's middle microphone to collect sound signals is different from the second collection time length used by the satellite microphone to collect sound signals. Therefore, it is necessary to perform time delay processing on the host's middle microphone so that the first collection time length used by the host's middle microphone to collect sound signals is aligned with the second collection time length used by the satellite microphone to collect sound signals.

[0290] First, the maximum number of delayed sampling points for delay processing of the host intermediate microphone is as shown in the above formula (10). max The above formula (16) is subjected to time delay processing, see the above formula (18).

[0291] Then, the above formula (18) is processed by STFT to obtain The above formula (16) is subjected to STFT processing to obtain Y′(l,k), where k=1, 2, ..., K, and K represents the number of FFT points. For the expression of , see the following formula (19).

[0292] Finally, the first delay and the second delay are obtained by using a method of mutually interfering with the delay. For details, please refer to the processing process of the above formulas (20) to (24).

[0293] Step S6033: Based on the first delay and the second delay, verify whether the placement position of the satellite microphone is centered.

[0294] During implementation, a first difference between the first time delay and the second time delay is obtained; based on the comparison between the first difference and a preset first threshold, it is determined whether the placement position of the satellite microphone is centered.

[0295] Wherein, based on the comparison between the first difference and the preset first threshold, see the above formula (25):

[0296] According to formula (25), when the absolute value of the first difference is less than the first threshold (i.e., the second target threshold), the satellite microphone is placed in the middle of the host, and when the absolute value of the first difference is greater than or equal to the first threshold, the satellite microphone is not placed in the middle of the host. Further, when the satellite microphone is not placed in the middle of the host, a prompt message is output to prompt the user that the current position of the satellite microphone is abnormal and the position of the satellite microphone needs to be adjusted.

[0297] Next, the specific implementation process of whether the results of the mutual interference method ranging are credible is introduced:

[0298] Step S6034: Based on the first time delay, determine the distance D between the satellite microphone and the left channel speaker L , and based on the second delay, determine the distance D between the satellite microphone and the right channel speaker R .

[0299] Here, D L For the determination method of D, refer to the above formula (26); R For the determination method of , see the following formula (27).

[0300] Step S6035: Through D R , D L and D 1 , find the distance D′ between the satellite microphone and the midpoint of the host 0 .

[0301] Here, D′ 0 For the determination of , see the above formula (28).

[0302] Step S6036: Through D′ 0 , and D measured by UWB 0 , determine whether the results of the mutual interference method distance measurement are reliable.

[0303] When implementing, obtain D′ 0 With D 0 Based on the comparison between the second difference and the preset second threshold (ie, the third target threshold), it is determined whether the result of the mutual interference method ranging is credible, see the above formula (29).

[0304] According to formula (29), when the absolute value of the second difference is less than the second threshold, the result of the mutual interference method distance measurement is credible, and when the absolute value of the second difference is greater than or equal to the second threshold, the result of the mutual interference method distance measurement is unreliable. Further, when the result of the mutual interference method distance measurement is unreliable, the current calculation result is discarded, and a prompt message is output to prompt the adjustment of the position of the satellite microphone so that the sound channel between the satellite microphone and the speaker is unobstructed.

[0305] Step S604: obtaining the angle between the host and the sound source (ie the above-mentioned angle information) through the host microphone array.

[0306] Here, the angle between the host and the sound source can be obtained by a sound source localization algorithm, wherein the sound source localization algorithm can be a TDOA algorithm, a GCC-PHAT algorithm, or a SRP-PHAT algorithm.

[0307] The following is an explanation of the specific implementation process using the SRP-PHAT algorithm as an example:

[0308] Step S6041: Control the host microphone array and the satellite microphones to collect the sound signal of the current place, and process the sound signal of the current place through voice activity detection (Voice Activity Detection, VAD) to obtain a non-silent sound signal.

[0309] Here, the non-silent third sound signal (ie, the first audio signal) collected by the host microphone array is referred to in the above formula (1).

[0310] The above formula (1) is subjected to STFT processing to obtain X(l,k). The representation of X(l,k) is shown in the above formula (2).

[0311] Step S6042: Use the frequency domain signal X(l,k) to determine the direction of the sound source, and select N direction vectors d uniformly in the space according to a preset angle range, such as 0-180. n′ , find the SRP-PHAT value corresponding to the N direction vectors, and then find the direction vector d corresponding to the maximum value in the obtained SRP-PHAT value peak .

[0312] During implementation, a frame of data X(l,k) is stored in d n′ For the calculation of the SRP-PHAT value of the direction, refer to the above formulas (3) to (6).

[0313] Step S6043: Smooth the SRP-PHAT result of the L frame (may be an average method), and search for the maximum value to obtain its corresponding direction d peak .

[0314] Here, the microphone collects the voice signal, which is collected frame by frame and sent to the host. When the host processes the voice signal, it can obtain the SRP-PHAT result once for each frame. Since each SRP-PHAT result may be interfered with, such as jitter, it is usually necessary to cache the results of multiple frames for smoothing, and then obtain a relatively stable result in the end. Among them, L is obtained based on the parameter adjustment. For example, in this embodiment, the smoothing results of 10 frames are selected for output.

[0315] In some embodiments, the direction vector d n′ In spatial coordinates, it can be decomposed into the pitch angle and azimuth angle θ, so that we can obtain The relationship table of the three is laid out on a two-dimensional plane for visualization, and an angular spectrum diagram can be obtained, thereby obtaining d peak .

[0316] like Figure 2 As shown, the horizontal axis represents the azimuth, and the vertical axis represents the SRP-PHAT result. It can be seen that Figure 2 There is a peak value d peak , that is to say, the location of the sound source (i.e. the target object mentioned above) is:

[0317] It should be noted that when the host microphone array is implemented as a linear array, since it is impossible to distinguish the pitch angle, the pitch angle is used as the Fixed at 90°.

[0318] It should be noted that in actual application, the method can be implemented by selecting the frequency band with the strongest speech to reduce the influence of reverberation and noise. The frequency band with the strongest speech can be within 100 Hz to 4000 Hz.

[0319] Step S605: Based on the third sound signal obtained by the VAD, determine the distance from the sound source to the host center microphone and the distance from the sound source to the satellite microphones.

[0320] Here, from the third sound signal, the fourth sound signal (ie, the third audio signal) collected by the middle microphone of the host is obtained as x c (l,n), the fourth sound signal (i.e. the second audio signal mentioned above) collected by the satellite microphone is y(l,n).

[0321] The following is a specific implementation process for obtaining the distance difference between the sound source and the host intermediate microphone, and between the sound source and the satellite microphone: Step S6051: Obtain the maximum number of delayed sampling points, see the above formula (10). Step S6052: Based on the maximum number of delayed sampling points, calculate x c(l,n) is processed with time delay to obtain x c Step S6053: For x c (l,nt) and y(l,n) are processed by STFT respectively, and we get and Y(l,k), where For the representation of , refer to the above formula (11). Step S6054: Determine the distance difference between the sound source and the host intermediate microphone, and between the sound source and the satellite microphone by using the mutual interference delay method, refer to the above formula (12) to formula (15). Further, if Figure 3 As shown in the figure, the triangular relationship between the host, satellite microphone and sound source is shown, see the above formula (7). Solving the above formula (7), we can get D m and D s , see the above formulas (8) and (9).

[0322] Step S606: According to the distance D m and D s , select the sound signal to be processed.

[0323] Here, the method for determining the sound signal to be processed (i.e. the fifth processed audio signal mentioned above) can be found in the following formula (30).

[0324] Step S607: performing uplink virtual spatial sound effect processing (ie, first processing) on ​​the sound signal to be processed.

[0325] During implementation, the single-channel Z(l,k) is expanded into a dual-channel sound signal (i.e., the target audio signal) Z L (l,k) and Z R (l, k), see the following formulas (31) to (34).

[0326] Step S608: performing downlink virtual azimuth anti-crosstalk processing (ie, the second processing described above) on the dual-channel sound signal.

[0327] Here, when the remote playback device is a non-external type such as headphones or earmuffs, the dual-channel sound signal obtained by processing in step S607 can be transmitted to the remote end and played after being received by the remote device. When the remote playback device (i.e., the fifth component mentioned above) is an external type (i.e., the first type mentioned above) such as a speaker or a sound bar, since the dual-channel sound signal is played externally, the sound played by the remote left channel speaker (i.e., the fifth sub-component) will be received not only by the left ear, but also by the right ear, resulting in crosstalk. Therefore, it is necessary to perform anti-crosstalk processing on the dual-channel sound signal obtained by processing in step S607 in the downstream, so that the sound played by the remote left channel speaker is transmitted to the left ear, and the sound played by the remote right channel speaker (i.e., the sixth sub-component) is transmitted to the right ear, thereby realizing the directional rendering of the sound signal.

[0328] During implementation, the remote device (i.e. the above-mentioned opposite device) is used for explanation, such as Figure 5B As shown, the remote device performs anti-crosstalk processing on the dual-channel sound signal, see formula (35) to formula (41).

[0329] The dual-channel frequency domain signal O obtained by the above formula (39) and the above formula (40) is L (l,k) and O R (l,k) is overlapped and added to convert it into the time domain to obtain o L (l,k) and o R (l,k), so that the far left speaker plays o L (l,k), plays through the far right speaker o R (l,k).

[0330] It should be noted that anti-crosstalk processing can only be performed on one position, so the position of the main speaker in the remote venue is given priority for anti-crosstalk processing. If no one is speaking in the venue, the middle position between the far-end right channel speaker and the far-end left channel speaker can be considered as the anti-crosstalk position. Furthermore, since the sense of direction rendering is the icing on the cake of the conference, giving users a sense of technology, if the anti-crosstalk position is not completely accurate, it only weakens the sense of direction of the sound signal, and does not affect the playback of the sound signal. Therefore, considering the anti-crosstalk of the main speaker position can meet the needs of sense of direction rendering.

[0331] Based on the above embodiments, the following beneficial effects can be obtained: 1. Through UWB positioning technology, there is no need to worry about objects on the conference table blocking the view, and the distance between the speaker and the satellite microphone can be accurately measured, so that the distance calculation accuracy reaches the centimeter level. 2. Use dual-channel sound signals to verify whether the satellite microphone is placed in the middle of the host, and use the UWB ranging results to verify whether the sound wave detection results are credible. 3. Use the host microphone array to locate the angle of the sound source, and measure the time difference from the sound source to the host microphone and the sound source to the satellite microphone, so as to construct a set of equations to obtain the distance between the sound source and the host. 4. The ultimate goal of this solution is to render the audio collected by the microphone with a virtual sense of space based on the angle information between the sound source and the host and the distance between the sound source and the host, and perform anti-crosstalk processing at the other end according to the position of the audience, so as to present a perfect sense of immersion to the audience at the other end.

[0332] Based on the foregoing embodiments, the embodiments of the present application provide an information processing device, which includes the modules included and the units included in the modules, which can be implemented by a processor in an electronic device; of course, it can also be implemented by a specific logic circuit; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP) or a field programmable gate array (FPGA), etc.

[0333] Figure 7 A schematic diagram of the structure of an information processing device provided in an embodiment of the present application is shown in FIG. Figure 7 As shown, the information processing device 700 includes:

[0334] A determination module 710, configured for the target device to determine first distance information between the target object and the target device based on the target signal;

[0335] An execution module 720, configured for the target device to execute a target strategy based on the first distance information;

[0336] Execution modules include any of the following:

[0337] A determination unit, configured for the target device to determine an audio collection component from at least two audio collection components based on the first distance information, so as to obtain target audio based on the audio collection component; the target audio is used to be output to the opposite end device;

[0338] An enhancement unit, configured to perform signal enhancement processing on the audio of the target object collected by the audio collection component when the target device determines that the target object is located in the target area based on the first distance information;

[0339] A processing unit is used for the target device to perform a first process on the audio of the target object collected by the audio collection component based on the first distance information to obtain the target audio, so that the opposite device can perform a second process on the target audio based on the type of the fifth component; the first process and the second process are different.

[0340] In some embodiments, the information processing device 700 also includes: an acquisition module, used by the opposite device to acquire the sixth audio collected by the audio collection component connected to the opposite device in response to the fifth component being of the first type; a first processing module, used by the opposite device to perform a second processing on the target audio based on the preset eighth distance information when the sixth audio is determined to be silent audio; a second processing module, used by the opposite device to determine the ninth distance information between the opposite device and the sound source object corresponding to the sixth audio when the sixth audio is determined to be non-silent audio; and perform a second processing on the target audio based on the ninth distance information.

[0341] The description of the above information processing device embodiment is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. In some embodiments, the functions or modules included in the information processing device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiment. For technical details not disclosed in the embodiments of the present application, please refer to the description of the method embodiment of the present application for understanding.

[0342] It should be noted that in the embodiment of the present application, if the above-mentioned information processing device is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, which is stored in a storage medium, including a number of instructions to enable an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods of each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific hardware, software or firmware, or any combination of hardware, software, and firmware.

[0343] An embodiment of the present application provides an electronic device, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, some or all of the steps in the above method are implemented.

[0344] The embodiment of the present application provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, some or all of the steps in the above method are implemented. The computer-readable storage medium can be transient or non-transient.

[0345] An embodiment of the present application provides a computer program, including a computer-readable code. When the computer-readable code runs in an electronic device, a processor in the electronic device executes some or all of the steps in the above method.

[0346] The embodiment of the present application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and when the computer program is read and executed by a computer, some or all of the steps in the above method are implemented. The computer program product can be implemented specifically by hardware, software or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium, and in other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK) and the like.

[0347] It should be noted here that the description of the various embodiments above tends to emphasize the differences between the various embodiments, and the same or similar aspects can be referenced to each other. The description of the above device, storage medium, computer program and computer program product embodiments is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. For technical details not disclosed in the embodiments of the device, storage medium, computer program and computer program product of this application, please refer to the description of the method embodiment of this application for understanding.

[0348] It should be noted that Figure 8 is a schematic diagram of a hardware entity of an electronic device in an embodiment of the present application, such as Figure 8 As shown, the hardware entity of the electronic device 800 includes: a processor 801, a communication interface 802 and a memory 803, wherein:

[0349] The processor 801 generally controls the overall operation of the electronic device 800 .

[0350] The communication interface 802 enables the electronic device to communicate with other terminals or servers through a network.

[0351] The memory 803 is configured to store instructions and applications executable by the processor 801, and can also cache data to be processed or processed by the processor 801 and each module in the electronic device 800 (for example, image data, audio data, voice communication data, and video communication data), which can be implemented by flash memory (FLASH) or random access memory (Random Access Memory, RAM). Data transmission between the processor 501, the communication interface 802 and the memory 803 can be carried out through the bus 804.

[0352] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the serial number of each step / process mentioned above does not mean the order of execution, and the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application. The serial numbers of the embodiments of the present application mentioned above are for description only and do not represent the advantages and disadvantages of the embodiments.

[0353] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0354] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.

[0355] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0356] In addition, all functional units in the embodiments of the present application may be integrated into one processing unit, or each unit may be a separate unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0357] A person skilled in the art can understand that all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, etc., various media that can store program codes.

[0358] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can essentially or in other words, the part that contributes to the relevant technology can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods of each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0359] The above are only implementation methods of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application.

Claims

1. An information processing device, comprising: Target device; A first component, the first component has a first positional relationship with the target device, and the first component includes a plurality of audio acquisition units; In the first position relationship, the distance between the first component and the target device is less than or equal to a first target threshold; a second component, the second component having a second positional relationship with the target device, the second component being used to collect audio; the distance between the second component and the target device in the second positional relationship is greater than the distance between the first component and the target device in the first positional relationship; A third component is capable of transmitting and collecting a target signal, wherein the target signal is used to determine first distance information between a target object and the target device.

2. The information processing device according to claim 1, wherein the third component includes a first subcomponent and a second subcomponent. The first subcomponent is used to transmit the target signal to the second subcomponent, and the second subcomponent is used to receive the target signal; The first subcomponent is arranged in the second component, and the second subcomponent is arranged in the target device; or, the first subcomponent is arranged in the target device, and the second subcomponent is arranged in the second component.

3. Based on the information processing apparatus according to claim 1 or 2, the target device is also used for: determining second distance information between the second component and the target device based on the target signal; Determining angle information between the target object and the target device; Determine third distance information based on the first audio collected by the first component and the second audio collected by the second component; The third distance information is used to represent the distance difference between the distance between the target device and the target object and the distance between the second component and the target object; The first distance information and the fourth distance information are determined based on the second distance information, the angle information and the third distance information; the fourth distance information is used to characterize the distance between the second component and the target object.

4. Based on the information processing apparatus according to claim 3, the target device is further used for: Determine, from the first audio, a third audio collected by a target audio collection unit; the target audio collection unit is an audio collection unit located in the middle of the multiple audio collection units constituting the first component; Determine first delay information based on the third audio and the second audio; The first delay information is used to represent the delay between the target acquisition unit acquiring the third audio and the second component acquiring the second audio; The third distance information is determined based on the first time delay information.

5. The information processing apparatus according to claim 3, wherein the information processing apparatus comprises a fourth component, the fourth component comprises a third subcomponent and a fourth subcomponent, the third subcomponent is used to output a fourth audio, the fourth subcomponent is used to output a fifth audio, the fourth audio and the fifth audio have different signal frequencies, and the target device is further used to: Based on the fourth audio collected by the first component and the second component, determine second time delay information between the fourth audio collected by the second component and the fourth audio collected by the first component; Based on the fifth audio collected by the first component and the second component, determine third time delay information between the second component collecting the fifth audio and the first component collecting the fifth audio; In a case where it is determined that the second component is located at a target position based on the second delay information, the third delay information, and the second target threshold, verifying whether the second distance information meets a target condition based on the second delay information, the third delay information, and fifth distance information between the third subcomponent and the fourth subcomponent; The target location is related to the location of the target device.

6. Based on the information processing apparatus according to claim 5, the target device is further used for: When the second component is not located at the target position, a first prompt is output; the first prompt is used to prompt adjustment of the position of the second component.

7. Based on the information processing apparatus according to claim 5, the target device is further used for: Determine sixth distance information between the third subcomponent and the target device based on the second delay information and the fifth distance information; Determine seventh distance information between the fourth subcomponent and the target device based on the third delay information and the fifth distance information; When it is determined that the second distance information meets the target condition based on the sixth distance information, the seventh distance information and the third target threshold, the third distance information is determined based on the first audio collected by the first component and the second audio collected by the second component.

8. An information processing method, comprising: The target device determines first distance information between the target object and the target device based on the target signal; The target device executes a target strategy based on the first distance information; The target device executes a target strategy based on the first distance information, including any of the following: The target device determines an audio collection component from at least two audio collection components based on the first distance information, so as to obtain target audio based on the audio collection component; the target audio is used to output to the opposite end device; When the target device determines that the target object is located in the target area based on the first distance information, the target device performs signal enhancement processing on the audio of the target object collected by the audio collection component; The target device performs a first process on the audio of the target object collected by the audio collection component based on the first distance information to obtain the target audio, so that the opposite device can perform a second process on the target audio based on the type of the fifth component; the first process and the second process are different.

9. The method according to claim 8, wherein the fifth component is used to output audio; The peer device can perform a second process on the target audio based on the type of the fifth component, including: In response to the fifth component being of the first type, the peer device acquires a sixth audio collected by an audio collection component connected to the peer device; When determining that the sixth audio is a silent audio, the opposite-end device performs the second processing on the target audio based on the preset eighth distance information; When determining that the sixth audio is non-silent audio, the opposite-end device determines ninth distance information between the opposite-end device and the sound source object corresponding to the sixth audio; and performs the second processing on the target audio based on the ninth distance information.

10. An information processing device, comprising: A determination module, configured for a target device to determine first distance information between a target object and the target device based on a target signal; An execution module, configured for the target device to execute a target strategy based on the first distance information; The execution module includes any of the following: a determining unit, configured for the target device to determine an audio collecting component from at least two audio collecting components based on the first distance information, so as to obtain target audio based on the audio collecting component; The target audio is used to output to the peer device; an enhancement unit, configured to perform signal enhancement processing on the audio of the target object collected by the audio collection component when the target device determines that the target object is located in the target area based on the first distance information; A processing unit is used for the target device to perform a first process on the audio of the target object collected by the audio collection component based on the first distance information to obtain the target audio, so that the opposite device can perform a second process on the target audio based on the type of the fifth component; the first process and the second process are different.