A method, apparatus, system, and electronic device for generating a recording signal

By obtaining the direct sound signals of the left and right ear pickup devices of the stereo earphones, determining the sound source orientation and rendering and generating recording signals, the problem that the stereo recording signal cannot express spatial information is solved, and a better user experience is achieved.

CN117812523BActive Publication Date: 2025-08-01HONOR DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211167230.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-23
Publication Date
2025-08-01
Estimated Expiration
2042-09-23

AI Technical Summary

Technical Problem

The stereo recording signal cannot accurately express the spatial information of the sound signal, resulting in the inability to restore the spatial sense during playback, and the user experience is poor.

Method used

By obtaining the direct sound signal of the audio signal after the left ear pickup device and the right ear pickup device, the sound source orientation information is determined, and the direct sound signal is rendered based on the information to generate a recording signal.

Benefits of technology

The generated recording signal fully integrates the direction information of the sound source, and after playback, it can achieve the hearing effect of the human ear and improve the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117812523B_ABST
    Figure CN117812523B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a method, apparatus, system, and electronic device for generating a recording signal. The method includes: obtaining a first audio signal obtained by a left ear pickup device after picking up a sound signal to be processed, and a second audio signal obtained by a right ear pickup device after picking up the sound signal to be processed; obtaining a first direct sound signal of the first audio signal and a second direct sound signal of the second audio signal; determining first azimuth information of a sound source that emits the sound signal to be processed according to the first direct sound signal and the second direct sound signal; and rendering the first direct sound signal and the second direct sound signal according to the first azimuth information to generate a recording signal. The recording signal obtained by this method fully integrates the azimuth information of the sound source, can well express the spatial information of the sound signal, and after being played back, can achieve the effect of human ear hearing, providing a better user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of terminals, and in particular, to a method, device, system, and electronic device for generating a recording signal. Background Art

[0002] Currently, in the application scenario of stereo recording, it is usually that a user wears stereo headphones, such as true wireless stereo (TWS) headphones, and then the left ear sound pickup device (such as a microphone) in the user's left ear and the right ear sound pickup device in the user's right ear of the stereo headphones simultaneously pick up the sound signals in the surrounding environment. Then, an electronic device communicatively connected to the stereo headphones saves the audio signals obtained after sound pickup, so as to obtain a stereo recording signal.

[0003] However, due to the limited spatial feeling that the hardware devices of stereo headphones can present, the recording signals obtained in the above manner cannot accurately express the spatial information of the sound signals, cannot achieve the effect of the human ear's listening feeling, and cannot restore the sense of space during subsequent playback, resulting in poor user experience. Summary of the Invention

[0004] Embodiments of the present application provide a method, device, system, and electronic device for generating a recording signal, so as to solve the problem that the recording signals obtained by current stereo recording cannot accurately express the spatial information of the sound signals, resulting in the inability to restore the sense of space during subsequent playback and poor user experience.

[0005] In a first aspect, embodiments of the present application provide a method for generating a recording signal, and the method includes:

[0006] Obtain a first audio signal and a second audio signal, where the first audio signal is an audio signal obtained by a left ear sound pickup device picking up a to-be-processed sound signal, and the second audio signal is an audio signal obtained by a right ear sound pickup device picking up the to-be-processed sound signal;

[0007] Obtain a first direct sound signal of the first audio signal and a second direct sound signal of the second audio signal;

[0008] Determine first azimuth information of a sound source that emits the to-be-processed sound signal according to the first direct sound signal and the second direct sound signal;

[0009] Render the first direct sound signal and the second direct sound signal according to the first azimuth information to generate a recording signal of the to-be-processed sound signal.

[0010] In this way, through this method, the electronic device can separately extract the direct sound signals from the first audio signal obtained after the left ear pickup device picks up the sound signal to be processed and the second audio signal obtained after the right ear pickup device picks up the same sound signal to be processed, so as to obtain the first direct sound signal and the second direct sound signal. Then, the azimuth information of the sound source that emits the sound signal can be determined according to the first direct sound signal and the second direct sound signal. After that, the first direct sound signal and the second direct sound signal can be rendered according to the azimuth information to obtain the recording signal of the sound signal to be processed. It can be seen that the recording signal obtained through this method fully integrates the azimuth information of the sound source, can well express the spatial information of the sound signal, and after playback, it can achieve the effect of human ear hearing, and the user experience is better.

[0011] In a possible implementation manner, the determining of the first azimuth information of the sound source that emits the sound signal to be processed according to the first direct sound signal and the second direct sound signal includes:

[0012] Generating binaural difference information according to the first direct sound signal and the second direct sound signal, where the binaural difference information includes a first interaural time difference and / or a first interaural intensity difference;

[0013] Determining the first azimuth information according to the binaural difference information.

[0014] In this way, after the electronic device extracts the first direct sound signal of the first audio signal and the second direct sound signal of the second audio signal, it can obtain the difference information between the two ears according to these two direct sound signals. After that, the azimuth information of the sound source can be determined according to the binaural difference information, so that the subsequent generated recording signal fully integrates the azimuth information of the sound source, can well express the spatial information of the sound signal, and after playback, the user's hearing experience is better.

[0015] In a possible implementation manner, the binaural difference information is the first interaural time difference, and the determining of the first azimuth information according to the binaural difference information includes:

[0016] Obtaining a first preset azimuth information and a second preset azimuth information corresponding to the first interaural time difference;

[0017] Determining a second azimuth information, where the azimuth indicated by the second azimuth information is the azimuth jointly indicated by the first preset azimuth information and the second preset azimuth information;

[0018] Obtaining a third audio signal, where the third audio signal is an audio signal obtained after a first auxiliary pickup device picks up the sound signal to be processed, and the first auxiliary pickup device is located at the azimuth indicated by the second azimuth information and is not in the same plane as the left ear pickup device and the right ear pickup device;

[0019] Obtain a first energy and a second energy, where the first energy is the energy of the audio signal corresponding to the second azimuth information in the first audio signal and the second audio signal, and the second energy is the energy of the third audio signal;

[0020] Determine the first azimuth information according to the first energy, the second energy, the first preset azimuth information, and the second preset azimuth information.

[0021] In this way, in a scenario where the frequencies of the first audio signal and the second audio signal are neither too high nor too low, the azimuth information of the sound source can be accurately obtained according to the interaural time difference, so that the subsequent generated recording signal can more accurately express the spatial information of the sound signal, and after playback, the user's listening experience is better.

[0022] In a possible implementation, the binaural difference information is the first interaural intensity difference, and determining the first azimuth information according to the binaural difference information includes:

[0023] Obtain a third preset azimuth information and a fourth preset azimuth information corresponding to the first interaural intensity difference;

[0024] Determine a third azimuth information, where the azimuth indicated by the third azimuth information is the azimuth jointly indicated by the third preset azimuth information and the fourth preset azimuth information;

[0025] Obtain a fourth audio signal, where the fourth audio signal is an audio signal obtained after the second auxiliary sound pickup device picks up the sound signal to be processed, and the second auxiliary sound pickup device is located at the azimuth indicated by the third azimuth information and is not in the same plane as the left ear sound pickup device and the right ear sound pickup device;

[0026] Obtain a third energy and a fourth energy, where the third energy is the energy of the audio signal corresponding to the third azimuth information in the first audio signal and the second audio signal, and the fourth energy is the energy of the fourth audio signal;

[0027] Determine the first azimuth information according to the third energy, the fourth energy, the third preset azimuth information, and the fourth preset azimuth information.

[0028] In this way, in a scenario where the frequencies of the first audio signal and the second audio signal are high, the azimuth information of the sound source can be accurately obtained according to the interaural intensity difference, so that the subsequent generated recording signal can more accurately express the spatial information of the sound signal, and after playback, the user's listening experience is better.

[0029] In a possible implementation, the binaural difference information is the first interaural time difference and the first interaural intensity difference. Determining the first azimuth information according to the binaural difference information includes:

[0030] Obtain a first preset azimuth information and a second preset azimuth information corresponding to the first interaural time difference, and a third preset azimuth information and a fourth preset azimuth information corresponding to the first interaural intensity difference;

[0031] Determine a second azimuth information and a third azimuth information. The azimuth indicated by the second azimuth information is the azimuth jointly indicated by the first preset azimuth information and the second preset azimuth information, and the azimuth indicated by the third azimuth information is the azimuth jointly indicated by the third preset azimuth information and the fourth preset azimuth information;

[0032] Obtain a third audio signal and a fourth audio signal. The third audio signal is an audio signal obtained after a first auxiliary sound pickup device picks up the sound signal to be processed, and the fourth audio signal is an audio signal obtained after a second auxiliary sound pickup device picks up the sound signal to be processed. The first auxiliary sound pickup device is located at the azimuth indicated by the second azimuth information, and the second auxiliary sound pickup device is located at the azimuth indicated by the third azimuth information. The first auxiliary sound pickup device and the second auxiliary sound pickup device are not in the same plane as the left ear sound pickup device and the right ear sound pickup device;

[0033] Obtain a first energy, a second energy, a third energy and a fourth energy. The first energy is the energy of the audio signal corresponding to the second azimuth information in the first audio signal and the second audio signal, the second energy is the energy of the third audio signal, the third energy is the energy of the audio signal corresponding to the third azimuth information in the first audio signal and the second audio signal, and the fourth energy is the energy of the fourth audio signal;

[0034] Determine a fourth azimuth information according to the first energy, the second energy, the first preset azimuth information and the second preset azimuth information;

[0035] Determine a fifth azimuth information according to the third energy, the fourth energy, the third preset azimuth information and the fourth preset azimuth information;

[0036] Generate the first azimuth information according to the fourth azimuth information and the fifth azimuth information.

[0037] In this way, in any application scenario, the azimuth information of the sound source can be jointly determined according to the interaural time difference and the interaural intensity difference, with better applicability and more accurate determined azimuth information. As a result, the subsequent generated recording signal can more accurately express the spatial information of the sound signal, and after playback, the user's listening experience is better.

[0038] In a possible implementation manner, the rendering of the first direct sound signal and the second direct sound signal according to the first azimuth information to generate the recording signal of the to-be-processed sound signal includes:

[0039] Obtain a first response information, where the first response information is the preset binaural room impulse response information corresponding to the first azimuth information;

[0040] Render the first direct sound signal and the second direct sound signal according to the first response information to generate the recording signal of the to-be-processed sound signal.

[0041] In this way, the spatial sense of the recording signal can be enhanced, making the recording signal closer to the human ear's listening feeling and providing a better user experience.

[0042] In a second aspect, an embodiment of the present application provides a generating device for a recording signal, and the device includes:

[0043] A first obtaining module, configured to obtain a first audio signal and a second audio signal, where the first audio signal is an audio signal obtained after a left ear pickup device picks up the to-be-processed sound signal, and the second audio signal is an audio signal obtained after a right ear pickup device picks up the to-be-processed sound signal;

[0044] A second obtaining module, configured to obtain a first direct sound signal of the first audio signal and a second direct sound signal of the second audio signal;

[0045] A determining module, configured to determine a first azimuth information of a sound source that emits the to-be-processed sound signal according to the first direct sound signal and the second direct sound signal;

[0046] A generating module, configured to render the first direct sound signal and the second direct sound signal according to the first azimuth information to generate the recording signal of the to-be-processed sound signal.

[0047] In a possible implementation manner, the determining module is configured to determine a first azimuth information of a sound source that emits the to-be-processed sound signal according to the first direct sound signal and the second direct sound signal, specifically:

[0048] The determining module is configured to:

[0049] Generate binaural difference information based on the first direct sound signal and the second direct sound signal, where the binaural difference information includes a first interaural time difference and / or a first interaural intensity difference;

[0050] Determine the first azimuth information according to the binaural difference information.

[0051] In a possible implementation, the binaural difference information is the first interaural time difference, and the determining module is configured to determine the first azimuth information according to the binaural difference information, specifically:

[0052] The determining module is configured to:

[0053] Obtain a first preset azimuth information and a second preset azimuth information corresponding to the first interaural time difference;

[0054] Determine a second azimuth information, where the azimuth indicated by the second azimuth information is the azimuth jointly indicated by the first preset azimuth information and the second preset azimuth information;

[0055] Obtain a third audio signal, where the third audio signal is an audio signal obtained by a first auxiliary sound pickup device picking up the sound signal to be processed, and the first auxiliary sound pickup device is located at the azimuth indicated by the second azimuth information and is not in the same plane as the left ear sound pickup device and the right ear sound pickup device;

[0056] Obtain a first energy and a second energy, where the first energy is the energy of the audio signal corresponding to the second azimuth information in the first audio signal and the second audio signal, and the second energy is the energy of the third audio signal;

[0057] Determine the first azimuth information according to the first energy, the second energy, the first preset azimuth information, and the second preset azimuth information.

[0058] In a possible implementation, the binaural difference information is the first interaural intensity difference, and the determining module is configured to determine the first azimuth information according to the binaural difference information, specifically:

[0059] The determining module is configured to:

[0060] Obtain a third preset azimuth information and a fourth preset azimuth information corresponding to the first interaural intensity difference;

[0061] Determine a third azimuth information, where the azimuth indicated by the third azimuth information is the azimuth jointly indicated by the third preset azimuth information and the fourth preset azimuth information;

[0062] Obtain a fourth audio signal, where the fourth audio signal is an audio signal obtained after the second auxiliary sound pickup device picks up the sound signal to be processed, and the second auxiliary sound pickup device is located at the azimuth indicated by the third azimuth information and is not in the same plane as the left ear sound pickup device and the right ear sound pickup device;

[0063] Obtain a third energy and a fourth energy, where the third energy is the energy of the audio signal corresponding to the third azimuth information in the first audio signal and the second audio signal, and the fourth energy is the energy of the fourth audio signal;

[0064] Determine the first azimuth information according to the third energy, the fourth energy, the third preset azimuth information, and the fourth preset azimuth information.

[0065] In a possible implementation, the binaural difference information is the first interaural time difference and the first interaural intensity difference, and the determining module is configured to determine the first azimuth information according to the binaural difference information, specifically:

[0066] The determining module is configured to:

[0067] Obtain the first preset azimuth information and the second preset azimuth information corresponding to the first interaural time difference, and the third preset azimuth information and the fourth preset azimuth information corresponding to the first interaural intensity difference;

[0068] Determine a second azimuth information and a third azimuth information, where the azimuth indicated by the second azimuth information is the azimuth jointly indicated by the first preset azimuth information and the second preset azimuth information, and the azimuth indicated by the third azimuth information is the azimuth jointly indicated by the third preset azimuth information and the fourth preset azimuth information;

[0069] Obtain a third audio signal and a fourth audio signal, where the third audio signal is an audio signal obtained after the first auxiliary sound pickup device picks up the sound signal to be processed, and the fourth audio signal is an audio signal obtained after the second auxiliary sound pickup device picks up the sound signal to be processed. The first auxiliary sound pickup device is located at the azimuth indicated by the second azimuth information, and the second auxiliary sound pickup device is located at the azimuth indicated by the third azimuth information. Both the first auxiliary sound pickup device and the second auxiliary sound pickup device are not in the same plane as the left ear sound pickup device and the right ear sound pickup device;

[0070] Obtain a first energy, a second energy, a third energy, and a fourth energy, where the first energy is the energy of the audio signal corresponding to the second azimuth information in the first audio signal and the second audio signal, the second energy is the energy of the third audio signal, the third energy is the energy of the audio signal corresponding to the third azimuth information in the first audio signal and the second audio signal, and the fourth energy is the energy of the fourth audio signal;

[0071] Determine a fourth azimuth information according to the first energy, the second energy, the first preset azimuth information, and the second preset azimuth information;

[0072] Determine a fifth azimuth information according to the third energy, the fourth energy, the third preset azimuth information, and the fourth preset azimuth information;

[0073] Generate the first azimuth information according to the fourth azimuth information and the fifth azimuth information.

[0074] In a possible implementation manner, the generating module is configured to render the first direct sound signal and the second direct sound signal according to the first azimuth information to generate a recording signal of the to-be-processed sound signal, specifically:

[0075] The generating module is configured to:

[0076] Obtain a first response information, where the first response information is the preset binaural room impulse response information corresponding to the first azimuth information;

[0077] Render the first direct sound signal and the second direct sound signal according to the first response information to generate a recording signal of the to-be-processed sound signal.

[0078] In a third aspect, an embodiment of the present application provides an electronic device, which includes: one or more processors and one or more memories; the one or more memories store computer programs or instructions, and when the computer programs or instructions are executed by the one or more processors, the electronic device executes the method described in any item of the first aspect.

[0079] In a fourth aspect, an embodiment of the present application provides a recording signal generating system, which includes: a left ear sound pickup device and a right ear sound pickup device for picking up the to-be-processed sound signal, and the electronic device described in the third aspect; a first auxiliary sound pickup device and a second auxiliary sound pickup device for picking up the to-be-processed sound signal are arranged in the electronic device; the left ear sound pickup device and the right ear sound pickup device are communicatively connected to the electronic device through a wireless or wired manner.

[0080] Fifth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program or instruction is stored. When the computer program or instruction is executed, the method according to any one of the first aspect is executed.

[0081] It can be seen that through the method, device, system and electronic device for generating a recording signal provided by the present application, the electronic device can respectively perform direct sound extraction on the first audio signal obtained by the left ear pickup device picking up the sound signal to be processed and the second audio signal obtained by the right ear pickup device picking up the same sound signal to be processed, to obtain a first direct sound signal and a second direct sound signal. Then, the azimuth information of the sound source emitting the sound signal can be determined according to the first direct sound signal and the second direct sound signal. After that, the first direct sound signal and the second direct sound signal can be rendered according to the azimuth information to obtain the recording signal of the sound signal to be processed. It can be seen that the recording signal obtained by this method fully integrates the azimuth information of the sound source, can well express the spatial information of the sound signal, and after playback, can achieve the effect of human ear hearing, and the user experience is better. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] Figure 1 is a schematic diagram of an application scenario provided by an embodiment of the present application;

[0083] Figure 2 is a schematic structural diagram of an electronic device provided by an embodiment of the present application;

[0084] Figure 3 is a software structure block diagram of an electronic device provided by an embodiment of the present application;

[0085] Figure 4 is a schematic flowchart of a method for generating a recording signal provided by an embodiment of the present application;

[0086] Figure 5 is a schematic flowchart of a method for obtaining the first direct sound signal of the first audio signal and the second direct sound signal of the second audio signal provided by an embodiment of the present application;

[0087] Figure 6 is a schematic flowchart of a method for determining the first azimuth information provided by an embodiment of the present application;

[0088] Figure 7 is another schematic diagram of an application scenario provided by an embodiment of the present application;

[0089] Figure 8 is a schematic flowchart of another method for determining the first azimuth information provided by an embodiment of the present application;

[0090] Figure 9It is a schematic flowchart of another method for determining the first orientation information provided by an embodiment of the present application;

[0091] Figure 10 It is a structural block diagram of a recording signal generation device provided by an embodiment of the present application;

[0092] Figure 11 It is a structural block diagram of a chip provided by an embodiment of the present application. Detailed implementation manners

[0093] Next, the technical solutions of the present application will be described with reference to the accompanying drawings.

[0094] In the description of the present application, unless otherwise specified, "and / or" is merely an association relationship describing associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, "at least one" means one or more, "at least two" means two or more, and "multiple" also means two or more. The terms "first", "second", etc. do not limit the quantity and execution order, and the terms "first", "second", etc. do not necessarily limit to be different.

[0095] It should be noted that in the present application, words such as "exemplarily" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplarily" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplarily" or "for example" aims to present relevant concepts in a specific manner.

[0096] To facilitate the understanding of the technical solutions of the present application, the application scenarios of the technical solutions provided by the present application will be exemplarily described below.

[0097] See Figure 1 , Figure 1 which is a schematic diagram of an application scenario provided by the present application. The technical solutions of the present application can be applied to the application scenario shown in Figure 1 . As shown in Figure 1 , this application scenario may include: an electronic device 100 and a stereo headset 200.

[0098] Among them, the stereo headset 200 may include a left ear headset 201 and a right ear headset 202. The left ear headset 201 can be used to be worn on the user's left ear. A left ear sound pickup device (such as a microphone) may be provided in the left ear headset 201. The left ear headset 201 can pick up sound signals in the surrounding environment through the left ear sound pickup device therein to generate a left ear audio signal. The right ear headset 202 can be used to be worn on the user's right ear. A right ear sound pickup device may be provided in the right ear headset 202. The right ear headset 202 can pick up sound signals in the surrounding environment through the right ear sound pickup device therein to generate a right ear audio signal.

[0099] The electronic device 100 can be communicatively connected to the left ear headset 201 and / or the right ear headset 202 of the stereo headset 200 in a wired connection manner or a wireless connection manner. After the electronic device 100 is communicatively connected to the left ear headset 201, it can obtain the left ear audio signal from the left ear headset 201. Correspondingly, after the electronic device 100 is communicatively connected to the right ear headset 202, it can obtain the right ear audio signal from the right ear headset 202.

[0100] One or more auxiliary sound pickup devices may also be provided in the electronic device 100. For example, a microphone (abbreviated as top microphone) provided on the top of the electronic device 100, a microphone (abbreviated as bottom microphone) provided on the bottom of the electronic device 100, and a microphone (abbreviated as back microphone) provided on the back of the electronic device 100, etc. The electronic device 100 can also pick up sound signals in the surrounding environment through the auxiliary sound pickup devices provided therein to generate corresponding audio signals.

[0101] In addition, for the specific functions or roles of the electronic device 100 and the stereo headset 200, reference may also be made to the content of the subsequent embodiments, which will not be elaborated here.

[0102] Currently, after the user wears the stereo headset 200 on both ears, that is, wears the left ear headset 201 on the left ear and at the same time wears the right ear headset 202 on the right ear, the left ear headset 201 and the right ear headset 202 of the stereo headset 200 can pick up sound signals in the surrounding environment at the same time to respectively generate a left ear audio signal and a right ear audio signal. During this process, the electronic device 100 can establish a communication connection with the left ear headset 201 and the right ear headset 202 of the stereo headset 200, obtain the left ear audio signal and the right ear audio signal from the stereo headset 200 in real time, and store the obtained left ear audio signal and right ear audio signal generated by the same sound signal as a recording signal correspondingly. Subsequently, the recording signal can be replayed through the speaker in the electronic device 100.

[0103] However, due to the limited spatial perception that can be presented by the hardware of the stereo headset 200, the generated recording signal cannot well express the spatial information of the sound signal, failing to achieve the effect of human ear hearing. Subsequently, during playback, the restoration of the sense of space cannot be realized, resulting in a poor user experience.

[0104] To solve the above problems, the embodiments of the present application provide a method, device, system, and electronic device for generating a recording signal. In this method, the electronic device can separately perform direct sound extraction on the first audio signal obtained by the left ear pickup device picking up the sound signal to be processed and the second audio signal obtained by the right ear pickup device picking up the same sound signal to be processed, to obtain a first direct sound signal and a second direct sound signal. Then, the azimuth information of the sound source that emits the sound signal can be determined based on the first direct sound signal and the second direct sound signal. After that, the first direct sound signal and the second direct sound signal can be rendered according to the azimuth information to obtain the recording signal of the sound signal to be processed. The recording signal obtained by this method fully integrates the azimuth information of the sound source, can well express the spatial information of the sound signal, and after playback, can achieve the effect of human ear hearing, providing a better user experience.

[0105] It should be noted that the electronic device of the present application can be stationary or mobile. The electronic device can include devices at the outermost periphery of a computer network such as a communication terminal, a vehicle-mounted device, a mobile device, a user terminal, a mobile terminal, a wireless communication device, a portable terminal, a user agent, a user device, a service device, or a user equipment (UE), etc., mainly used for data input and output or display of processing results, etc. For example, the terminal device can be a mobile phone, a tablet device, a handheld device with wireless communication function, a computing device, a vehicle-mounted communication module, or other processing devices connected to a wireless modem, etc.

[0106] Exemplarily, reference can be made to Figure 2 , Figure 2 which shows a schematic structural diagram of the electronic device 100.

[0107] The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0108] It can be understood that the structure schematically shown in the embodiments of the present invention does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0109] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0110] The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.

[0111] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can hold the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can be directly called from the said memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0112] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0113] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 may include multiple groups of I2C buses. The processor 110 can be respectively coupled to the touch sensor 180K, the charger, the flash, the camera 193, etc. through different I2C bus interfaces. For example: The processor 110 can be coupled to the touch sensor 180K through the I2C interface, enabling the processor 110 and the touch sensor 180K to communicate through the I2C bus interface to implement the touch function of the electronic device 100.

[0114] The I2S interface can be used for audio communication. In some embodiments, the processor 110 may include multiple groups of I2S buses. The processor 110 can be coupled to the audio module 170 through the I2S bus to implement communication between the processor 110 and the audio module 170. In some embodiments, the audio module 170 can transmit an audio signal to the wireless communication module 160 through the I2S interface to implement the function of answering a call through a Bluetooth headset.

[0115] The PCM interface can also be used for audio communication to sample, quantize, and encode analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled through the PCM bus interface. In some embodiments, the audio module 170 can also transmit audio signals to the wireless communication module 160 through the PCM interface to implement the function of answering a call through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.

[0116] The UART interface is a general-purpose serial data bus for asynchronous communication. This bus can be a two-way communication bus. It converts the data to be transmitted between serial communication and parallel communication. In some embodiments, the UART interface is generally used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 through the UART interface to implement the Bluetooth function. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 through the UART interface to implement the function of playing music through a Bluetooth headset.

[0117] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display screen 194 and the camera 193. The MIPI interface includes a camera serial interface (CSI), a display serial interface (DSI), etc. In some embodiments, the processor 110 and the camera 193 communicate through the CSI interface to implement the shooting function of the electronic device 100. The processor 110 and the display screen 194 communicate through the DSI interface to implement the display function of the electronic device 100.

[0118] The GPIO interface can be configured by software. The GPIO interface can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to the camera 193, the display screen 194, the wireless communication module 160, the audio module 170, the sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.

[0119] The USB interface 130 is an interface that complies with the USB standard specification, and can specifically be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 130 can be used to connect a charger to charge the electronic device 100, and can also be used for data transmission between the electronic device 100 and peripheral devices. It can also be used to connect headphones to play audio. This interface can also be used to connect other electronic devices, such as AR devices, etc.

[0120] It can be understood that the interface connection relationships among the modules illustrated in the embodiments of the present invention are only illustrative descriptions and do not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection manners in the above embodiments, or a combination of multiple interface connection manners.

[0121] The charging management module 140 is configured to receive a charging input from a charger. The charger may be a wireless charger or a wired charger. In some embodiments of wired charging, the charging management module 140 may receive the charging input from the wired charger through the USB interface 130. In some embodiments of wireless charging, the charging management module 140 may receive the wireless charging input through the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 may also supply power to the electronic device through the power management module 141.

[0122] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives the inputs from the battery 142 and / or the charging management module 140 and supplies power to the processor 110, the internal memory 121, the display screen 194, the camera 193, the wireless communication module 160, etc. The power management module 141 may also be used to monitor parameters such as the battery capacity, the number of battery charge cycles, and the battery health status (leakage, impedance). In some other embodiments, the power management module 141 may also be disposed in the processor 110. In some other embodiments, the power management module 141 and the charging management module 140 may also be disposed in the same device.

[0123] The wireless communication function of the electronic device 100 may be implemented by the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modulation and demodulation processor, and the baseband processor, etc.

[0124] The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 100 may be used to cover a single or multiple communication frequency bands. Different antennas may also be multiplexed to improve the utilization rate of the antennas. For example, the antenna 1 may be multiplexed as the diversity antenna of the wireless local area network. In some other embodiments, the antenna may be used in combination with a tuning switch.

[0125] The mobile communication module 150 may provide solutions for wireless communications such as 2G / 3G / 4G / 5G applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 may receive electromagnetic waves through the antenna 1, filter, amplify, and process the received electromagnetic waves, and then transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 may also amplify the signal modulated by the modulation and demodulation processor and convert it into electromagnetic waves through the antenna 1 for radiation. In some embodiments, at least some functional modules of the mobile communication module 150 may be provided in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be provided in the same device.

[0126] The modulation and demodulation processor may include a modulator and a demodulator. Among them, the modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. Subsequently, the demodulator transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, receiver 170B, etc.), or displays an image or video through the display screen 194. In some embodiments, the modulation and demodulation processor may be an independent device. In other embodiments, the modulation and demodulation processor may be independent of the processor 110 and provided in the same device as the mobile communication module 150 or other functional modules.

[0127] The wireless communication module 160 may provide solutions for wireless communications applied to the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite systems (GNSSs), frequency modulation (FM), near field communication (NFC), infrared (IR), etc. The wireless communication module 160 may be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, frequency-modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 may also receive signals to be sent from the processor 110, frequency-modulate them, amplify them, and convert them into electromagnetic waves through the antenna 2 for radiation.

[0128] In some embodiments, the antenna 1 of the electronic device 100 is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, so that the electronic device 100 can communicate with the network and other devices through wireless communication technologies. The wireless communication technologies may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology, etc. The GNSS may include global positioning system (GPS), global navigation satellite system (GLONASS), beidou navigation satellite system (BDS), quasi-zenith satellite system (QZSS), and / or satellite based augmentation systems (SBAS).

[0129] The electronic device 100 implements the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, and is connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.

[0130] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or N display screens 194, where N is a positive integer greater than 1.

[0131] The electronic device 100 can implement the shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, an application processor, etc.

[0132] The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, and light passes through the lens and is transmitted to the camera photosensitive element. The optical signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also optimize the noise, brightness, and skin color of the image through algorithms. The ISP can also optimize parameters such as the exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.

[0133] The camera 193 is used to capture static images or videos. An object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in standard RGB, YUV, etc. formats. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0134] The digital signal processor is used to process digital signals. In addition to being able to process digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.

[0135] The video codec is used to compress or decompress digital videos. The electronic device 100 can support one or more video codecs. In this way, the electronic device 100 can play or record videos in multiple coding formats, such as: Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0136] The NPU is a neural-network (NN) computing processor. By learning from the structure of biological neural networks, such as learning from the transmission mode between human brain neurons, it can quickly process input information and can also continuously self-learn. Through the NPU, applications such as intelligent cognition of the electronic device 100 can be realized, such as: image recognition, face recognition, speech recognition, text understanding, etc.

[0137] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.

[0138] The internal memory 121 can be used to store computer-executable program code, and the executable program code includes instructions. The internal memory 121 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.). The data storage area can store data created during the use of the electronic device 100 (such as audio data, phone book, etc.). In addition, the internal memory 121 can include high-speed random access memory and can also include non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121 and / or the instructions stored in the memory provided in the processor.

[0139] The electronic device 100 can implement audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone interface 170D, and the application processor, etc. For example, music playback, recording, etc.

[0140] The audio module 170 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 can be disposed in the processor 110, or some functional modules of the audio module 170 can be disposed in the processor 110.

[0141] The speaker 170A, also known as a "loudspeaker", is used to convert an audio electrical signal into a sound signal. The electronic device 100 can listen to music or hands-free calls through the speaker 170A. Multiple speakers 170A can be provided in the electronic device 100. For example, one speaker 170A can be provided at the top of the electronic device 100, and another speaker 170A can be provided at the bottom, etc.

[0142] The receiver 170B, also known as an "earpiece", is used to convert an audio electrical signal into a sound signal. When the electronic device 100 answers a call or a voice message, the voice can be listened to by placing the receiver 170B close to the human ear. In some embodiments, the speaker 170A and the receiver 170B can also be set as one component, and the present application does not limit this.

[0143] The microphone 170C, also known as a "microphone" or "transmitter", is used to convert a sound signal into an electrical signal. When making a call or sending a voice message, the user can speak close to the microphone 170C with the mouth to input the sound signal into the microphone 170C. At least one microphone 170C can be provided in the electronic device 100. In some other embodiments, two microphones 170C can be provided in the electronic device 100, which can not only collect sound signals but also implement a noise reduction function. In some other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C to collect sound signals, reduce noise, identify the sound source, and implement a directional recording function, etc.

[0144] The headphone jack 170D is used to connect a wired headphone. The headphone jack 170D can be a USB interface 130, or a 3.5 mm open mobile terminal platform (OMTP) standard interface, or a Cellular Telecommunications Industry Association of the USA (CTIA) standard interface.

[0145] The pressure sensor 180A is used to sense pressure signals and can convert pressure signals into electrical signals. In some embodiments, the pressure sensor 180A may be disposed on the display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. The capacitive pressure sensor may include at least two parallel plates having conductive materials. When a force acts on the pressure sensor 180A, the capacitance between the electrodes changes. The electronic device 100 determines the intensity of the pressure based on the change in capacitance. When a touch operation acts on the display screen 194, the electronic device 100 detects the intensity of the touch operation according to the pressure sensor 180A. The electronic device 100 can also calculate the position of the touch based on the detection signal of the pressure sensor 180A. In some embodiments, touch operations acting on the same touch position but with different touch operation intensities may correspond to different operation instructions. For example: When a touch operation with a touch operation intensity less than the first pressure threshold acts on the short message application icon, the instruction to view the short message is executed. When a touch operation with a touch operation intensity greater than or equal to the first pressure threshold acts on the short message application icon, the instruction to create a new short message is executed.

[0146] The gyroscope sensor 180B can be used to determine the motion posture of the electronic device 100. In some embodiments, the angular velocity of the electronic device 100 around three axes (i.e., the x, y, and z axes) can be determined by the gyroscope sensor 180B. The gyroscope sensor 180B can be used for anti-shake shooting. Exemplarily, when the shutter is pressed, the gyroscope sensor 180B detects the angle of jitter of the electronic device 100, calculates the distance that the lens module needs to compensate based on the angle, and enables the lens to offset the jitter of the electronic device 100 through reverse movement to achieve anti-shake. The gyroscope sensor 180B can also be used for navigation and somatosensory game scenarios.

[0147] The barometric pressure sensor 180C is used to measure barometric pressure. In some embodiments, the electronic device 100 calculates the altitude based on the barometric pressure value measured by the barometric pressure sensor 180C to assist in positioning and navigation.

[0148] The magnetic sensor 180D includes a Hall sensor. The electronic device 100 can use the magnetic sensor 180D to detect the opening and closing of the flip leather case. In some embodiments, when the electronic device 100 is a flip phone, the electronic device 100 can detect the opening and closing of the flip according to the magnetic sensor 180D. Furthermore, according to the detected opening and closing state of the leather case or the opening and closing state of the flip, features such as automatic flip unlocking are set.

[0149] The acceleration sensor 180E can detect the magnitude of the acceleration of the electronic device 100 in various directions (generally three axes). When the electronic device 100 is stationary, the magnitude and direction of gravity can be detected. It can also be used to identify the posture of the electronic device and is applied to applications such as horizontal and vertical screen switching and pedometers.

[0150] A distance sensor 180F is used to measure distance. The electronic device 100 can measure distance through infrared or laser. In some embodiments, when shooting a scene, the electronic device 100 can use the distance sensor 180F to measure distance to achieve fast focusing.

[0151] The proximity light sensor 180G may include, for example, a light-emitting diode (LED) and a light detector, such as a photodiode. The light-emitting diode may be an infrared light-emitting diode. The electronic device 100 emits infrared light outward through the light-emitting diode. The electronic device 100 uses the photodiode to detect the infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the electronic device 100. When insufficient reflected light is detected, the electronic device 100 can determine that there is no object near the electronic device 100. The electronic device 100 can use the proximity light sensor 180G to detect when the user holds the electronic device 100 close to the ear for a call, so as to automatically turn off the screen to save power. The proximity light sensor 180G can also be used for automatic unlocking and locking of the holster mode and pocket mode.

[0152] The ambient light sensor 180L is used to sense the ambient light brightness. The electronic device 100 can adaptively adjust the brightness of the display screen 194 according to the sensed ambient light brightness. The ambient light sensor 180L can also be used to automatically adjust the white balance when taking pictures. The ambient light sensor 180L can also cooperate with the proximity light sensor 180G to detect whether the electronic device 100 is in the pocket to prevent accidental touch.

[0153] The fingerprint sensor 180H is used to collect fingerprints. The electronic device 100 can use the collected fingerprint characteristics to achieve fingerprint unlocking, access application locks, fingerprint photography, fingerprint answering of incoming calls, etc.

[0154] The temperature sensor 180J is used to detect temperature. In some embodiments, the electronic device 100 uses the temperature detected by the temperature sensor 180J to execute a temperature processing strategy. For example, when the temperature reported by the temperature sensor 180J exceeds a threshold, the electronic device 100 reduces the performance of the processor near the temperature sensor 180J to reduce power consumption and implement thermal protection. In other embodiments, when the temperature is lower than another threshold, the electronic device 100 heats the battery 142 to avoid abnormal shutdown of the electronic device 100 caused by low temperature. In other embodiments, when the temperature is lower than yet another threshold, the electronic device 100 boosts the output voltage of the battery 142 to avoid abnormal shutdown caused by low temperature.

[0155] The touch sensor 180K, also known as the "touch control device". The touch sensor 180K can be disposed on the display screen 194. The touch sensor 180K and the display screen 194 together form a touch screen, also known as the "touch control screen". The touch sensor 180K is used to detect touch operations acting thereon or in its vicinity. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 194. In some other embodiments, the touch sensor 180K can also be disposed on the surface of the electronic device 100, at a different position from that of the display screen 194.

[0156] The bone conduction sensor 180M can acquire vibration signals. In some embodiments, the bone conduction sensor 180M can acquire the vibration signals of the vibrating bone mass of the human vocal part. The bone conduction sensor 180M can also contact the human pulse to receive blood pressure pulsation signals. In some embodiments, the bone conduction sensor 180M can also be disposed in the earphone to form a bone conduction earphone. The audio module 170 can parse out voice signals based on the vibration signals of the vibrating bone mass of the human vocal part acquired by the bone conduction sensor 180M to implement the voice function. The application processor can parse out heart rate information based on the blood pressure pulsation signals acquired by the bone conduction sensor 180M to implement the heart rate detection function.

[0157] The button 190 includes a power-on button, a volume button, etc. The button 190 can be a mechanical button or a touch button. The electronic device 100 can receive button inputs to generate key signal inputs related to the user settings and function control of the electronic device 100.

[0158] The motor 191 can generate vibration prompts. The motor 191 can be used for incoming call vibration prompts and can also be used for touch vibration feedback. For example, touch operations acting on different applications (such as taking pictures, playing audio, etc.) can correspond to different vibration feedback effects. Touch operations acting on different areas of the display screen 194 can also correspond to different vibration feedback effects for the motor 191. Different application scenarios (such as time reminder, receiving information, alarm clock, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.

[0159] The indicator 192 can be an indicator light and can be used to indicate the charging state, power change, and can also be used to indicate messages, missed calls, notifications, etc.

[0160] The SIM card interface 195 is used to connect to the SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to achieve contact and separation from the electronic device 100. The electronic device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 195 simultaneously. The types of the multiple cards can be the same or different. The SIM card interface 195 can also be compatible with different types of SIM cards. The SIM card interface 195 can also be compatible with external memory cards. The electronic device 100 interacts with the network through the SIM card to implement functions such as calls and data communication. In some embodiments, the electronic device 100 uses an eSIM, that is, an embedded SIM card. The eSIM card can be embedded in the electronic device 100 and cannot be separated from the electronic device 100.

[0161] The software system of the electronic device 100 can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. In the embodiments of the present invention, the Android system with a layered architecture is taken as an example to exemplarily illustrate the software structure of the electronic device 100.

[0162] Figure 3 It is a software structure block diagram of the electronic device 100 in the embodiments of the present application.

[0163] The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom, namely the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.

[0164] The application layer can include a series of application packages.

[0165] As Figure 3 shown, the application packages can include applications such as a camera, a gallery, a calendar, a call, a map, a navigation, a WLAN, a Bluetooth, music, a video, a short message, etc.

[0166] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions.

[0167] As Figure 3 shown, the application framework layer can include a window manager, a content provider, a view system, a telephone manager, a resource manager, a notification manager, etc.

[0168] The window manager is used to manage window programs. The window manager can obtain the display screen size, determine whether there is a status bar, lock the screen, capture the screen, etc.

[0169] The content provider is used to store and obtain data, and make this data accessible to application programs. The data may include videos, images, audio, incoming and outgoing calls, browsing history and bookmarks, phone books, etc.

[0170] The view system includes visual controls, such as controls for displaying text, controls for displaying pictures, etc. The view system can be used to build application programs. The display interface can be composed of one or more views. For example, a display interface including a text message notification icon may include a view for displaying text and a view for displaying pictures.

[0171] The phone manager is used to provide the communication function of the electronic device 100. For example, the management of call states (including connection, disconnection, etc.).

[0172] The resource manager provides various resources for application programs, such as localized strings, icons, pictures, layout files, video files, etc.

[0173] The notification manager enables application programs to display notification information in the status bar. It can be used to convey message types of notifications, which can automatically disappear after a short stay without user interaction. For example, the notification manager is used to inform that the download is completed, message reminders, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or scroll bar text, such as a notification of a background running application program, and can also be a notification that appears on the screen in the form of a dialogue window. For example, prompt text information in the status bar, emit a prompt tone, the electronic device vibrates, the indicator light flashes, etc.

[0174] AndroidRuntime includes a core library and a virtual machine. Androidruntime is responsible for the scheduling and management of the Android system.

[0175] The core library contains two parts: one part is the functional functions that need to be called by the Java language, and the other part is the core library of Android.

[0176] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as the management of object life cycles, stack management, thread management, security and exception management, and garbage collection.

[0177] The system library may include multiple functional modules. For example: surface manager, Media Libraries, 3D graphics processing library (e.g., OpenGL ES), 2D graphics engine (e.g., SGL), etc.

[0178] The surface manager is used to manage the display subsystem and provides the fusion of 2D and 3D layers for multiple applications.

[0179] The media library supports the playback and recording of various common audio and video formats, as well as static image files, etc. The media library can support multiple audio and video coding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.

[0180] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing, etc.

[0181] The 2D graphics engine is a drawing engine for 2D drawing.

[0182] The kernel layer is the layer between hardware and software. The kernel layer at least includes a display driver, a camera driver, an audio driver, and a sensor driver.

[0183] Next, an embodiment of the method for generating a recording signal provided in this application will be described.

[0184] See Figure 4 , Figure 4 is a schematic flowchart of a method for generating a recording signal provided in an embodiment of this application. This method can be applied to an electronic device, for example, it can be applied to Figures 1 to 3 the electronic device 100 shown. As Figure 4 shown, this method may include the following steps:

[0185] Step S101, obtain a first audio signal and a second audio signal.

[0186] Combined with the foregoing content, after the user wears stereo headphones (such as TWS headphones) on both ears, the left-ear pickup device and the right-ear pickup device of the stereo headphones worn on both ears can simultaneously pick up the same sound signal, and then send the left-ear audio signal and the right-ear audio signal obtained after pickup to the electronic device for storage by the electronic device to achieve the purpose of recording, thereby obtaining a recording signal. That is to say, the electronic device can obtain the left-ear audio signal obtained by the left-ear pickup device picking up the sound signal and the right-ear audio signal obtained by the right-ear pickup device picking up the same sound signal through the left-ear pickup device and the right-ear pickup device communicatively connected thereto.

[0187] For example, in an application scenario where a user uses TWS earphones and a mobile phone to shoot a video blog (VOLG), after the user wears the TWS earphones on both ears, the TWS earphones can be communicatively connected to the mobile phone via Bluetooth or WIFI. Then, the user can turn on the binaural recording function in the VOLG shooting mode of the mobile phone. After that, the TWS earphones can pick up the same sound signal simultaneously through the microphone in the left ear and the microphone in the right ear, obtaining the left ear audio signal and the right ear audio signal corresponding to the same sound signal. During this process, the mobile phone can obtain the left ear audio signal and the right ear audio signal generated after the TWS earphones pick up the sound in real time, and the mobile phone can store the left ear audio signal and the right ear audio signal from the same sound signal as the recording signal correspondingly.

[0188] It should be noted that during the process of recording with stereo earphones worn on both ears by the user, both the left ear sound pickup device and the right ear sound pickup device can pick up the sound signal in real time and continuously. And the electronic device can also obtain the left ear audio signal and the right ear audio signal from the left ear sound pickup device and the right ear sound pickup device in real time and continuously. After the electronic device obtains the left ear audio signal and the right ear audio signal, it can perform subsequent processing on the obtained left ear audio signal and right ear audio signal generated by the same sound signal at a certain processing frequency to generate a recording signal. Since the implementation method of the electronic device for processing the left ear audio signal and the right ear audio signal generated by the same sound signal to generate a recording signal is the same each time, in the subsequent embodiments of this application, the left ear audio signal and the right ear audio signal generated by any same sound signal (subsequently, this sound signal will be denoted as the sound signal to be processed) are taken as an example to illustrate the embodiments of the method for generating a recording signal provided by this application. The processing method for generating a recording signal based on the left ear audio signal and the right ear audio signal obtained from other sound signals can refer to this processing method.

[0189] On this basis, in the subsequent embodiments of this application, the audio signal (left ear audio signal) obtained after the left ear sound pickup device picks up the sound signal to be processed is denoted as the first audio signal. The audio signal (right ear audio signal) obtained after the right ear sound pickup device picks up the sound signal to be processed is denoted as the second audio signal.

[0190] After the left ear sound pickup device and the right ear sound pickup device obtain the first audio signal and the second audio signal respectively, they can send the first audio signal and the second audio signal to the electronic device. That is to say, the electronic device can obtain the first audio signal and the second audio signal from the left ear sound pickup device and the right ear sound pickup device communicatively connected to it.

[0191] Step S102: Obtain the first direct sound signal of the first audio signal and the second direct sound signal of the second audio signal.

[0192] An audio signal usually includes a direct sound signal and an ambient sound signal. For the convenience of distinction, in subsequent embodiments of this application, the direct sound signal included in the first audio signal is denoted as the first direct sound signal, and the ambient sound signal included in the first audio signal is denoted as the first ambient sound signal. The direct sound signal included in the second audio signal is denoted as the second direct sound signal, and the ambient sound signal included in the second audio signal is denoted as the second ambient sound signal.

[0193] After the electronic device obtains the first audio signal through the left ear pickup device communicatively connected thereto and the second audio signal through the right ear pickup device communicatively connected thereto, the first direct sound signal can be extracted from the first audio signal, and the second direct sound signal can be extracted from the second audio signal.

[0194] In a possible implementation manner, for the implementation manner of extracting the first direct sound signal from the first audio signal and the second direct sound signal from the second audio signal, reference can be made to Figure 5 , Figure 5 which is a schematic flowchart of a method for obtaining the first direct sound signal of the first audio signal and the second direct sound signal of the second audio signal provided in an embodiment of this application. As Figure 5 shown, the method may include the following steps:

[0195] Step S201: Generate a first frequency domain signal and a second frequency domain signal.

[0196] Among them, the first frequency domain signal refers to the signal obtained after the first audio signal is converted to the frequency domain coordinate system. The second frequency domain signal refers to the signal obtained after the second audio signal is converted to the frequency domain coordinate system.

[0197] Optionally, the first frequency domain signal can be obtained by performing a Fourier transform on the first audio signal. Similarly, the second frequency domain signal can be obtained by performing a Fourier transform on the second audio signal.

[0198] Step S202: Generate a covariance matrix according to the first frequency domain signal and the second frequency domain signal.

[0199] In specific implementation, the sound signal to be processed may include multiple sound signals. For each sound signal included in the sound signal to be processed, there is a corresponding signal in the first audio signal and the second audio signal respectively. Similarly, for each sound signal included in the sound signal to be processed, there is also a corresponding signal in the first frequency domain signal and the second frequency domain signal respectively. Therefore, when performing step S202, a covariance matrix needs to be generated respectively for each sound signal included in the sound signal to be processed.

[0200] Optionally, when generating the covariance matrix corresponding to any one of the sound signals included in the sound signal to be processed, the sound signal may be determined as the target signal, and then the covariance matrix corresponding to the target signal may be generated in the following manner: Obtain the first signal X included in the first frequency-domain signal l and the second signal X included in the second frequency-domain signal r ; where the first signal X l is the signal corresponding to the target signal included in the first frequency-domain signal, and the second signal X r is the signal corresponding to the target signal included in the second frequency-domain signal; Obtain the matrix formed by the first signal X l and the second signal X r , denoted as the first matrix X H , X H = [X l , X r ; Obtain the conjugate transpose matrix X of the first matrix; Generate the covariance matrix C according to the first matrix X H and the conjugate transpose matrix X of the first matrix, C = XX H .

[0201] Step S203: Perform eigenvalue decomposition on the covariance matrix to generate a plurality of eigenvalues and eigenvectors corresponding to each eigenvalue.

[0202] When performing step S203, it is necessary to perform eigenvalue decomposition on each covariance matrix to generate a plurality of eigenvalues corresponding to the covariance matrix and eigenvectors corresponding to each eigenvalue.

[0203] Step S204: Obtain the normalized vector of the target eigenvector.

[0204] When performing step S204, it is necessary to obtain the normalized vector of the target eigenvector corresponding to each covariance matrix.

[0205] Among them, for any covariance matrix, the target eigenvector corresponding to it refers to the eigenvector corresponding to the target eigenvalue of the covariance matrix. The target eigenvalue corresponding to the covariance matrix refers to the largest eigenvalue among the plurality of eigenvalues generated corresponding to the covariance matrix.

[0206] Optionally, to obtain the normalized vector of any target eigenvector, it may be implemented in the following manner: Perform normalization processing on the target eigenvector ν1 to generate the normalized vector P

[0207] Step S205: Obtain the first direct sound signal and the second direct sound signal according to the normalized vector.

[0208] For each voice signal included in the voice signal to be processed, after determining the voice signal as the target signal, generating the covariance matrix corresponding to the target signal, and obtaining the normalized vector of the target eigenvector corresponding to the covariance matrix, the direct sound signal corresponding to the target signal can be obtained in the following manner: According to the normalized vector P and the first signal X l , according to the following formula ω l = X l P, calculate the first weight coefficient ω l ; According to the normalized vector P and the second signal X r , according to the following formula ω r = X r P, calculate the second weight coefficient ω r ; According to the first weight coefficient ω l and the normalized vector P, according to the following formula P l = ω l P, calculate the direct sound signal P l included in the first signal X l ; According to the second weight coefficient ω r and the normalized vector P, according to the following formula P r = ω r P, calculate the direct sound signal P r included in the second signal X r .

[0209] Obtain the direct sound signal P l corresponding to each voice signal included in the voice signal to be processed and the direct sound signal P r in the above manner. After that, the first direct sound signal and the second direct sound signal can be obtained. The first direct sound signal includes the direct sound signal P l corresponding to all voice signals included in the voice signal to be processed, and the second direct sound signal includes the direct sound signal P r corresponding to all voice signals included in the voice signal to be processed. Both the first direct sound signal and the second direct sound signal are signals in the frequency domain coordinate system.

[0210] Step S103: Determine the first azimuth information of the sound source that emits the voice signal to be processed according to the first direct sound signal and the second direct sound signal.

[0211] After obtaining the first direct sound signal of the first audio signal and the second direct sound signal of the second audio signal, the binaural difference information, simply referred to as binaural difference information, can be generated based on the first direct sound signal and the second direct sound signal. Exemplarily, the binaural difference information may include a first interaural time difference (ITD) and / or a first interaural intensity difference (IID). Then, based on the binaural difference information, the first azimuth information of the sound source emitting the sound signal to be processed can be determined.

[0212] Optionally, after obtaining the first direct sound signal and the second direct sound signal, the following formula can be used: ITD = argmax τ R lr (τ) = argmax τ (ifft(p l (f)p r (f))), to calculate the first interaural time difference ITD. Where p l (f) represents the signal of a certain frequency point included in the first direct sound signal. p r (f) represents the signal of a certain frequency point included in the second direct sound signal. The signal p l (f) and the signal p r (f) correspond to the same sound signal in the sound signal to be processed, and τ represents the time delay corresponding to the maximum cross-correlation between the first direct sound signal and the second direct sound signal.

[0213] Optionally, after obtaining the first direct sound signal and the second direct sound signal, the first direct sound signal can be first transformed to the time domain coordinate system through inverse Fourier transform to obtain the third direct sound signal in the time domain coordinate system, and the second direct sound signal can be transformed to the time domain coordinate system through inverse Fourier transform to obtain the fourth direct sound signal in the time domain coordinate system. Then, the following formula can also be used: ITD = argmax τ R lr (τ) = argmax τ ∫p l (t)p r (t + τ)dt, to calculate the first interaural time difference ITD. Where p l (t) represents the signal of a certain moment included in the third direct sound signal, p r (t + τ) represents the signal of a certain moment included in the fourth direct sound signal, and the meaning of τ is as above and will not be elaborated here. When the value of t is the same, p l (t) and p r (t + τ) correspond to the same sound signal in the sound signal to be processed.

[0214] Optionally, after obtaining the first direct sound signal and the second direct sound signal, the following formula can be used First, calculate the IID(f) corresponding to each sound signal in the sound signal to be processed, and then calculate the average value of the IID(f) corresponding to all the sound signals in the sound signal to be processed to obtain the first interaural intensity difference IID. Where p l (f) represents the signal of a certain frequency point included in the first direct sound signal. p r (f) represents the signal of a certain frequency point included in the second direct sound signal, and the signal p l (f) and the signal p r (f) correspond to the same sound signal in the sound signal to be processed.

[0215] After calculating the binaural difference information based on the first direct sound signal and the second direct sound signal, according to the binaural difference information, the first azimuth information of the sound source that emits the sound signal to be processed can be determined in multiple ways. For example:

[0216] In the first method, if the binaural difference information is the first interaural time difference, then according to the binaural difference information, to determine the first azimuth information, reference can be made to Figure 6 the implementation method shown, Figure 6 which is a schematic flowchart of a method for determining the first azimuth information provided by an embodiment of the present application. As Figure 6 shown, the method may include the following steps:

[0217] Step S301: Obtain the first preset azimuth information and the second preset azimuth information corresponding to the first interaural time difference.

[0218] In specific implementation, according to the head-related transfer function (HRTF), the corresponding relationship between the interaural time difference and the azimuth information can be established in advance.

[0219] The azimuth information determined according to the interaural time difference is usually located within a confusion cone as Figure 7 shown. Usually, one interaural time difference can correspond to two azimuth information, but these two azimuth information can only indicate the left or right azimuth, and cannot indicate the front or rear azimuth. That is to say, usually the left or right azimuth indicated by these two azimuth information is the same, but the front or rear azimuth indicated is different. Here, the front, rear, left, and right are divided with the front of the user's both eyes as the front.

[0220] For the convenience of distinction, in the embodiment of the present application, the interaural time difference in the corresponding relationship between the interaural time difference and the azimuth information established in advance is denoted as the second interaural time difference, and the azimuth information in the corresponding relationship is denoted as the fifth preset azimuth information.

[0221] Based on this, obtaining the first preset azimuth information and the second preset azimuth information corresponding to the first interaural time difference can be achieved in the following manner: find the second interaural time difference that is the same as the first interaural time difference among the pre-stored second interaural time differences, and denote it as the third interaural time difference; respectively determine the two fifth preset azimuth information corresponding to the third interaural time difference as the first preset azimuth information and the second preset azimuth information corresponding to the first interaural time difference. The left or right azimuth indicated by the first preset azimuth information and the second preset azimuth information is the same, but the front or rear azimuth indicated may be different.

[0222] Step S302: Determine the second azimuth information.

[0223] Among them, the azimuth indicated by the second azimuth information is the azimuth jointly indicated by the first preset azimuth information and the second preset azimuth information. For example, if the first preset azimuth information indicates 30 degrees to the left and the second preset azimuth information indicates 60 degrees to the left, then the azimuth jointly indicated by the first preset azimuth information and the second preset azimuth information is left. Then, the azimuth indicated by the second azimuth information is left.

[0224] Step S303: Obtain the third audio signal.

[0225] Among them, the third audio signal is the audio signal obtained after the first auxiliary sound pickup device picks up the sound signal to be processed. The first auxiliary sound pickup device is located at the azimuth indicated by the second azimuth information and is not in the same plane as the left ear sound pickup device and the right ear sound pickup device.

[0226] For example, in the application scenario where a user uses TWS earphones and a mobile phone for VOLG shooting, if the azimuth indicated by the second azimuth information is left, the top microphone of the mobile phone can be used as the first auxiliary sound pickup device, and the audio signal generated after the top microphone of the mobile phone picks up the sound signal to be processed is the third audio signal. Or, if the azimuth indicated by the second azimuth information is right, the bottom microphone of the mobile phone can be used as the first auxiliary sound pickup device, and the audio signal generated after the bottom microphone of the mobile phone picks up the sound signal to be processed is the third audio signal.

[0227] Step S304: Obtain the first energy and the second energy.

[0228] Among them, the first energy is the energy of the audio signal corresponding to the second azimuth information in the first audio signal and the second audio signal. If the azimuth indicated by the second azimuth information is left, the audio signal corresponding to the second azimuth information in the first audio signal and the second audio signal is the first audio signal, and the first energy is the energy of the first audio signal. Or, if the azimuth indicated by the second azimuth information is right, the audio signal corresponding to the second azimuth information in the first audio signal and the second audio signal is the second audio signal, and the first energy is the energy of the second audio signal. The second energy is the energy of the third audio signal.

[0229] Step S305: Determine the first azimuth information according to the first energy, the second energy, the first preset azimuth information, and the second preset azimuth information.

[0230] In specific implementation, if the first energy is greater than the second energy, it means that the sound source emitting the sound signal to be processed is close to the user. Then, determine the preset azimuth information that is closer to the user among the first preset azimuth information and the second preset azimuth information as the first azimuth information. That is, if the azimuth indicated by the first preset azimuth information is closer to the user than the azimuth indicated by the second preset azimuth information, determine the first preset azimuth information as the first azimuth information. Or, if the azimuth indicated by the second preset azimuth information is closer to the user than the azimuth indicated by the first preset azimuth information, determine the second preset azimuth information as the first azimuth information. Or,

[0231] If the first energy is less than the second energy, it means that the sound source emitting the sound signal to be processed is far from the user and close to the first auxiliary sound pickup device. Then, determine the preset azimuth information that is farther from the user among the first preset azimuth information and the second preset azimuth information as the first azimuth information. That is, if the azimuth indicated by the first preset azimuth information is farther from the user than the azimuth indicated by the second preset azimuth information, determine the first preset azimuth information as the first azimuth information. Or, if the azimuth indicated by the second preset azimuth information is farther from the user than the azimuth indicated by the first preset azimuth information, determine the second preset azimuth information as the first azimuth information.

[0232] In the second method, the binaural difference information is the first interaural intensity difference. Then, according to the binaural difference information, determine the first azimuth information. The implementation method can be referred to Figure 8 the implementation method shown Figure 8 which is a schematic flowchart of another method for determining the first azimuth information provided by the embodiments of the present application. As Figure 8 shown, this method may include the following steps:

[0233] Step S401: Obtain the third preset azimuth information and the fourth preset azimuth information corresponding to the first interaural intensity difference.

[0234] In specific implementation, a corresponding relationship between the interaural intensity difference and the azimuth information can also be established in advance according to the head-related functions. And, the azimuth information determined according to the interaural intensity difference is usually also located within a confusion cone as shown in Figure 7 That is, within a confusion cone as shown in Figure 7 . Usually, one interaural intensity difference can correspond to two azimuth information, and these two azimuth information can only indicate the left or right azimuth, and cannot indicate the front or rear azimuth. That is to say, the left or right azimuths indicated by these two azimuth information are the same, but the front or rear azimuths indicated may be different.

[0235] For the convenience of distinction, in the embodiments of the present application, the interaural intensity difference in the pre-established corresponding relationship between the interaural intensity difference and the azimuth information is denoted as the second interaural intensity difference, and the azimuth information in this corresponding relationship is denoted as the sixth preset azimuth information.

[0236] Based on this, obtaining the third preset azimuth information and the fourth preset azimuth information corresponding to the first interaural intensity difference can be achieved in the following manner: find the second interaural intensity difference that is the same as the first interaural intensity difference among the pre-stored second interaural intensity differences, and denote it as the third interaural intensity difference; respectively determine the two sixth preset azimuth information corresponding to the third interaural intensity difference as the third preset azimuth information and the fourth preset azimuth information corresponding to the first interaural intensity difference. The left or right azimuths indicated by the third preset azimuth information and the fourth preset azimuth information are the same, but the front or rear azimuths indicated may be different.

[0237] Step S402, determine the third azimuth information.

[0238] Among them, the azimuth indicated by the third azimuth information is the azimuth jointly indicated by the third preset azimuth information and the fourth preset azimuth information. For example, if the azimuth indicated by the third preset azimuth information is 20 degrees to the left and the azimuth indicated by the fourth preset azimuth information is 70 degrees to the left, then the azimuth jointly indicated by the third preset azimuth information and the fourth preset azimuth information is to the left. Then, the azimuth indicated by the third azimuth information is to the left.

[0239] Step S403, obtain the fourth audio signal.

[0240] Among them, the fourth audio signal is the audio signal obtained after the second auxiliary sound pickup device picks up the sound signal to be processed. The second auxiliary sound pickup device is located at the azimuth indicated by the third azimuth information and is not in the same plane as the left ear sound pickup device and the right ear sound pickup device.

[0241] For example, in an application scenario where a user uses TWS earphones and a mobile phone for VOLG shooting, if the azimuth indicated by the third azimuth information is left, the top microphone of the mobile phone can be used as the second auxiliary sound pickup device, and the audio signal generated after the top microphone of the mobile phone picks up the sound signal to be processed is the fourth audio signal. Or, if the azimuth indicated by the third azimuth information is right, the bottom microphone of the mobile phone can be used as the second auxiliary sound pickup device, and the audio signal generated after the bottom microphone of the mobile phone picks up the sound signal to be processed is the fourth audio signal.

[0242] Step S404: Obtain the third energy and the fourth energy.

[0243] Among them, the third energy is the energy of the audio signal corresponding to the third azimuth information in the first audio signal and the second audio signal. If the azimuth indicated by the third azimuth information is left, the audio signal corresponding to the third azimuth information in the first audio signal and the second audio signal is the first audio signal, then the third energy is the energy of the first audio signal. Or, if the azimuth indicated by the third azimuth information is right, the audio signal corresponding to the third azimuth information in the first audio signal and the second audio signal is the second audio signal, then the third energy is the energy of the second audio signal. The fourth energy is the energy of the fourth audio signal.

[0244] Step S405: Determine the first azimuth information according to the third energy, the fourth energy, the third preset azimuth information, and the fourth preset azimuth information.

[0245] In specific implementation, if the third energy is greater than the fourth energy, it means that the sound source emitting the sound signal to be processed is close to the user, then the preset azimuth information indicating the azimuth closer to the user among the third preset azimuth information and the fourth preset azimuth information is determined as the first azimuth information. That is, if the azimuth indicated by the third preset azimuth information is closer to the user than the azimuth indicated by the fourth preset azimuth information, the third preset azimuth information is determined as the first azimuth information. Or, if the azimuth indicated by the fourth preset azimuth information is closer to the user than the azimuth indicated by the third preset azimuth information, the fourth preset azimuth information is determined as the first azimuth information. Or,

[0246] If the third energy is less than the fourth energy, it means that the sound source emitting the sound signal to be processed is far from the user and close to the second auxiliary sound pickup device, then the preset azimuth information indicating the azimuth farther from the user among the third preset azimuth information and the fourth preset azimuth information is determined as the first azimuth information. That is, if the azimuth indicated by the third preset azimuth information is farther from the user than the azimuth indicated by the fourth preset azimuth information, the third preset azimuth information is determined as the first azimuth information. Or, if the azimuth indicated by the fourth preset azimuth information is farther from the user than the azimuth indicated by the third preset azimuth information, the fourth preset azimuth information is determined as the first azimuth information.

[0247] In the third method, the binaural difference information is the first interaural time difference and the first interaural intensity difference. Then, according to the binaural difference information, the first azimuth information can be determined. Refer to Figure 9 the implementation method shown. Figure 9 It is a schematic flowchart of another method for determining the first azimuth information provided by an embodiment of the present application. As Figure 9 shown, the method may include the following steps:

[0248] Step S501: Obtain the first preset azimuth information and the second preset azimuth information corresponding to the first interaural time difference, and the third preset azimuth information and the fourth preset azimuth information corresponding to the first interaural intensity difference.

[0249] For the specific implementation method of step S501, reference may be made to the content of the foregoing embodiments, which will not be elaborated here.

[0250] Step S502: Determine the second azimuth information and the third azimuth information.

[0251] Among them, the azimuth indicated by the second azimuth information is the azimuth jointly indicated by the first preset azimuth information and the second preset azimuth information. The azimuth indicated by the third azimuth information is the azimuth jointly indicated by the third preset azimuth information and the fourth preset azimuth information.

[0252] For the specific implementation method of step S502, reference may be made to the content of the foregoing embodiments, which will not be elaborated here.

[0253] Step S503: Obtain the third audio signal and the fourth audio signal.

[0254] Among them, the third audio signal is the audio signal obtained after the first auxiliary sound pickup device picks up the sound signal to be processed. The fourth audio signal is the audio signal obtained after the second auxiliary sound pickup device picks up the sound signal to be processed. The first auxiliary sound pickup device is located at the azimuth indicated by the second azimuth information. The second auxiliary sound pickup device is located at the azimuth indicated by the third azimuth information. The first auxiliary sound pickup device and the second auxiliary sound pickup device are not in the same plane as the left ear sound pickup device and the right ear sound pickup device.

[0255] It should be noted that the first auxiliary sound pickup device and the second auxiliary sound pickup device may be the same or different. That is, the first auxiliary sound pickup device and the second auxiliary sound pickup device may be the same sound pickup device or different sound pickup devices.

[0256] For the specific implementation method of step S503, reference may be made to the content of the foregoing embodiments, which will not be elaborated here.

[0257] Step S504: Obtain the first energy, the second energy, the third energy, and the fourth energy.

[0258] Wherein, the first energy is the energy of the audio signal corresponding to the second azimuth information in the first audio signal and the second audio signal. The second energy is the energy of the third audio signal. The third energy is the energy of the audio signal corresponding to the third azimuth information in the first audio signal and the second audio signal. The fourth energy is the energy of the fourth audio signal.

[0259] The specific implementation manner of step S504 may refer to the content of the foregoing embodiments and will not be elaborated here.

[0260] Step S505: Determine a fourth azimuth information according to the first energy, the second energy, the first preset azimuth information, and the second preset azimuth information.

[0261] The specific implementation manner of step S505 may refer to Figure 6 the specific implementation manner of step S305 in the illustrated embodiment and will not be elaborated here.

[0262] Step S506: Determine a fifth azimuth information according to the third energy, the fourth energy, the third preset azimuth information, and the fourth preset azimuth information.

[0263] The specific implementation manner of step S506 may refer to Figure 8 the specific implementation manner of step S405 in the illustrated embodiment and will not be elaborated here.

[0264] Step S507: Generate a first azimuth information according to the fourth azimuth information and the fifth azimuth information.

[0265] When determining the first azimuth information solely based on the interaural time difference, if the frequencies of the first audio signal and the second audio signal are too high or too low, it will cause a large error in positioning. When determining the first azimuth information solely based on the interaural intensity difference, if the frequencies of the first audio signal and the second audio signal are relatively low, the low-frequency signal may bypass the head, resulting in an insignificant difference between the two ears and inaccurate positioning.

[0266] Therefore, in order to further improve the accuracy of determining the first azimuth information according to the binaural difference information, after determining the fourth azimuth information through the first interaural time difference and determining the fifth azimuth information through the first interaural intensity difference, the fourth azimuth information and the fifth azimuth information can be averaged, and the calculated average value can be determined as the first azimuth information. Alternatively, the fourth azimuth information and the fifth azimuth information can be weighted-averaged, and the calculated weighted average value can be determined as the first azimuth information.

[0267] Step S104: Render the first direct sound signal and the second direct sound signal according to the first azimuth information to generate a recording signal of the sound signal to be processed.

[0268] In specific implementation, a correspondence between binaural room impulse response (BRIR) information and azimuth information can be established in advance. For the convenience of distinction, in this application, the binaural room impulse response information in the correspondence between the pre-established binaural room impulse response information and the azimuth information is denoted as the preset binaural room impulse response information, and the azimuth information in this correspondence is denoted as the seventh preset azimuth information.

[0269] Based on this, according to the first azimuth information, rendering the first direct sound signal and the second direct sound signal to generate a recording signal of the sound signal to be processed can be implemented in the following manner: Obtain the first response information, where the first response information is the preset binaural room impulse response information corresponding to the first azimuth information. According to the first response information, render the first direct sound signal and the second direct sound signal to generate a recording signal of the sound signal to be processed.

[0270] Further, obtaining the first response information can be implemented in the following manner: Find the seventh preset azimuth information that is the same as the first azimuth information in the pre-stored seventh preset azimuth information, and denote it as the sixth azimuth information; Determine the preset binaural room impulse response information corresponding to the sixth azimuth information as the first response information.

[0271] The first response information may include left-ear response information and right-ear response information. Both the left-ear response information and the right-ear response information may include multiple response information. Any one of the response information included in the left-ear response information can be expressed as h l (t′), where t′ represents the acquisition time (or generation time) of this response information. Any one of the response information included in the right-ear response information can be expressed as h r (t′), where t′ represents the acquisition time (or generation time) of this response information.

[0272] In specific implementation, the implementation manner of rendering the first direct sound signal and the second direct sound signal according to the first response information to generate a recording signal of the sound signal to be processed may include various types. For example:

[0273] In the first method, the first direct sound signal can be first transformed to the time-domain coordinate system through inverse Fourier transform to obtain the third direct sound signal in the time-domain coordinate system, and the second direct sound signal can be transformed to the time-domain coordinate system through inverse Fourier transform to obtain the fourth direct sound signal in the time-domain coordinate system. Then, the recording signal of each sound signal in the sound signal to be processed can be generated according to the following formula where, represents the recording signal of any one sound signal in the sound signal to be processed, and p l(t - t′) represents the signal corresponding to the sound signal included in the third direct sound signal, p r (t - t′) represents the signal corresponding to the sound signal included in the fourth direct sound signal.

[0274] After generating the recording signal of each sound signal in the sound signal to be processed in the above manner, the recording signal of the sound signal to be processed can be obtained. The recording signal of the sound signal to be processed includes the recording signals of all the sound signals included in the sound signal to be processed.

[0275] In the second method, the recording signal of each sound signal in the sound signal to be processed can be generated according to the following formula wherein, represents the recording signal of any one sound signal in the sound signal to be processed, H l (k) represents the response information of the frequency point corresponding to the sound signal included in the left-ear frequency-domain response information obtained after converting the left-ear response information to the frequency-domain coordinate system, H r (k) represents the response information of the frequency point corresponding to the sound signal included in the right-ear frequency-domain response information obtained after converting the right-ear response information to the frequency-domain coordinate system, P l (k) represents the signal of the frequency point corresponding to the sound signal included in the first direct sound signal, P r (k) represents the signal of the frequency point corresponding to the sound signal included in the second direct sound signal.

[0276] After generating the recording signal of each sound signal in the sound signal to be processed in the above manner, the recording signal of the sound signal to be processed can be obtained. The recording signal of the sound signal to be processed includes the recording signals of all the sound signals included in the sound signal to be processed.

[0277] In the third method, the recording signal of each sound signal in the sound signal to be processed can be generated according to the following formula wherein, represents the recording signal of any one sound signal in the sound signal to be processed, P l (k) represents the signal of the frequency point corresponding to the sound signal included in the first direct sound signal, P r (k) represents the signal of the frequency point corresponding to the sound signal included in the second direct sound signal, represents the azimuth angle corresponding to the first response information, is equal to the first interaural time difference, N represents the number of frequency points corresponding to the sound signal to be processed, A l represents the result of normalizing the left-ear response information at different azimuths under the frequency point corresponding to the sound signal, that is, wherein, N aIndicates the number of directions corresponding to the preset binaural room impulse response information, A r Indicates the result of normalizing the right-ear response information in different directions at the frequency points corresponding to the sound signal, that is,

[0278] After generating the recording signals of each sound signal in the sound signal to be processed in the above manner, the recording signals of the sound signal to be processed can be obtained. The recording signals of the sound signal to be processed include the recording signals of all the sound signals included in the sound signal to be processed.

[0279] Compared with the first method and the second method, the third method can enhance the sense of space while reducing the modification strength of the original sound signal to be processed, reduce the distortion of the rendered recording signal, make the recording signal closer to the human ear's sense of hearing, and provide a better user experience.

[0280] It can be seen that through the method for generating a recording signal provided by the embodiments of the present application, the electronic device can separately perform direct sound extraction on the first audio signal obtained by the left-ear pickup device picking up the sound signal to be processed and the second audio signal obtained by the right-ear pickup device picking up the same sound signal to be processed, to obtain a first direct sound signal and a second direct sound signal. Then, the direction information of the sound source emitting the sound signal can be determined according to the first direct sound signal and the second direct sound signal. After that, the first direct sound signal and the second direct sound signal can be rendered according to the direction information to obtain the recording signal of the sound signal to be processed. The recording signal obtained by this method fully integrates the direction information of the sound source, can well express the spatial information of the sound signal, and after playback, can achieve the effect of the human ear's sense of hearing, providing a better user experience.

[0281] Each method embodiment described herein can be an independent solution or can be combined according to the internal logic, and all these solutions fall within the protection scope of the present application.

[0282] It can be understood that in the above method embodiments, the methods and operations implemented by the electronic device can also be implemented by components (such as chips, modules, or circuits) available for the electronic device.

[0283] The above embodiments have introduced the method for generating a recording signal provided by the present application. It can be understood that, in order to implement the above functions, an electronic device includes corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0284] Embodiments of the present application can perform functional module division on an electronic device according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiments of the present application is illustrative, only a logical function division, and there may be other division methods in actual implementation.

[0285] Above, in combination with Figures 1 to 9 the method for generating a recording signal provided by the embodiments of the present application has been described in detail. Below, in combination with Figure 10 and Figure 11 the device provided by the embodiments of the present application will be described in detail. It should be understood that the description of the device embodiments corresponds to the description of the method embodiments. Therefore, the content not described in detail can be referred to the above method embodiments. For the sake of brevity, it will not be repeated here.

[0286] Refer to Figure 10 , Figure 10 which is a structural block diagram of a device for generating a recording signal provided by an embodiment of the present application. This device can be a part of an electronic device and be applied in the electronic device. It can also be an electronic device, and the present application does not limit this. As Figure 10 shown, the device 1000 may include: a first acquisition module 1001, a second acquisition module 1002, a determination module 1003, and a generation module 1004. The device 1000 can execute the operations performed by the electronic device in any one of the method embodiments shown in the above Figures 4 to 9 .

[0287] For example, in an alternative embodiment of the present application, the first acquisition module 1001 may be configured to acquire a first audio signal and a second audio signal. The first audio signal is an audio signal obtained by a left ear pickup device picking up the sound signal to be processed, and the second audio signal is an audio signal obtained by a right ear pickup device picking up the sound signal to be processed.

[0288] The second acquisition module 1002 may be configured to acquire a first direct sound signal of the first audio signal and a second direct sound signal of the second audio signal.

[0289] The determination module 1003 may be configured to determine first azimuth information of a sound source that emits the sound signal to be processed according to the first direct sound signal and the second direct sound signal.

[0290] The generation module 1004 may be configured to render the first direct sound signal and the second direct sound signal according to the first azimuth information to generate a recording signal of the sound signal to be processed.

[0291] In a possible implementation manner, the determination module 1003 is configured to determine first azimuth information of a sound source that emits the sound signal to be processed according to the first direct sound signal and the second direct sound signal. Specifically, the determination module 1003 is configured to: generate binaural difference information according to the first direct sound signal and the second direct sound signal, where the binaural difference information includes a first interaural time difference and / or a first interaural intensity difference; determine the first azimuth information according to the binaural difference information.

[0292] In a possible implementation manner, the binaural difference information is the first interaural time difference. The determination module 1003 is configured to determine the first azimuth information according to the binaural difference information. Specifically, the determination module 1003 is configured to: acquire first preset azimuth information and second preset azimuth information corresponding to the first interaural time difference; determine second azimuth information, where the azimuth indicated by the second azimuth information is the azimuth jointly indicated by the first preset azimuth information and the second preset azimuth information; acquire a third audio signal, where the third audio signal is an audio signal obtained by a first auxiliary pickup device picking up the sound signal to be processed, and the first auxiliary pickup device is located at the azimuth indicated by the second azimuth information and is not in the same plane as the left ear pickup device and the right ear pickup device; acquire a first energy and a second energy, where the first energy is the energy of the audio signal corresponding to the second azimuth information in the first audio signal and the second audio signal, and the second energy is the energy of the third audio signal; determine the first azimuth information according to the first energy, the second energy, the first preset azimuth information, and the second preset azimuth information.

[0293] In a possible implementation manner, the binaural difference information is the first interaural intensity difference, and the determining module 1003 is configured to determine the first azimuth information according to the binaural difference information. Specifically: the determining module 1003 is configured to: obtain a third preset azimuth information and a fourth preset azimuth information corresponding to the first interaural intensity difference; determine a third azimuth information, where the azimuth indicated by the third azimuth information is the azimuth jointly indicated by the third preset azimuth information and the fourth preset azimuth information; obtain a fourth audio signal, where the fourth audio signal is an audio signal obtained after the second auxiliary sound pickup device picks up the sound signal to be processed, and the second auxiliary sound pickup device is located at the azimuth indicated by the third azimuth information and is not in the same plane as the left ear sound pickup device and the right ear sound pickup device; obtain a third energy and a fourth energy, where the third energy is the energy of the audio signal corresponding to the third azimuth information in the first audio signal and the second audio signal, and the fourth energy is the energy of the fourth audio signal; and determine the first azimuth information according to the third energy, the fourth energy, the third preset azimuth information, and the fourth preset azimuth information.

[0294] In a possible implementation, the binaural difference information is the first interaural time difference and the first interaural intensity difference. The determining module 1003 is configured to determine the first azimuth information according to the binaural difference information, specifically: The determining module 1003 is configured to: obtain a first preset azimuth information and a second preset azimuth information corresponding to the first interaural time difference, and a third preset azimuth information and a fourth preset azimuth information corresponding to the first interaural intensity difference; determine a second azimuth information and a third azimuth information, where the azimuth indicated by the second azimuth information is the azimuth jointly indicated by the first preset azimuth information and the second preset azimuth information, and the azimuth indicated by the third azimuth information is the azimuth jointly indicated by the third preset azimuth information and the fourth preset azimuth information; obtain a third audio signal and a fourth audio signal, where the third audio signal is an audio signal obtained after a first auxiliary sound pickup device picks up the to-be-processed sound signal, and the fourth audio signal is an audio signal obtained after a second auxiliary sound pickup device picks up the to-be-processed sound signal, the first auxiliary sound pickup device is located at the azimuth indicated by the second azimuth information, the second auxiliary sound pickup device is located at the azimuth indicated by the third azimuth information, and both the first auxiliary sound pickup device and the second auxiliary sound pickup device are not in the same plane as the left ear sound pickup device and the right ear sound pickup device; obtain a first energy, a second energy, a third energy, and a fourth energy, where the first energy is the energy of the audio signal corresponding to the second azimuth information in the first audio signal and the second audio signal, the second energy is the energy of the third audio signal, the third energy is the energy of the audio signal corresponding to the third azimuth information in the first audio signal and the second audio signal, and the fourth energy is the energy of the fourth audio signal; determine a fourth azimuth information according to the first energy, the second energy, the first preset azimuth information, and the second preset azimuth information; determine a fifth azimuth information according to the third energy, the fourth energy, the third preset azimuth information, and the fourth preset azimuth information; generate the first azimuth information according to the fourth azimuth information and the fifth azimuth information.

[0295] In a possible implementation, the generating module 1004 is configured to render the first direct sound signal and the second direct sound signal according to the first azimuth information to generate a recording signal of the to-be-processed sound signal, specifically: The generating module 1004 is configured to: obtain a first response information, where the first response information is a preset binaural room impulse response information corresponding to the first azimuth information; render the first direct sound signal and the second direct sound signal according to the first response information to generate a recording signal of the to-be-processed sound signal.

[0296] That is to say, the device 1000 can implement corresponding to Figures 4 to 9In the method embodiments for generating any of the recorded signals shown, the steps or processes performed by the electronic device. The apparatus 1000 may include modules for performing Figures 4 to 9 the method performed by the electronic device in any of the method embodiments for generating a recorded signal shown. It should be understood that the specific processes of each module performing the corresponding steps above have been described in detail in the above method embodiments for game operation. For the sake of brevity, they will not be repeated here.

[0297] An embodiment of the present application also provides a processing device, which includes at least one processor and a communication interface. The communication interface is used to provide information input and / or output for the at least one processor, and the at least one processor is used to execute the method in the above method embodiments.

[0298] It should be understood that the above processing device may be a chip. For example, referring to Figure 11 , Figure 11 is a structural block diagram of a chip provided in an embodiment of the present application. Figure 11 The chip shown may be a general-purpose processor or a dedicated processor. The chip 1100 may include at least one processor 1101. Among them, the at least one processor 1101 may be used to support Figure 10 the device shown to execute Figures 4 to 9 the technical solutions shown in any of the embodiments.

[0299] Optionally, the chip 1100 may further include a transceiver 1102. The transceiver 1102 is used to receive the control of the processor 1101 and is used to support Figure 10 the device shown to execute Figures 4 to 9 the technical solutions shown in any of the embodiments. Optionally, Figure 11 the chip 1100 shown may further include a storage medium 1103. Specifically, the transceiver 1102 may be replaced by a communication interface, and the communication interface provides information input and / or output for the at least one processor 1101.

[0300] It should be noted that Figure 11The illustrated chip 1100 may be implemented using the following circuits or devices: one or more field programmable gate arrays (FPGAs), programmable logic devices (PLDs), application specific integrated circuits (ASICs), system on chips (SoCs), central processor units (CPUs), network processors (NPs), digital signal processing circuits (DSPs), micro controller units (MCUs), controllers, state machines, gate logic, discrete hardware components, any other suitable circuits, or any combination of circuits capable of performing the various functions described throughout this application.

[0301] In the implementation process, the steps of the above method may be completed by the integrated logic circuit of the hardware in the processor or the instructions in the form of software. The steps of the method disclosed in combination with the embodiments of the present application may be directly embodied as being executed and completed by the hardware processor, or executed and completed by the combination of the hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as random access memory, flash memory, read only memory, programmable read only memory, or electrically erasable programmable memory, registers, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0302] It can be understood that the memory in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include but not be limited to these and any other suitable types of memory.

[0303] According to the method provided by the embodiments of the present application, the embodiments of the present application further provide a computer program product, which includes: a computer program or instruction, when the computer program or instruction runs on a computer, enabling the computer to execute Figures 4 to 9 the method of any one of the embodiments shown.

[0304] According to the method provided by the embodiments of the present application, the embodiments of the present application further provide a computer-readable storage medium, which stores a computer program or instruction, when the computer program or instruction runs on a computer, enabling the computer to execute Figures 4 to 9 the method of any one of the embodiments shown.

[0305] According to the method provided by the embodiments of the present application, the embodiments of the present application also provide an electronic device. The electronic device includes, but is not limited to, mobile phones, tablet computers, personal computers, workstation devices, large-screen devices (such as smart screens, smart TVs, etc.), handheld game consoles, home game consoles, virtual reality devices, augmented reality devices, mixed reality devices, vehicle-mounted intelligent terminals, etc. The electronic device may include the generating device for the recording signal provided in the above embodiments of the present application. The electronic device can be used to execute Figures 4 to 9 the method of any one of the embodiments shown

[0306] According to the method provided by the embodiments of the present application, the embodiments of the present application also provide a system for generating a recording signal. The system may include: a left-ear sound pickup device and a right-ear sound pickup device for picking up the sound signal to be processed, and the above-mentioned electronic device; a first auxiliary sound pickup device and a second auxiliary sound pickup device for picking up the sound signal to be processed are arranged in the electronic device; the left-ear sound pickup device and the right-ear sound pickup device are communicatively connected to the electronic device by wireless or wired means. Among them, the specific functions of the left-ear sound pickup device, the right-ear sound pickup device, the first auxiliary sound pickup device, the second auxiliary sound pickup device and the system can refer to the content of the foregoing method embodiments, which will not be elaborated here. The system can be used to execute Figures 4 to 9 the method of any one of the embodiments shown and obtain corresponding beneficial effects.

[0307] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system, device modules, and electronic devices can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated here.

[0308] In several embodiments provided by the present application, it should be understood that the disclosed system, device, and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0309] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0310] In addition, in each embodiment of the present application, each functional module can be integrated into a processing unit, can exist physically alone for each module, or two or more modules can be integrated into one unit.

[0311] If the above-mentioned function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0312] The audio signal generation device, system, processing device, chip, computer storage medium, computer program product, and electronic device provided in the above embodiments of the present application are all used to execute the methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects corresponding to the methods provided above, and will not be elaborated here.

[0313] It should be understood that in each embodiment of the present application, the execution order of each step should be determined according to its function and internal logic. The size of each step number does not mean the sequence of execution, and it does not limit the implementation process of the embodiment.

[0314] Each part of this specification is described in a progressive manner. For the same or similar parts between each embodiment, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the audio signal generation device, system, chip, computer storage medium, computer program product, and electronic device, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the descriptions in the method embodiments for the relevant parts.

[0315] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0316] The above-described embodiments of the present application do not constitute a limitation to the protection scope of the present application.

Claims

1. A method for generating a recording signal, characterized in that, The method includes: Obtaining a first audio signal and a second audio signal, where the first audio signal is an audio signal obtained by a left ear pickup device picking up a sound signal to be processed, and the second audio signal is an audio signal obtained by a right ear pickup device picking up the sound signal to be processed; Obtaining a first direct sound signal of the first audio signal and a second direct sound signal of the second audio signal; Generating binaural difference information according to the first direct sound signal and the second direct sound signal, where the binaural difference information is a first interaural time difference; Determining first azimuth information of a sound source that emits the sound signal to be processed according to the binaural difference information; specifically including: obtaining first preset azimuth information and second preset azimuth information corresponding to the first interaural time difference; determining second azimuth information, where the azimuth indicated by the second azimuth information is the azimuth jointly indicated by the first preset azimuth information and the second preset azimuth information; obtaining a third audio signal, where the third audio signal is an audio signal obtained by a first auxiliary pickup device picking up the sound signal to be processed, and the first auxiliary pickup device is located at the azimuth indicated by the second azimuth information and is not in the same plane as the left ear pickup device and the right ear pickup device; obtaining a first energy and a second energy, where the first energy is the energy of the audio signal corresponding to the second azimuth information in the first audio signal and the second audio signal, and the second energy is the energy of the third audio signal; determining the first azimuth information according to the first energy, the second energy, the first preset azimuth information and the second preset azimuth information; rendering the first direct sound signal and the second direct sound signal according to the first azimuth information to generate a recording signal of the sound signal to be processed.

2. The method according to claim 1, characterized in that, The rendering the first direct sound signal and the second direct sound signal according to the first azimuth information to generate a recording signal of the sound signal to be processed includes: Obtaining first response information, where the first response information is preset binaural room impulse response information corresponding to the first azimuth information; Rendering the first direct sound signal and the second direct sound signal according to the first response information to generate a recording signal of the sound signal to be processed.

3. A method for generating a recording signal, characterized in that, The method includes: Obtaining a first audio signal and a second audio signal, where the first audio signal is an audio signal obtained by a left ear pickup device picking up a sound signal to be processed, and the second audio signal is an audio signal obtained by a right ear pickup device picking up the sound signal to be processed; Obtaining a first direct sound signal of the first audio signal and a second direct sound signal of the second audio signal; Generating binaural difference information according to the first direct sound signal and the second direct sound signal, where the binaural difference information is a first interaural intensity difference; Determine the first azimuth information of the sound source that emits the to-be-processed sound signal according to the binaural difference information; specifically including: obtaining the third preset azimuth information and the fourth preset azimuth information corresponding to the first interaural intensity difference; determining the third azimuth information, where the azimuth indicated by the third azimuth information is the azimuth jointly indicated by the third preset azimuth information and the fourth preset azimuth information; obtaining a fourth audio signal, where the fourth audio signal is an audio signal obtained after the second auxiliary sound pickup device picks up the to-be-processed sound signal, and the second auxiliary sound pickup device is located at the azimuth indicated by the third azimuth information and is not in the same plane as the left ear sound pickup device and the right ear sound pickup device; obtaining a third energy and a fourth energy, where the third energy is the energy of the audio signal corresponding to the third azimuth information in the first audio signal and the second audio signal, and the fourth energy is the energy of the fourth audio signal; determining the first azimuth information according to the third energy, the fourth energy, the third preset azimuth information, and the fourth preset azimuth information; Render the first direct sound signal and the second direct sound signal according to the first azimuth information to generate a recording signal of the to-be-processed sound signal.

4. The method according to claim 3, characterized in that The step of rendering the first direct sound signal and the second direct sound signal according to the first azimuth information to generate a recording signal of the to-be-processed sound signal includes: Obtaining a first response information, where the first response information is the preset binaural room impulse response information corresponding to the first azimuth information; Render the first direct sound signal and the second direct sound signal according to the first response information to generate a recording signal of the to-be-processed sound signal.

5. A method for generating a recording signal, characterized in that, The method includes: Obtaining a first audio signal and a second audio signal, where the first audio signal is an audio signal obtained after the left ear sound pickup device picks up the to-be-processed sound signal, and the second audio signal is an audio signal obtained after the right ear sound pickup device picks up the to-be-processed sound signal; Obtaining a first direct sound signal of the first audio signal and a second direct sound signal of the second audio signal; Generating binaural difference information according to the first direct sound signal and the second direct sound signal, where the binaural difference information is a first interaural time difference and a first interaural intensity difference; Determine the first azimuth information of the sound source that emits the to-be-processed sound signal according to the binaural difference information; specifically including: obtaining the first preset azimuth information and the second preset azimuth information corresponding to the first interaural time difference, and the third preset azimuth information and the fourth preset azimuth information corresponding to the first interaural intensity difference; determining the second azimuth information and the third azimuth information, where the azimuth indicated by the second azimuth information is the azimuth jointly indicated by the first preset azimuth information and the second preset azimuth information, and the azimuth indicated by the third azimuth information is the azimuth jointly indicated by the third preset azimuth information and the fourth preset azimuth information; obtaining a third audio signal and a fourth audio signal, where the third audio signal is an audio signal obtained after the first auxiliary pickup device picks up the to-be-processed sound signal, and the fourth audio signal is an audio signal obtained after the second auxiliary pickup device picks up the to-be-processed sound signal, the first auxiliary pickup device is located at the azimuth indicated by the second azimuth information, the second auxiliary pickup device is located at the azimuth indicated by the third azimuth information, and the first auxiliary pickup device and the second auxiliary pickup device are not in the same plane as the left ear pickup device and the right ear pickup device; obtaining a first energy, a second energy, a third energy, and a fourth energy, where the first energy is the energy of the audio signal corresponding to the second azimuth information in the first audio signal and the second audio signal, the second energy is the energy of the third audio signal, the third energy is the energy of the audio signal corresponding to the third azimuth information in the first audio signal and the second audio signal, and the fourth energy is the energy of the fourth audio signal; determining a fourth azimuth information according to the first energy, the second energy, the first preset azimuth information, and the second preset azimuth information; determining a fifth azimuth information according to the third energy, the fourth energy, the third preset azimuth information, and the fourth preset azimuth information; generating the first azimuth information according to the fourth azimuth information and the fifth azimuth information; Render the first direct sound signal and the second direct sound signal according to the first azimuth information to generate a recording signal of the to-be-processed sound signal.

6. The method according to claim 5, wherein The rendering the first direct sound signal and the second direct sound signal according to the first azimuth information to generate a recording signal of the to-be-processed sound signal includes: Obtaining a first response information, where the first response information is a preset binaural room impulse response information corresponding to the first azimuth information; Render the first direct sound signal and the second direct sound signal according to the first response information to generate a recording signal of the to-be-processed sound signal.

7. A generating device for a recording signal, characterized in that, The device includes: A first acquisition module, configured to acquire a first audio signal and a second audio signal, where the first audio signal is an audio signal obtained after the left ear pickup device picks up the to-be-processed sound signal, and the second audio signal is an audio signal obtained after the right ear pickup device picks up the to-be-processed sound signal; A second acquisition module, configured to acquire a first direct sound signal of the first audio signal and a second direct sound signal of the second audio signal; A determination module, configured to generate binaural difference information according to the first direct sound signal and the second direct sound signal, where the binaural difference information is a first interaural time difference; The determination module is further configured to determine first azimuth information of a sound source that emits the to-be-processed sound signal according to the binaural difference information; specifically including: acquiring first preset azimuth information and second preset azimuth information corresponding to the first interaural time difference; determining second azimuth information, where the azimuth indicated by the second azimuth information is the azimuth jointly indicated by the first preset azimuth information and the second preset azimuth information; acquiring a third audio signal, where the third audio signal is an audio signal obtained by a first auxiliary sound pickup device picking up the to-be-processed sound signal, and the first auxiliary sound pickup device is located at the azimuth indicated by the second azimuth information and is not in the same plane as the left ear sound pickup device and the right ear sound pickup device; acquiring a first energy and a second energy, where the first energy is the energy of the audio signals corresponding to the second azimuth information in the first audio signal and the second audio signal, and the second energy is the energy of the third audio signal; determining the first azimuth information according to the first energy, the second energy, the first preset azimuth information, and the second preset azimuth information; A generation module, configured to render the first direct sound signal and the second direct sound signal according to the first azimuth information to generate a recording signal of the to-be-processed sound signal.

8. The recording signal generation device according to claim 7, wherein The generation module is specifically configured to: acquire first response information, where the first response information is preset binaural room impulse response information corresponding to the first azimuth information; Render the first direct sound signal and the second direct sound signal according to the first response information to generate a recording signal of the to-be-processed sound signal.

9. A generating device for a recording signal, characterized in that, The device includes: A first acquisition module, configured to acquire a first audio signal and a second audio signal, where the first audio signal is an audio signal obtained by a left ear sound pickup device picking up the to-be-processed sound signal, and the second audio signal is an audio signal obtained by a right ear sound pickup device picking up the to-be-processed sound signal; A second acquisition module, configured to acquire a first direct sound signal of the first audio signal and a second direct sound signal of the second audio signal; A determination module, configured to generate binaural difference information according to the first direct sound signal and the second direct sound signal, where the binaural difference information is a first interaural intensity difference; The determining module is further configured to determine first azimuth information of a sound source that emits the to-be-processed sound signal according to the binaural difference information; specifically, it includes: obtaining third preset azimuth information and fourth preset azimuth information corresponding to the first interaural intensity difference; determining third azimuth information, where the azimuth indicated by the third azimuth information is the azimuth jointly indicated by the third preset azimuth information and the fourth preset azimuth information; obtaining a fourth audio signal, where the fourth audio signal is an audio signal obtained after the second auxiliary sound pickup device picks up the to-be-processed sound signal, and the second auxiliary sound pickup device is located at the azimuth indicated by the third azimuth information and is not in the same plane as the left ear sound pickup device and the right ear sound pickup device; obtaining third energy and fourth energy, where the third energy is the energy of the audio signal corresponding to the third azimuth information in the first audio signal and the second audio signal, and the fourth energy is the energy of the fourth audio signal; determining the first azimuth information according to the third energy, the fourth energy, the third preset azimuth information, and the fourth preset azimuth information; The generating module is configured to render the first direct sound signal and the second direct sound signal according to the first azimuth information to generate a recording signal of the to-be-processed sound signal.

10. The generating device for the recording signal according to claim 9, wherein The generating module is specifically configured to: obtain first response information, where the first response information is preset binaural room impulse response information corresponding to the first azimuth information; Render the first direct sound signal and the second direct sound signal according to the first response information to generate a recording signal of the to-be-processed sound signal.

11. A generating device for a recording signal, characterized in that, The device includes: The first obtaining module is configured to obtain a first audio signal and a second audio signal, where the first audio signal is an audio signal obtained after the left ear sound pickup device picks up the to-be-processed sound signal, and the second audio signal is an audio signal obtained after the right ear sound pickup device picks up the to-be-processed sound signal; The second obtaining module is configured to obtain a first direct sound signal of the first audio signal and a second direct sound signal of the second audio signal; The determining module is configured to generate binaural difference information according to the first direct sound signal and the second direct sound signal, where the binaural difference information is a first interaural time difference and a first interaural intensity difference; The determining module is further configured to determine first azimuth information of a sound source that emits the to-be-processed sound signal according to the binaural difference information. Specifically, it includes: obtaining first preset azimuth information and second preset azimuth information corresponding to the first interaural time difference, and third preset azimuth information and fourth preset azimuth information corresponding to the first interaural intensity difference; determining second azimuth information and third azimuth information, where the azimuth indicated by the second azimuth information is the azimuth jointly indicated by the first preset azimuth information and the second preset azimuth information, and the azimuth indicated by the third azimuth information is the azimuth jointly indicated by the third preset azimuth information and the fourth preset azimuth information; obtaining a third audio signal and a fourth audio signal, where the third audio signal is an audio signal obtained after the first auxiliary sound pickup device picks up the to-be-processed sound signal, and the fourth audio signal is an audio signal obtained after the second auxiliary sound pickup device picks up the to-be-processed sound signal, the first auxiliary sound pickup device is located at the azimuth indicated by the second azimuth information, the second auxiliary sound pickup device is located at the azimuth indicated by the third azimuth information, and the first auxiliary sound pickup device and the second auxiliary sound pickup device are not in the same plane as the left ear sound pickup device and the right ear sound pickup device; obtaining first energy, second energy, third energy, and fourth energy, where the first energy is the energy of the audio signal corresponding to the second azimuth information in the first audio signal and the second audio signal, the second energy is the energy of the third audio signal, the third energy is the energy of the audio signal corresponding to the third azimuth information in the first audio signal and the second audio signal, and the fourth energy is the energy of the fourth audio signal; determining fourth azimuth information according to the first energy, the second energy, the first preset azimuth information, and the second preset azimuth information; determining fifth azimuth information according to the third energy, the fourth energy, the third preset azimuth information, and the fourth preset azimuth information; generating the first azimuth information according to the fourth azimuth information and the fifth azimuth information; The generating module is configured to render the first direct sound signal and the second direct sound signal according to the first azimuth information to generate a recording signal of the to-be-processed sound signal.

12. The generating device for the recording signal according to claim 11, wherein: The generating module is specifically configured to: obtain first response information, where the first response information is preset binaural room impulse response information corresponding to the first azimuth information; Render the first direct sound signal and the second direct sound signal according to the first response information to generate a recording signal of the to-be-processed sound signal.

13. An electronic device, characterized in that, The electronic device includes: one or more processors and one or more memories; the one or more memories store computer programs or instructions, and when the computer programs or instructions are executed by the one or more processors, the electronic device executes the method according to any one of claims 1-6.

14. A generation system for recording signals, characterized in that, including: A left ear sound pickup device and a right ear sound pickup device for picking up a sound signal to be processed, and the electronic device according to claim 13; a first auxiliary sound pickup device and a second auxiliary sound pickup device for picking up the sound signal to be processed are provided in the electronic device; the left ear sound pickup device and the right ear sound pickup device are communicatively connected to the electronic device by wireless or wired means.

15. A computer-readable storage medium, characterized in that, A computer program or instruction is stored in the computer-readable storage medium, and when the computer program or instruction is executed, the method according to any one of claims 1-6 is executed.

Citation Information

Patent Citations

  • Method and apparatus for generating a binaural audio signal

    CN101933344A

  • Audio processing method and electronic equipment

    CN114727212A