A video recording method and electronic device
By acquiring image and audio signals and using facial recognition and bone conduction audio signals to generate stereo audio signals, the problem of mismatch between sound and visual orientation in TWS earphone video recording has been solved, achieving clear and stereo video recording of user voice.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-17
- Publication Date
- 2026-04-03
AI Technical Summary
When recording video while the user is wearing TWS earphones, the sound image location of the user's voice does not match the image location, resulting in the user's voice not having a stereo effect in the recorded video.
The first electronic device acquires image and audio signals, uses facial recognition and bone conduction audio signals to determine the acoustic image location of the user's voice, and generates stereo audio signals to match the acoustic image location with the visual image location.
In the recorded video, the user's voice is clear and the audio-visual orientation matches the visual orientation, highlighting the user's audio information.
Smart Images

Figure CN115225840B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminal and communication technology, and in particular to video recording methods and electronic devices. Background Technology
[0002] With the improvement of mobile phone shooting functions, more and more users like to record videos, including themselves, when they are far away from electronic devices. In many cases, users hope that their voices are clear and have a sense of direction in the recorded videos.
[0003] With the increasing popularity of true wireless stereo (TWS) earbuds, when a user wears TWS earbuds and is at a distance from an electronic device while recording video, the TWS earbuds can minimize ambient sound signals and capture a clear audio signal from the user. This audio signal is then converted into an audio signal and transmitted wirelessly to the electronic device. The electronic device can then process the received audio signal and the recorded image to produce the video. In this video, the user's voice is clear.
[0004] However, when recording videos using the above method, although the user's voice is stereo when speaking, it is no longer stereo after being recorded and processed by the TWS earphones. As a result, the user's voice in the video does not have a stereo feel, and the sound image location of the user's voice cannot match the user's visual location. Summary of the Invention
[0005] This application provides a video recording method and an electronic device that, when the user is far away from the electronic device and the electronic device is recording video including the user, ensures that the user's voice is clear in the recorded video even when the sound image location of the user's voice matches the user's visual location.
[0006] In a first aspect, this application provides a video recording method, comprising: during the recording of video by a first electronic device, the first electronic device acquires an image and a first audio signal; the first electronic device determines the acoustic image location of a user's voice in the image based on the image and the first audio signal; the first electronic device generates a stereo audio signal based on the acoustic image location of the user's voice and the user's audio signal; the user's audio signal is acquired by a second electronic device and sent to the first electronic device; and the first electronic device generates a video based on the image and the stereo audio signal.
[0007] In the above embodiments, the first electronic device calculates the acoustic image location of the user's voice and performs stereo processing on the user's audio signal, ensuring that the acoustic image location of the user's voice reconstructed from the stereo audio signal in the generated video matches the user's visual location. The user's audio signal is acquired by a second electronic device, which can be headphones worn by the user. This ensures the acquired user audio signal is clear, making the stereo user audio signal generated by the first electronic device also clear. Therefore, in the generated video, the acoustic image location of the user's voice matches the user's visual location, and the user's voice is clear. Furthermore, the stereo audio signal generated based on the acoustic image location of the user's voice and the user's audio signal primarily includes the user's voice information without mixing in other environmental sound information. In scenarios where it is necessary to highlight the user's voice information in the video, the method in this embodiment can achieve the goal of emphasizing the user's voice.
[0008] In conjunction with some embodiments of the first aspect, in some embodiments, the first electronic device determines the acoustic image location of a user's voice in an image based on the image and the first audio signal, specifically including: the first electronic device performing face recognition on the image to obtain the pixel position of the face; the first electronic device determining the sound source location of the user based on the first audio signal; the first electronic device determining the pixel position of the user in the image based on the sound source location of the user and the pixel position of the face; and the first electronic device determining the acoustic image location of the user's voice in the image based on the pixel position of the user in the image.
[0009] In the above embodiment, the first electronic device calculates the acoustic image location of the user's voice by collecting the first audio signal and the image. The acoustic image location of the user's voice is used as a parameter of the algorithm when generating the stereo audio signal, so that the acoustic image location of the user's voice restored by the stereo audio signal matches the visual image location of the user in the image.
[0010] In conjunction with some embodiments of the first aspect, in some embodiments, the first electronic device determines the acoustic image location of the user's voice in the image based on the image and the first audio signal, specifically including: the first electronic device performing face recognition on the image to obtain the pixel position of the face; the first electronic device using bone conduction audio signals, combined with the first audio signal, to determine the sound source location of the user; the bone conduction audio signal being acquired by the second electronic device and sent to the first electronic device; the first electronic device determining the pixel position of the user in the image based on the sound source location of the user and the pixel position of the face; and the first electronic device determining the acoustic image location of the user's voice in the image based on the pixel position of the user in the image.
[0011] In the above embodiments, during the process of calculating the acoustic image location of the user's voice, the bone conduction audio signal can be used to filter out the part of the first audio signal that is strongly correlated with the user's voice information, thereby improving the accuracy of calculating the acoustic image location of the user's voice.
[0012] In conjunction with some embodiments of the first aspect, in some embodiments, the first electronic device generates a stereo audio signal based on the sound image location of the user's voice and the user's audio signal, specifically including: the first electronic device generating an ambient stereo audio signal; the first electronic device generating the stereo audio signal based on the sound image location of the user's voice, the user's audio signal, and the ambient stereo audio signal.
[0013] In the above embodiments, during the generation of stereo audio signals, not only user audio signals but also ambient audio signals are utilized, so that the generated stereo audio signals contain not only user voice information but also other sound information from the environment.
[0014] In conjunction with some embodiments of the first aspect, in some embodiments, the first electronic device generates an ambient stereo audio signal, specifically including: the first electronic device performing adaptive blocking filtering on the first audio signal based on the user's audio signal to filter out the user's voice information in the first audio signal; and the first electronic device generating the ambient stereo audio signal based on the filtered first audio signal.
[0015] In the above embodiments, the first audio signal acquired by the first electronic device may contain clearer information about other sounds in the environment. The first electronic device uses the acquired first audio signal to filter out the user's voice information, thus obtaining other sound information from the actual environment.
[0016] In conjunction with some embodiments of the first aspect, in some embodiments, the first electronic device generates an ambient stereo audio signal, specifically including: the first electronic device uses the first audio signal to generate a stereo first audio signal; the first electronic device performs adaptive blocking filtering on the stereo first audio signal according to the user audio signal to filter out the user's voice information in the stereo first audio signal, thereby obtaining an ambient stereo audio signal.
[0017] In the above embodiments, the first audio signal acquired by the first electronic device may contain clearer information about other sounds in the environment. The first electronic device uses the acquired first audio signal to filter out the user's voice information, thus obtaining other sound information from the actual environment.
[0018] In conjunction with some embodiments of the first aspect, in some embodiments, the first electronic device generates the stereo audio signal based on the sound image location of the user's voice, the user's audio signal, and the ambient stereo audio signal. Specifically, this includes: the first electronic device generating a user stereo audio signal based on the sound image location of the user's voice and the user's audio signal; the first electronic device enhancing the user stereo audio signal without altering the ambient stereo audio signal; and the first electronic device generating the stereo audio signal based on the enhanced user stereo audio signal and the ambient stereo audio signal.
[0019] In the above embodiments, when the user stereo audio signal includes both user stereo audio signal and ambient stereo audio signal, audio zoom can be performed. When the user is closer to the first electronic device in the image, the user's voice can become louder, while the ambient sound remains unchanged. When the user is farther away from the first electronic device in the image, the user's voice can become quieter, while the ambient sound remains unchanged.
[0020] In conjunction with some embodiments of the first aspect, in some embodiments, the first electronic device generates the stereo audio signal based on the sound image location of the user's voice, the user's audio signal, and the ambient stereo audio signal. Specifically, this includes: the first electronic device generating a user stereo audio signal based on the sound image location of the user's voice and the user's audio signal; the first electronic device enhancing the user stereo audio signal while suppressing the ambient stereo audio signal; and the first electronic device generating the stereo audio signal based on the enhanced user stereo audio signal and the suppressed ambient stereo audio signal.
[0021] In the above embodiments, when the user stereo audio signal includes both user stereo audio signal and ambient stereo audio signal, audio zoom can be performed. When the user is closer to the first electronic device in the image, the user's voice can become louder, while the ambient sound becomes quieter. When the user is farther away from the first electronic device in the image, the user's voice can become quieter, while the ambient sound becomes quieter.
[0022] In conjunction with some embodiments of the first aspect, in some embodiments, the user audio signal is obtained by the second electronic device performing joint noise reduction processing on the second audio signal based on the bone conduction audio signal to remove other sound information from the environment surrounding the second electronic device in the second audio signal.
[0023] In the above embodiments, the second audio signal is filtered using bone conduction audio signals, so that the resulting user audio signal is mainly composed of the user's voice information. When the first electronic device uses the user audio signal to calculate the sound image location of the user's voice, the influence of other sound information in the environment on the calculation result is smaller, making the calculation result more accurate.
[0024] In conjunction with some embodiments of the first aspect, in some embodiments, the user audio signal is obtained by the second electronic device performing noise reduction processing on the second audio signal to remove other sound information from the environment surrounding the second electronic device.
[0025] In the above embodiment, the second audio signal is filtered to remove some sound information in the environment, so that the user's voice information is preserved in the obtained user audio signal. When the first electronic device uses the user audio signal to calculate the sound image location of the user's voice, the influence of other sound information in the environment on the calculation result is small, so that the calculation result is accurate.
[0026] Secondly, this application provides an electronic device comprising: one or more processors and a memory; the memory is coupled to the one or more processors and is used to store computer program code, the computer program code including computer instructions, wherein the one or more processors call the computer instructions to cause the electronic device to perform: during video recording, acquiring an image and a first audio signal; determining the acoustic image location of a user's voice in the image based on the image and the first audio signal; generating a stereo audio signal based on the acoustic image location of the user's voice and the user's audio signal; the user's audio signal being acquired by a second electronic device and sent to the electronic device; and generating video based on the image and the stereo audio signal.
[0027] In the above embodiments, the first electronic device calculates the acoustic image location of the user's voice and performs stereo processing on the user's audio signal, ensuring that the acoustic image location of the user's voice reconstructed from the stereo audio signal in the generated video matches the user's visual location. The user's audio signal is acquired by a second electronic device, which can be headphones worn by the user. This ensures the acquired user audio signal is clear, making the stereo user audio signal generated by the first electronic device also clear. Therefore, in the generated video, the acoustic image location of the user's voice matches the user's visual location, and the user's voice is clear. Furthermore, the stereo audio signal generated based on the acoustic image location of the user's voice and the user's audio signal primarily includes the user's voice information without mixing in other environmental sound information. In scenarios where it is necessary to highlight the user's voice information in the video, the method in this embodiment can achieve the goal of emphasizing the user's voice.
[0028] In conjunction with some embodiments of the second aspect, in some embodiments, the one or more processors are specifically configured to invoke the computer instructions to cause the electronic device to perform: face recognition on the image to obtain the pixel position of the face; determining the location of the user's sound source based on the first audio signal; determining the pixel position of the user in the image based on the location of the user's sound source and the pixel position of the face; and determining the sound image location of the user's voice in the image based on the pixel position of the user in the image.
[0029] In the above embodiment, the first electronic device calculates the acoustic image location of the user's voice by collecting the first audio signal and the image. The acoustic image location of the user's voice is used as a parameter of the algorithm when generating the stereo audio signal, so that the acoustic image location of the user's voice restored by the stereo audio signal matches the visual image location of the user in the image.
[0030] In conjunction with some embodiments of the second aspect, in some embodiments, the one or more processors are specifically configured to invoke the computer instructions to cause the electronic device to perform: face recognition on the image to obtain the pixel position of the face; using a bone conduction audio signal, combined with the first audio signal, to determine the location of the user's sound source; the bone conduction audio signal being acquired by the second electronic device and sent to the electronic device; determining the pixel position of the user in the image based on the location of the user's sound source and the pixel position of the face; and determining the sound image location of the user's voice in the image based on the pixel position of the user in the image.
[0031] In the above embodiments, during the process of calculating the acoustic image location of the user's voice, the bone conduction audio signal can be used to filter out the part of the first audio signal that is strongly correlated with the user's voice information, thereby improving the accuracy of calculating the acoustic image location of the user's voice.
[0032] In conjunction with some embodiments of the second aspect, in some embodiments, the one or more processors are specifically configured to invoke the computer instructions to cause the electronic device to perform: generating an ambient stereo audio signal; generating the stereo audio signal based on the sound image location of the user's voice and the user's audio signal, as well as the ambient stereo audio signal.
[0033] In the above embodiments, during the generation of stereo audio signals, not only user audio signals but also ambient audio signals are utilized, so that the generated stereo audio signals contain not only user voice information but also other sound information from the environment.
[0034] In conjunction with some embodiments of the second aspect, in some embodiments, the one or more processors are specifically configured to invoke the computer instructions to cause the electronic device to perform: adaptive blocking filtering of the first audio signal based on the user audio signal to filter out the user's voice information in the first audio signal; and generating the ambient stereo audio signal based on the filtered first audio signal.
[0035] In the above embodiments, the first audio signal acquired by the first electronic device may contain clearer information about other sounds in the environment. The first electronic device uses the acquired first audio signal to filter out the user's voice information, thus obtaining other sound information from the actual environment.
[0036] In conjunction with some embodiments of the second aspect, in some embodiments, the one or more processors are specifically configured to invoke the computer instructions to cause the electronic device to perform: adaptive blocking filtering on the stereo first audio signal based on the user audio signal, filtering out the user's voice information in the stereo first audio signal, and obtaining an ambient stereo audio signal.
[0037] In the above embodiments, the first audio signal acquired by the first electronic device may contain clearer information about other sounds in the environment. The first electronic device uses the acquired first audio signal to filter out the user's voice information, thus obtaining other sound information from the actual environment.
[0038] In conjunction with some embodiments of the second aspect, in some embodiments, the one or more processors are specifically configured to invoke the computer instructions to cause the electronic device to perform: generating a user stereo audio signal based on the sound image location of the user's voice and the user's audio signal; enhancing the user stereo audio signal without changing the ambient stereo audio signal; and generating the stereo audio signal based on the enhanced user stereo audio signal and the ambient stereo audio signal.
[0039] In the above embodiments, when the user stereo audio signal includes both user stereo audio signal and ambient stereo audio signal, audio zoom can be performed. When the user is closer to the first electronic device in the image, the user's voice can become louder, while the ambient sound remains unchanged. When the user is farther away from the first electronic device in the image, the user's voice can become quieter, while the ambient sound remains unchanged.
[0040] In conjunction with some embodiments of the second aspect, in some embodiments, the one or more processors are specifically configured to invoke the computer instructions to cause the electronic device to perform: generating a user stereo audio signal based on the sound image location of the user's voice and the user's audio signal; enhancing the user stereo audio signal while suppressing the ambient stereo audio signal; and generating the stereo audio signal based on the enhanced user stereo audio signal and the suppressed ambient stereo audio signal.
[0041] In the above embodiments, when the user stereo audio signal includes both user stereo audio signal and ambient stereo audio signal, audio zoom can be performed. When the user is closer to the first electronic device in the image, the user's voice can become louder, while the ambient sound becomes quieter. When the user is farther away from the first electronic device in the image, the user's voice can become quieter, while the ambient sound becomes quieter.
[0042] Thirdly, embodiments of this application provide a chip system applied to an electronic device. The chip system includes one or more processors, which are used to invoke computer instructions to cause the electronic device to perform the method described in any of the embodiments of the first aspect.
[0043] Fourthly, embodiments of this application provide a computer program product containing instructions, characterized in that, when the computer program product is run on an electronic device, it causes the electronic device to perform the method described in any embodiment of the first aspect.
[0044] Fifthly, embodiments of this application provide a computer-readable storage medium including instructions, characterized in that, when the instructions are executed on an electronic device, the electronic device performs the method described in any embodiment of the first aspect.
[0045] Understandably, the chip system provided in the third aspect, the computer program product provided in the fourth aspect, and the computer storage medium provided in the fifth aspect are all used to execute the methods provided in the embodiments of this application. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods, and will not be repeated here. Attached Figure Description
[0046] Figure 1a This application provides a schematic diagram of the structure of the world coordinate system, camera coordinate system, and image plane coordinate system in an embodiment.
[0047] Figure 1b , Figure 1c This is a schematic diagram of the image pixel coordinate system provided in an embodiment of this application;
[0048] Figure 2a , Figure 2b This is a schematic diagram of video recording in one of the solutions provided in this application embodiment;
[0049] Figure 3 This is a schematic diagram of the structure of the communication system 100 provided in the embodiments of this application;
[0050] Figure 4 This is a schematic diagram of the structure of the first electronic device provided in the embodiments of this application;
[0051] Figure 5 This is a schematic diagram of the structure of the second electronic device provided in the embodiments of this application;
[0052] Figure 6 This is a signaling interaction diagram of the video recording method in an embodiment of this application;
[0053] Figure 7 This is a flowchart illustrating how the first electronic device in this application determines the acoustic image location of the user's voice corresponding to the user's audio signal. Detailed Implementation
[0054] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this application refers to and includes any or all possible combinations of one or more of the listed items.
[0055] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.
[0056] To facilitate understanding, the relevant terms and concepts involved in the embodiments of this application will be introduced below.
[0057] (1) Image orientation:
[0058] In this embodiment, the user's visual orientation refers to the position of the user in the real world relative to the center of the camera, determined from the image including the user, when the electronic device acquires the image. The reference coordinate system for this position can be the camera coordinate system.
[0059] Understandably, in some cases, even if the user's actual position relative to the camera center remains unchanged, the image may perceive that the user's position has changed. Examples include images obtained by electronic devices altering camera parameters (such as focal length), or images obtained by cropping images that include the user.
[0060] Determining the orientation of an image involves the world coordinate system, camera coordinate system, image plane coordinate system, and image pixel coordinate system.
[0061] Figure 1a Examples of the world coordinate system, camera coordinate system, and image plane coordinate system provided in embodiments of this application are shown.
[0062] The world coordinate system is represented by O3-Xw-Yw-Zw. Using the world coordinate system, the user's coordinates in the real world can be obtained.
[0063] The camera coordinate system is represented by O1-Xc-Yc-Zc, where O1 is the optical center of the camera and the origin of the camera coordinate system. The Xc, Yc, and Zc axes are the coordinate axes of the camera coordinate system, with the Zc axis being the principal optical axis.
[0064] The coordinates of a point in the world coordinate system can be transformed into the camera coordinate system through rigid body transformation.
[0065] The camera captures the light reflected from the user and projects this light onto an imaging plane to obtain an optical image of the user. An image plane coordinate system can be established on this imaging plane.
[0066] The image plane coordinate system is represented by O2-XY, where O2 is the center of the optical image and the origin of the image plane coordinate system. The X-axis and Y-axis are parallel to the Xc-axis and Yc-axis.
[0067] The coordinates of a point in the camera coordinate system can be transformed to the image plane coordinate system through perspective projection.
[0068] Figure 1b , Figure 1c A schematic diagram of the image pixel coordinate system provided in an embodiment of this application is shown.
[0069] Electronic devices can process optical images in an image plane to obtain images that can be displayed on a screen.
[0070] like Figure 1bAs shown, in some embodiments, an image plane coordinate system, denoted by O2-XY, is established in the image plane. The electronic device can directly display the image corresponding to the image on the display screen without cropping the image in the image plane. An image pixel coordinate system, denoted by OUV, is established on this image, with units in pixels. O is a vertex of the image, and the U-axis and V-axis are parallel to the X-axis and Y-axis, respectively.
[0071] like Figure 1c As shown, in some embodiments, an image plane coordinate system, denoted as O2-XY, is established in the image plane. The electronic device can crop the image in the image plane and display the image corresponding to the cropped image on the display screen. An image pixel coordinate system, denoted as OUV, is established on this image. O is a vertex of the image corresponding to the cropped image, and the U-axis and V-axis are parallel to the X-axis and Y-axis, respectively.
[0072] In other embodiments, the electronic device can also perform other processing on the image in the image plane, such as zooming, to obtain an image that can be displayed on the screen. Then, an image pixel coordinate system can be established with the vertices of the image as the origin.
[0073] An electronic device can determine the user's position relative to the camera's center using a point on the user's body. Once the electronic device determines this position relative to the camera's center, it can convert it into the pixel coordinates of the corresponding pixel in the image's pixel coordinate system. Similarly, when the electronic device obtains the pixel coordinates of a pixel in the image's pixel coordinate system, it can also use these pixel coordinates to determine the position of the corresponding point relative to the camera's center. Thus, the electronic device can obtain the user's pixel coordinates in the image from the user's position relative to the camera's center, and vice versa. The specific process will be described below and will not be elaborated upon here.
[0074] In this embodiment, a user's view orientation in the camera coordinate system can be represented in several ways, including the azimuth angle relative to the center of the camera coordinate system. This azimuth angle can include azimuth and pitch angles. For example, the pitch angle is the angle from the object's image to the Zc axis, which is m°, and the azimuth angle is the angle from the object's image to the Yc axis, which is n°. In this case, the user's view orientation can be recorded as (m°, n°). Alternatively, the distances relative to the camera's Xc, Yc, and Zc axes can be a, b, and c, respectively. In this case, the user's view orientation can be recorded as (a, b, c). Other representations can also be used to define a user's view orientation relative to the camera coordinate system; this embodiment does not limit this method.
[0075] To match the location of a user's voice with their visual location in a video recorded by an electronic device, one approach is to capture sound signals using the device's microphone and convert them into electrical audio signals. The electronic device then focuses this audio signal onto a desired area (e.g., the area where the user is speaking) and reproduces it as sound. This ensures that the location of the user's voice in the recorded video matches their visual location.
[0076] Among them, sound image location refers to the location of the sound when the sound is transmitted to the electronic device through the sound signal, the electronic device collects the sound signal, converts the sound signal into an audio signal, and then reproduces the sound corresponding to the audio signal.
[0077] However, when this approach is adopted, if the user is far away from the electronic device, the electronic device will collect other sound signals in the environment while collecting the user's voice signal. Furthermore, the user's voice signal is easily affected by the attenuation of distance when it travels through the air, resulting in the user's voice often being unclear in the recorded video.
[0078] To address the issue of unclear user audio in recorded videos when the user is far from the electronic device, as described in the previous solution, another approach is for the user to wear TWS (True Wireless Stereo) earbuds. These earbuds then capture the user's audio signal. Because the microphone of the TWS earbuds is very close to the user, they can capture the user's audio signal at close range while isolating some ambient noise. The TWS earbuds can then perform noise reduction processing on the captured ambient noise, removing it and retaining only the user's audio signal. This audio signal is then transmitted wirelessly to the electronic device, which processes the received audio signal and the recorded video to obtain the video. In this video, the user's audio is clear.
[0079] However, with this approach, since the microphone of the TWS earphones always collects the user's sound signal at the user's ear, the direction of the collected sound signal remains unchanged for the microphone; it is always the direction of the user's vocalization point relative to the microphone's center. But for the camera of an electronic device, even if the direction of the user's sound signal remains unchanged, the direction of the captured image of the user can change.
[0080] like Figure 2aAs shown, at a certain moment, when the user speaks, their position is slightly to the left relative to the center of the camera on the electronic device. Assume that at this moment, in the resulting image, the user's visual orientation is slightly to the left relative to the camera center. Figure 2b As shown, at another moment, the user's position has clearly changed; at this moment, the user's direction relative to the camera center of the electronic device is shifted to the right. Assume that at this point, in the obtained image, the user's visual orientation is shifted to the right relative to the camera center. However, the user's direction relative to the microphone center of the TWS earphone remains essentially unchanged between the two moments. Therefore, the direction of the user's sound signal captured by the microphone remains essentially unchanged. Consequently, in the recorded video, the sound image orientation of the user's voice remains unchanged.
[0081] Therefore, when the sound signal captured by TWS earbuds is reproduced in a video, the location of the user's voice image remains unchanged, but the user's visual location in the video has changed. Thus, in the resulting video, the location of the user's voice image does not match the user's visual location.
[0082] When using the video recording method provided in this application embodiment, when the user is far away from the electronic device and uses the electronic device to record video, the sound image location of the user's voice in the recorded video matches the user's visual location, and the user's voice is clear in the video.
[0083] In this embodiment, a user wears TWS earphones and records a video including themselves when the user is far from the phone. This scenario can be referred to the above. Figure 2a and Figure 2b The difference lies in the method used in this application embodiment. While the direction of the user's sound signal collected by the microphone remains essentially unchanged during a change in the user's position, after the TWS earphones transmit the sound signal to the electronic device, the electronic device can obtain the sonic image location of the user's voice corresponding to the sound signal by comparing the sound signal with the user's visual orientation in the image. This sonic image location matches the user's visual orientation, and when the user's voice is reproduced using this sonic image location, the user's voice can match the user's visual orientation. For example, Figure 2a If the user's visual orientation is slightly to the left of the camera's center, then the user's audio image orientation is also slightly to the left relative to the camera. Figure 2b If the user's visual orientation is slightly to the right of the camera's center, then the user's audio image orientation relative to the camera is also slightly to the right.
[0084] In this way, the sound image location of the user's voice matches the user's visual location in the recorded video, and when the user's voice signal is collected using TWS earphones, the user's voice is clear in the video.
[0085] The communication system 100 used in the embodiments of this application will be introduced first below.
[0086] Figure 3 This is a schematic diagram of the structure of the communication system 100 provided in the embodiments of this application.
[0087] like Figure 3 As shown, the communication system 100 includes multiple electronic devices, such as a first electronic device, a second electronic device, and a third electronic device.
[0088] The first electronic device in this application embodiment may be a terminal device running Android, Huawei HarmonyOS, iOS, Microsoft or other operating systems, such as a smart screen, mobile phone, tablet computer, laptop computer, personal computer, etc.
[0089] The second and third electronic devices can collect sound signals and convert them into electrical signals, which are then transmitted to the first electronic device. For example, the second and third electronic devices can be TWS earphones, Bluetooth earphones, etc.
[0090] The wireless network is used to provide various services for the electronic devices involved in the embodiments of this application, such as communication services, connection services, transmission services, etc.
[0091] Wireless networks include Bluetooth (BT), wireless local area network (WLAN) technology, and wireless wide area network (WWAN) technology.
[0092] The first electronic device can establish a connection with the second and third electronic devices via a wireless network, and then transmit data.
[0093] For example, the first electronic device can search for the second electronic device. When the second electronic device is found, the first electronic device can send a connection request to the second electronic device. After receiving the request, the second electronic device can establish a connection with the first electronic device. At this time, the second electronic device can convert the collected sound signal into an electrical signal, that is, an audio signal, and transmit it to the first electronic device through a wireless network.
[0094] The first electronic device involved in the communication system 100 is described below.
[0095] Figure 4 This is a schematic diagram of the structure of the first electronic device provided in the embodiments of this application.
[0096] The following description uses a first electronic device as an example to illustrate the embodiments. It should be understood that the first electronic device may have more or fewer components than shown in the figures, may combine two or more components, or may have different component configurations. The various components shown in the figures can be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits.
[0097] The first electronic device may include: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0098] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0099] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0100] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0101] The I2C interface is a bidirectional synchronous serial bus that includes a serial data line (SDA) and a serial clock line (SCL).
[0102] The I2S interface can be used for audio communication. In some embodiments, the processor 110 may include multiple I2S buses. The processor 110 can be coupled to the audio module 170 via the I2S bus to realize communication between the processor 110 and the audio module 170.
[0103] The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled via the PCM bus interface.
[0104] The UART interface is a general-purpose serial data bus used for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial and parallel communication.
[0105] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display 194 and the camera 193.
[0106] The GPIO interface can be configured via software. The GPIO interface can be configured as either control signals or data signals.
[0107] The SIM interface can be used to communicate with the SIM card interface 195 to transmit data to or read data from the SIM card.
[0108] USB port 130 is a USB standard compliant interface, specifically a Mini USB port, Micro USB port, USB Type-C port, etc. USB port 130 can be used to connect a charger to charge a primary electronic device, or to transfer data between the primary electronic device and peripheral devices. It can also be used to connect headphones for audio playback. This interface can also be used to connect other electronic devices, such as AR devices.
[0109] It is understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are merely illustrative and do not constitute a structural limitation on the first electronic device. In other embodiments of this application, the first electronic device may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0110] The charging management module 140 is used to receive charging input from the charger. The charger can be a wireless charger or a wired charger.
[0111] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to power the processor 110, internal memory 121, external memory, display 194, camera 193, and wireless communication module 160, etc.
[0112] The wireless communication function of the first electronic device can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor, and baseband processor.
[0113] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the first electronic device can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with a tuning switch.
[0114] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to a first electronic device. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc.
[0115] A modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal.
[0116] The wireless communication module 160 can provide solutions for wireless communication applications in the first electronic device, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), and other wireless communication methods. The wireless communication module 160 can be one or more devices integrating at least one communication processing module.
[0117] In some embodiments, antenna 1 of the first electronic device is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, enabling the first electronic device to communicate with networks and other devices via wireless communication technology.
[0118] The first electronic device implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0119] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel.
[0120] The first electronic device can achieve shooting functions through an ISP, camera 193, video codec, GPU, display 194, and application processor.
[0121] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.
[0122] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the first electronic device may include one or N cameras 193, where N is a positive integer greater than 1.
[0123] A digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals.
[0124] Video codecs are used to compress or decompress digital video. A first electronic device may support one or more video codecs. Thus, the first electronic device can play or record video in various encoded formats, such as Moving Picture Experts Group (MPEG) 1, MPEG 2, MPEG 3, MPEG 4, etc.
[0125] An NPU (Neural Processing Unit) is a neural network (NN) computing processor that, by borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, rapidly processes input information and can continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.
[0126] The external storage interface 120 can be used to connect an external storage card, such as a Micro SD card, to expand the storage capacity of the first electronic device. The external storage card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external storage card.
[0127] Internal memory 121 can be used to store computer executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of the first electronic device by running the instructions stored in internal memory 121. Internal memory 121 may include a program storage area and a data storage area.
[0128] The first electronic device can implement audio functions, such as music playback and recording, through an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, and an application processor.
[0129] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.
[0130] The speaker 170A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The first electronic device can listen to music or make hands-free calls through the speaker 170A.
[0131] The receiver 170B, also known as the "earpiece," is used to convert audio electrical signals into sound signals. When the first electronic device answers a telephone call or voice message, it can listen to the voice by bringing the receiver 170B close to the ear.
[0132] Microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 170C, inputting the sound signal into microphone 170C. The first electronic device may have at least one microphone 170C. In some embodiments, the first electronic device may have two microphones 170C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, the first electronic device may also have three, four, or more microphones 170C, which can collect sound signals, reduce noise, identify the sound source, and perform directional recording, etc.
[0133] The 170D headphone jack is used to connect wired headphones. The 170D headphone jack can be a USB 130 interface or a 3.5mm Open Mobile Terminal Platform (OMTP) standard interface, a CTIA (Cellular Telecommunications Industry Association of the USA) standard interface.
[0134] Pressure sensor 180A is used to sense pressure signals and can convert the pressure signals into electrical signals. In some embodiments, pressure sensor 180A can be disposed on display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. Capacitive pressure sensors may include at least two parallel plates with conductive materials.
[0135] The gyroscope sensor 180B can be used to determine the motion attitude of the first electronic device. In some embodiments, the angular velocity of the first electronic device about three axes (i.e., the x, y, and z axes) can be determined by the gyroscope sensor 180B. The gyroscope sensor 180B can be used for image stabilization.
[0136] The barometric pressure sensor 180C is used to measure air pressure. In some embodiments, the first electronic device calculates altitude using the air pressure value measured by the barometric pressure sensor 180C to assist in positioning and navigation.
[0137] The magnetic sensor 180D includes a Hall effect sensor. The first electronic device can use the magnetic sensor 180D to detect the opening and closing of the flip cover. In some embodiments, when the first electronic device is a flip phone, the first electronic device can detect the opening and closing of the flip cover based on the magnetic sensor 180D. Then, based on the detected opening and closing state of the cover or the flip cover, features such as automatic flip unlocking can be set.
[0138] The 180E accelerometer can detect the magnitude of acceleration in various directions (typically three axes) of a first electronic device. When the first electronic device is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the posture of electronic devices, and is applicable to screen orientation switching, pedometers, and other applications.
[0139] A distance sensor 180F is used to measure distance. The first electronic device can measure distance via infrared or laser. In some embodiments, during scene capture, the first electronic device can utilize the distance sensor 180F to measure distance for rapid focusing.
[0140] The proximity light sensor 180G may include, for example, a light-emitting diode (LED) and a light detector, such as a photodiode. The LED may be an infrared LED. The first electronic device emits infrared light outward through the LED. The first electronic device uses a photodiode to detect infrared reflected light from nearby objects.
[0141] The ambient light sensor 180L is used to sense ambient light intensity. The first electronic device can adaptively adjust the brightness of the display screen 194 based on the sensed ambient light intensity. The ambient light sensor 180L can also be used to automatically adjust the white balance when taking pictures. The ambient light sensor 180L can also work with the proximity sensor 180G to detect whether the first electronic device is in a pocket to prevent accidental touches.
[0142] The fingerprint sensor 180H is used to collect fingerprints. The first electronic device can utilize the characteristics of the collected fingerprint to achieve fingerprint unlocking, accessing application locks, taking photos with fingerprints, answering calls with fingerprints, etc.
[0143] Temperature sensor 180J is used to detect temperature. In some embodiments, the first electronic device uses the temperature detected by temperature sensor 180J to execute a temperature processing strategy.
[0144] Touch sensor 180K, also known as a "touch panel". Touch sensor 180K can be set on display screen 194. Touch sensor 180K and display screen 194 together form a touch screen, also known as a "touch screen". Touch sensor 180K is used to detect touch operations on or near it.
[0145] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch-sensitive buttons. The first electronic device can receive button input and generate key signal inputs related to user settings and function control of the first electronic device.
[0146] Motor 191 can generate vibration alerts. Motor 191 can be used for incoming call vibration alerts or for touch vibration feedback.
[0147] Indicator 192 can be an indicator light, used to indicate charging status, power changes, or to indicate messages, missed calls, notifications, etc.
[0148] The SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to achieve contact and separation with the first electronic device.
[0149] In this embodiment of the application, the first electronic device may further include: a laser sensor (not shown).
[0150] Laser sensors are used to sense the vibration of an object and convert that vibration into an electrical signal. In some embodiments, the laser sensor can detect throat vibration, acquire the Doppler frequency shift signal during throat vibration, and convert the Doppler frequency shift signal into an electrical signal corresponding to the vibration frequency of the throat.
[0151] It is understood that the first electronic device may also include other devices for detecting vibrations of an object, such as vibration sensors (not shown), ultrasonic sensors (not shown), etc.
[0152] In this embodiment of the application, the processor 110 can call computer instructions stored in the internal memory 121 to cause the first electronic device to execute the video recording method in this embodiment of the application.
[0153] The second electronic device involved in the communication system 100 is described below.
[0154] Figure 5 This is a schematic diagram of the structure of the second electronic device provided in the embodiments of this application.
[0155] The following description uses a second electronic device as an example to illustrate the embodiment. It should be understood that the second electronic device may have more or fewer components than shown in the figures, may combine two or more components, or may have different component configurations. The various components shown in the figures can be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits.
[0156] In this embodiment, the second electronic device may include: a processor 151, a microphone 152, a bone conduction sensor 153, and a wireless communication processing module 154.
[0157] The processor 151 can be used to parse signals received by the wireless communication processing module 154. These signals include requests to establish a connection sent by the first electronic device. The processor 151 can also be used to generate signals to be transmitted outward by the wireless communication processing module 154, including requests to transmit audio signals to the first electronic device, etc.
[0158] The processor 151 may also include a memory for storing instructions. In some embodiments, the instructions may include instructions for enhancing noise reduction, instructions for transmitting signals, etc.
[0159] Microphone 152 utilizes the conductivity of air for sound to collect the user's voice signal, as well as some other sound signals from the surrounding environment. It then converts the sound signal into an electrical signal to obtain an audio signal. Microphone 152 is also called a "phone" or "microphone".
[0160] The second electronic device may include one or N microphones 152, where N is a positive integer greater than 1. When there are two or more microphones, different arrangements can be adopted to obtain different microphone arrays, which can be used to improve the quality of the acquired sound signal.
[0161] The bone conduction sensor 153 can collect sound signals by utilizing the sound conductivity of bones. Most sound signals in the surrounding environment are conducted through the air, while the bone conduction sensor can only collect sound signals conducted in direct contact with the bones, such as the user's voice signal, and then convert the sound signal into an electrical signal to obtain a bone conduction audio signal.
[0162] The wireless communication processing module 154 may include one or more of the Bluetooth (BT) communication processing module 154A and the WLAN communication processing module 154B, for providing services such as establishing a connection with the first electronic device and transmitting data.
[0163] The video recording method in this application embodiment will be described in detail below with reference to the hardware structure diagrams of the exemplary first electronic device and the second electronic device described above:
[0164] Figure 6 This is a signaling interaction diagram of the video recording method in the embodiments of this application.
[0165] Assume that at this point, the first electronic device and the second electronic device have already established a connection and can transmit data.
[0166] When the user is far from the electronic device and is using the first electronic device to record video, the first electronic device can continuously capture multiple frames of images while simultaneously capturing audio information from the shooting environment. At the same time, the second electronic device will also record audio signals from the shooting environment.
[0167] Steps S101-S109 describe the processing of the audio signal (including the first audio signal, the second audio signal, the bone conduction audio signal, and the user's audio signal) corresponding to the current frame image during video recording. It can be understood that by processing the audio signal corresponding to each frame image and each frame image according to the description in steps S101-S109, the video involved in the embodiments of this application can be obtained.
[0168] S101. The second electronic device acquires the second audio signal;
[0169] During video recording, the electronic device can continuously acquire a second audio signal for a period of time (the time corresponding to the playback of one frame of an image). This second audio signal may include the user's voice information as well as other sound information in the environment surrounding the second electronic device.
[0170] Understandably, the length of this time interval can vary depending on the circumstances. For example, it could be 1 / 24 of a second or 1 / 12 of a second.
[0171] S102. The second electronic device acquires bone conduction audio signals;
[0172] Optionally, in some embodiments, during video recording, the second electronic device may also use a bone conduction sensor to continuously collect the user's bone conduction audio signal over a period of time (the time corresponding to playing one frame of an image).
[0173] S103. The second electronic device processes the second audio signal to obtain the user audio signal;
[0174] In some embodiments, the second electronic device may sample, denoise, or otherwise process the second audio signal to remove other sound information in the environment, thereby enhancing the user's voice information in the second audio signal and obtaining the user's audio signal.
[0175] The user audio information obtained in the above method includes not only the user's voice information but also a small portion of other sound information from the environment surrounding the second electronic device. Therefore, optionally, in other embodiments, in order to remove other sound information from the environment surrounding the second electronic device so that the user audio signal obtained after processing the second audio signal only includes the user's voice information, the second electronic device can use bone conduction audio signals and the second audio signal for joint noise reduction processing to remove other sound information from the environment surrounding the second electronic device from the second audio signal, thus obtaining the user audio signal.
[0176] Specifically, the implementation method of joint noise reduction is the same as the existing technology's implementation method of joint noise reduction processing of multiple audio signals by the same electronic device. This embodiment provides one method of joint noise reduction processing: differential calculation is performed on the bone conduction audio signal and the second audio signal to cancel out the noise in both the bone conduction audio signal and the second audio signal, thereby achieving the effect of joint noise reduction. It should be noted that during the differential calculation process, the sound wave intensity of the two audio signals needs to be weighted to make the weighted noise intensity basically the same, achieving the maximum noise reduction. In addition, if the differential calculation weakens the normal audio signal, i.e., the non-noise signal, the differential audio signal can be amplified to obtain the user's audio signal.
[0177] S104. The first electronic device acquires images;
[0178] During video recording, the camera of the first electronic device can capture images, which may include a human image, such as the user's image.
[0179] S105. The first electronic device acquires the first audio signal;
[0180] During video recording, the microphone of the first electronic device begins to continuously collect a first audio signal for a period of time. The first audio signal may include the user's voice information as well as other sound information in the environment surrounding the first electronic device.
[0181] S106. The second electronic device sends a user audio signal to the first electronic device;
[0182] The second electronic device can send user audio signals to the first electronic device via a wireless network.
[0183] S107. The second electronic device sends a bone conduction audio signal to the first electronic device;
[0184] Optionally, the second electronic device can send bone conduction audio signals to the first electronic device via a wireless network.
[0185] S108. The first electronic device determines the acoustic image location of the user's voice corresponding to the user's audio signal;
[0186] Figure 7 A flowchart is shown showing how a first electronic device determines the acoustic image location of the user's voice corresponding to the user's audio signal.
[0187] In some embodiments, the first electronic device may first use a first audio signal to determine the location of the user's sound source. In other embodiments, to improve the accuracy of the obtained location of the user's sound source, the first electronic device may also use a bone conduction audio signal combined with the first audio signal to obtain the location of the user's sound source; a detailed description of this process can be found in step S201.
[0188] Then, the first electronic device can obtain the pixel position of the face in the image, and use the pixel position of each face, combined with the user's sound source location, to obtain the sound image location of the user's voice. A detailed description of this process can be found in steps S202-S204.
[0189] S201. The first electronic device uses the first audio signal to determine the location of the user's sound source;
[0190] In some embodiments, the user's sound source orientation can be the azimuth angle of the user's sound source relative to the center of the microphone of the first electronic device. This azimuth angle can include at least one of an azimuth angle and a pitch angle. The horizontal angle is denoted as α, and the pitch angle is denoted as β.
[0191] In other embodiments, the user's sound source orientation can be the azimuth angle of the user's sound source relative to the center of the microphone of the first electronic device.
[0192] It is understood that there may be other ways to represent the location of a user's sound source, and this application embodiment does not limit this.
[0193] Suppose that at this time, the user's sound source orientation is represented by the horizontal and vertical angles of the sound source relative to the microphone of the first electronic device, which can be denoted as θ = [α, β].
[0194] The horizontal angle α and the pitch angle β can be obtained from the first audio signal. For details on the implementation, please refer to the description of the algorithm below:
[0195] In some embodiments, the first electronic device may use the first audio signal to determine the horizontal angle α and the pitch angle β based on a high-resolution spatial spectrum estimation algorithm.
[0196] In other embodiments, the first electronic device may determine the horizontal angle α and the pitch angle β based on the beamforming of the N microphones and the first audio signal, using a beamforming algorithm with maximum output power.
[0197] It is understood that the first electronic device may also determine the horizontal angle α and the pitch angle β in other ways. This application does not limit this approach.
[0198] The following section uses a beamforming algorithm based on maximum output power to determine the horizontal angle α and the pitch angle β as an example. It details a possible implementation algorithm, and it should be understood that the algorithm does not limit this application.
[0199] The first electronic device, by comparing the output power of the first audio signal in various directions, can determine the direction of the beam with the maximum power as the azimuth of the target sound source, which is the user's sound source azimuth. The formula for obtaining the azimuth θ of this target sound source can be expressed as:
[0200]
[0201] In the formula, t represents a time frame, i.e., a processing frame for the audio signal. i represents the i-th microphone, and H... i (f, θ) represents the beam weight of the i-th microphone in beamforming, Y i (f, t) represents the audio signal in the time-frequency domain obtained from the sound information collected by the i-th microphone.
[0202] Beamforming refers to the response of N microphones to a narrowband audio signal. Since this response varies in different locations, beamforming is correlated with the location of the sound source. Therefore, beamforming can locate sound sources in real time and suppress background noise interference.
[0203] Beamforming can be represented as a 1×N matrix, denoted as H(f, θ), where N is the number of microphones. The value of the i-th element in beamforming can be represented as H... i (f, θ), this value is related to the position of the i-th microphone among the N microphones. Beamforming can be obtained using the power spectrum, which can be a capon spectrum, a barttle spectrum, etc.
[0204] For example, taking the Barttlett spectrum as an example, the first electronic device can use the Barttlett spectrum to obtain the i-th element in beamforming, which can be represented as: In the formula, j is an imaginary number. τ is the phase compensation value of the beamformer for this microphone. i This represents the time delay difference when the same sound information reaches the i-th microphone. This time delay difference is related to the location of the sound source and the position of the i-th microphone, as described below.
[0205] Establish a three-dimensional coordinate system by selecting the center of the first microphone among N microphones that can receive sound information as the origin. In this three-dimensional coordinate system, the position of the Nth microphone can be represented as P. i =[x i y i , z i ]. Then τ i The relationship between the location of the sound source and the position of the i-th microphone can be expressed by the following formula:
[0206]
[0207] Where c is the speed at which the sound signal travels.
[0208] The first audio signal includes audio signals obtained from sound information collected by N microphones, where N is a positive integer greater than 1.
[0209] The sound information collected by the i-th microphone can be converted into an audio signal in the time-frequency domain, represented as follows: Among them, s o (f, t) represents the change of time t. The microphone, which serves as the origin, collects sound information and converts it into an audio signal in the time-frequency domain.
[0210] In some embodiments, when the first audio signal is broadband information, to improve processing accuracy, the first audio signal can be divided into several narrowband audio signals by using a discrete Fourier transform (DFT). The processing results of the narrowband audio signals at each frequency point are then combined to obtain the location result of the broadband audio signal. For example, a broadband audio signal with a sampling rate of 48 kHz can be divided into 2049 narrowband audio signals using a 4096-point DFT. Then, by processing each narrowband audio signal or several of them using the above algorithm, the location of the target sound source can be determined. The formula for obtaining the location θ of the target sound source can be expressed as:
[0211]
[0212] In the formula, f represents the frequency value in the frequency domain.
[0213] Optionally, in other embodiments, since the first audio signal includes other sound information besides the user's voice information, in order to prevent other sound information from affecting the determination of the user's sound source location, the first electronic device can use bone conduction audio signals to filter out other sound signals in the first audio signal, enhance the user's voice information in the first audio signal, and make the obtained sound source location of the user more accurate.
[0214] For example, a correlation analysis can be performed on the first audio signal by combining it with the bone conduction audio signal. Audio information corresponding to time-frequency points in the first audio signal that are strongly correlated with the bone conduction audio signal can be assigned a larger weight, while audio information corresponding to time-frequency points that are weakly correlated can be assigned a smaller weight. This results in a weight matrix w(ft), where each element can be denoted as w. mn Let w(f,t) represent the weight of the audio signal with frequency n at time m. Then, using this weight matrix w(f,t) and the algorithm described above, the formula for obtaining the azimuth θ of the target sound source using the bone conduction audio signal and the first audio signal can be expressed as:
[0215]
[0216] S202. The first electronic device obtains the pixel position of the face in the image based on the image;
[0217] A face refers to all human faces in an image that can be recognized by the first electronic device. The pixel position of a face in an image can be represented by the pixel coordinates of the face in the image pixel coordinate system.
[0218] In some implementations, a single pixel can be selected from the face, and its pixel coordinates can be used to represent the pixel position of the face. For example, the pixel coordinates of the center point of the mouth can be used as the pixel position of the face, or the pixel coordinates of the center point of the face can be used as the pixel position of the face.
[0219] The first electronic device can sample N frames of images obtained over a period of time to determine the pixel position of a face in a certain frame.
[0220] In some embodiments, the first electronic device can perform face recognition on the image to obtain the pixel coordinates of the face, which are the pixel positions of the face in the image during this time period.
[0221] The pixel positions of a face can be represented by a matrix, denoted as H. The i-th element of this matrix represents the pixel position of the i-th face in the image, which can be represented as H. i =[u i v i ].
[0222] S203. The first electronic device obtains the user's visual orientation in the image based on the location of the sound source and the pixel position of the face;
[0223] The first electronic device can determine the correlation between the pixel position of each face and the sound source location based on the sound source location and the pixel position of the face. If the correlation between the pixel position of a certain face and the sound source location is stronger, the first electronic device determines the viewing location of that face as the user's viewing location.
[0224] Specifically, the first electronic device can obtain the approximate pixel coordinates of the user in the image based on the location of the sound source. Then, from the pixel positions of the face, it determines the pixel position of the face that is closest to the approximate pixel coordinates. This pixel position of the face is used as the user's pixel coordinates in the image. Then, using the user's pixel coordinates in the image, the user's viewing orientation is obtained.
[0225] The following provides an implementation of the algorithm. It should be understood that this algorithm does not constitute a limitation on the embodiments of this application.
[0226] Suppose that the user's sound source orientation is represented by the horizontal and vertical angles of the sound source relative to the microphone of the first electronic device, which can be denoted as θ = [α, β]. An algorithm for obtaining the user's visual orientation using this sound source orientation can be found in the description below.
[0227] The sound source orientation is obtained relative to the center of the microphone. Since the distance between the center of the microphone and the center of the camera in the first electronic device is much smaller than the distance between the first electronic device and the user, the sound source orientation can be considered to be obtained relative to the center of the camera. Using the horizontal angle α and pitch angle β corresponding to this sound source orientation as the approximate horizontal angle α and pitch angle β of the user relative to the camera, the user's visual orientation can be obtained using relevant algorithms that include camera parameters.
[0228] Specifically, firstly, using the horizontal angle α and the pitch angle β, we can obtain the user's coordinates g = [x, y] in the image coordinate system. The formula for x in g can be expressed as: x = f tanα, where f is the focal length of the camera. The formula for y in g can be expressed as: y = f cosαtanβ.
[0229] Then, the coordinates g = [x, y] in the image coordinate system are converted to pixel coordinates h = [u, v] in the image. These pixel coordinates h = [u, v] represent the approximate pixel coordinates of the user in the image. The conversion formula can be expressed as:
[0230]
[0231] In the formula, u0 and v0 are the pixel coordinates (u0 and v0) of the origin in the image plane coordinate system in the image pixel coordinate system, dx is the length of one pixel in the U-axis direction in the image, and dy is the length of one pixel in the V-axis direction in the image.
[0232] Then, using this approximate pixel coordinate h = [u, v], the pixel position H of the face is determined. i =[u i v i The system performs matching to obtain the correlation between each face and the approximate pixel coordinates h = [u, v]. For example, the smaller the distance between the face's pixel position and the approximate pixel coordinates, the stronger the correlation; therefore, the stronger the correlation between the face's pixel position and the location of the sound source. H... i =[u i v i The pixel position of the face that is most strongly correlated with h = [u, v] is the user's pixel coordinates in the image, which can be represented as h′ = [u′, v′].
[0233] Finally, the user's pixel coordinates in the image are converted into horizontal angles α′ and pitch angles β′ relative to the camera, which are taken as the user's viewing orientation θ′ = [α′, β′]. This conversion process is the reverse of the process described above, which uses horizontal angle α and pitch angle β to obtain the user's approximate pixel coordinates h = [u, v] in the image, and will not be repeated here.
[0234] It is understood that when the representation of the sound source location is different, the conversion method may be different. The embodiments of this application do not limit the algorithm in step S203.
[0235] In some embodiments, in order to further improve the accuracy of the first electronic device in obtaining the user's visual orientation, the correlation between each face obtained in steps S201-S203 and the orientation of the sound source can be used as the first determining factor.
[0236] A second determining factor is added to the first determining factor to jointly determine the user in the face. The second determining factor is the correlation between the first feature of each face and the first feature in the user's bone conduction audio signal.
[0237] Specifically, the first electronic device can use bone conduction audio signals to perform first feature extraction on the user's voice signal to obtain a first feature in the user's voice signal. Furthermore, it can perform the same first feature extraction on faces in an image to obtain a first feature for each face in the image.
[0238] The first feature of a face can be obtained in different ways.
[0239] For example, in some embodiments, the first electronic device can use an image to extract the first feature from a face, and the first electronic device can use this first feature as the first feature of the face. In this case, the first feature may include voice activity detection (VAD) features, phoneme features, etc.
[0240] In other embodiments, the first electronic device may extract a first feature using ultrasonic or laser echo signals generated by throat vibrations during facial vocalization. In this case, the first feature may be a pitch feature or the like. The first electronic device may use this first feature as the first feature of the face.
[0241] The ultrasonic echo signal generated by the vibration of the throat when the selected person speaks can be collected using a vibration sensor or an ultrasonic sensor in the first electronic device. The laser echo signal can be collected using a laser sensor in the first electronic device.
[0242] The first electronic device can perform correlation analysis between the first feature in the user's voice signal and the first feature of the face in the image to obtain the correlation between the first feature of each face and the first feature in the user's bone conduction audio signal, and use this correlation as the second determining factor.
[0243] Since the distance between the center of the microphone and the center of the camera is much smaller than the distance between the first electronic device and the user, the first electronic device can determine the user's visual orientation as the sound image orientation of the user's voice.
[0244] S109. The first electronic device uses the sound image location of the user's voice and the user's audio signal to obtain a stereo audio signal.
[0245] In some embodiments, the stereo audio signal includes only the user's stereo audio signal but excludes the ambient stereo audio signal; in this case, the user's stereo audio signal is the stereo audio signal. The user's stereo audio signal includes the user's voice information. The user's stereo audio signal can be used to reproduce the user's voice, and in the reproduced user's voice, the sound image location of the user's voice matches the user's visual location.
[0246] Among them, the user stereo audio signal refers to the dual-channel user audio signal.
[0247] The aforementioned user audio signal is a single-channel audio signal, and the user's voice restored from it is not stereo. In order to obtain a stereo user audio signal and more realistically restore the user's voice, this single-channel audio signal can be converted into a dual-channel user audio signal. The user's voice restored from this dual-channel user audio signal will then be stereo.
[0248] The first electronic device can utilize the sound image location of the user's voice and combine it with the user's audio signal to obtain a user stereo audio signal corresponding to the sound image location of the user's voice.
[0249] Specifically, the first electronic device convolves the user audio signal with the head-related impulse response (HRIR) corresponding to the user's voice image location to recover the inter-aural level difference (ILD), inter-aural time difference (ITD), and spectral clues. This allows the mono audio signal to be converted into a stereo audio signal, which may include a left channel and a right channel. The ILD, ITD, and spectral clues are used to enable the stereo audio signal of the stereo signal to determine the user's voice image location.
[0250] Understandably, in addition to the HRIR algorithm, the first electronic device can also use other algorithms to obtain the dual-channel user audio signal, such as binaural room impulse response (BRIR).
[0251] In some embodiments, the stereo audio information in the recorded video may include, in addition to the user's stereo audio signal, an ambient stereo audio signal. This ambient stereo audio signal can be used to reproduce sounds in the shooting environment other than the user's voice. This ambient stereo audio signal refers to a dual-channel ambient audio signal, which is an electrical signal converted from other sound signals in the environment.
[0252] Generally speaking, since the second electronic device filters out most of the other sound information in the environment when collecting the sound signal, the other sound information in the environment included in the first audio signal is clearer than that in the second signal. Therefore, the first electronic device can use the first audio signal to obtain the audio signals of other sounds in the environment, which is the environmental audio signal. Then, using this environmental audio signal, an environmental stereo audio signal is obtained.
[0253] Specifically, in some embodiments, firstly, the first electronic device can perform adaptive blocking filtering on the first audio signal and the user audio signal to filter out the user's voice information in the first audio signal and obtain an ambient audio signal. Then, a two-channel ambient audio signal is obtained through beamforming of the first electronic device. This two-channel ambient audio signal can be in X / Y, M / S, or A / B format.
[0254] In other embodiments, the first electronic device may also acquire the ambient stereo audio signal in other ways. For example, it may first obtain a stereo first audio signal using a first audio signal. This stereo first audio signal is a two-channel first audio signal. Then, it may perform adaptive blocking filtering on the stereo first audio signal and the user audio signal to filter out the user's voice information in the stereo first audio signal, thereby obtaining the ambient stereo audio signal. This application does not limit the scope of this embodiment.
[0255] Then, the ambient stereo audio signal is mixed with the user stereo audio signal to obtain a stereo audio signal that includes both the user's voice information and other sound information from the environment.
[0256] During the mixing of the ambient stereo audio signal and the user stereo audio signal, the first electronic device can also perform audio zoom on the user stereo audio signal to match the size of the user's sound image with the distance between the user and the first electronic device when the user speaks. The size of the user's sound image refers to the volume of the user's voice in the video.
[0257] The first electronic device can determine the volume of the user's stereo audio signal in the video based on the focus information provided by the user.
[0258] Specifically, in some embodiments, when the first electronic device is recording video, in response to the user increasing the shooting focal length, the first electronic device determines that the user has moved closer to it. At this point, the first electronic device can amplify the user's stereo audio signal while leaving the ambient stereo audio signal unchanged or suppressing it. Then, the user's audio signal is mixed with the ambient stereo audio signal to obtain a stereo audio signal. In the video, the volume of the user's voice in this stereo audio signal will be louder, while the volume of other ambient sounds will be relatively lower.
[0259] In this process, the suppression of the ambient stereo audio signal by the first electronic device can reduce the noise of other sounds around the first electronic device that reproduces the ambient stereo audio signal.
[0260] In some embodiments, when the first electronic device is recording video, in response to the user's action of reducing the shooting focal length, the first electronic device determines that the user has moved further away from it. At this point, the first electronic device can suppress the user's stereo audio signal while leaving the ambient stereo audio signal unchanged or enhancing it. Then, the user's audio signal is mixed with the ambient stereo audio signal to obtain a stereo audio signal. In the video, the volume of the user's voice in this stereo audio signal will be lower, while the volume of other ambient sounds will be relatively higher.
[0261] In other embodiments, in addition to using the shooting focal length set by the user, the first electronic device can also mix the ambient stereo audio signal and the user stereo audio signal in other ways according to a certain volume ratio. For example, a default mixing ratio can be set, which is not limited in this application embodiment.
[0262] It should be understood that the steps S101-S105 above are not in any particular order, as long as step S103 follows step S101. Steps S106 and S107 are also not in any particular order and can be performed simultaneously. That is, in some embodiments, the second electronic device can encode the user's audio signal and the bone conduction audio signal together and send them to the first electronic device simultaneously.
[0263] It is understandable that when a user uses the second electronic device, the first electronic device can repeatedly obtain the user's stereo audio signal corresponding to multiple frames of images according to steps S101-S109, encode the stereo audio signal corresponding to the multiple frames of images to obtain an audio stream. Simultaneously, the multiple frames of images are encoded to obtain a video stream. Then, the audio stream and video stream are mixed to obtain the recorded video. The electronic device can process the multiple frames of images in the video to obtain multiple frames of images.
[0264] When the first electronic device plays the video, it can display one frame of image at a certain moment, and this image will remain on the screen for a period of time. During this period, the first electronic device can play the stereo audio signal corresponding to that frame of image. The stereo audio signal can reproduce the user's voice, and at this time, the sound image location of the user's voice matches the user's visual location. After the current frame of image finishes playing, the next frame of image can be played, and the stereo audio signal corresponding to the next frame of image can be played simultaneously, until the video finishes playing.
[0265] For example, at a first moment, the first electronic device can play the first frame of an image and simultaneously begin playing the stereo audio signal corresponding to that first frame. This stereo audio signal can reconstruct the user's voice. At this time, the user's viewing position in the first frame is to the left, so the sound image of the user's voice reconstructed by the stereo audio signal corresponding to that frame is also to the left. The first frame can remain on the electronic device's display screen for a period of time, during which the first electronic device can continuously play the stereo audio signal corresponding to that frame. Then, at a second moment, the first electronic device can play the second frame of an image and simultaneously begin playing the stereo audio signal corresponding to that second frame. This stereo audio signal can reconstruct the user's voice. At this time, the user's viewing position in the second frame is to the right, so the sound image of the user's voice reconstructed by the stereo audio signal corresponding to that second frame is also to the right.
[0266] The sound image location of the user's voice in the video matches the user's visual location. The user's audio signal in the stereo audio signal in the video is collected by a second electronic device. This stereo user audio signal can reproduce the user's voice, and the user's voice is clear.
[0267] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
[0268] As used in the above embodiments, depending on the context, the term "when..." can be interpreted as meaning "if...", "after...", "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if (the stated condition or event) is interpreted as meaning "if determining...", "in response to determining...", "when (the stated condition or event) is detected", or "in response to detecting (the stated condition or event)".
[0269] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.
[0270] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A video recording method, characterized in that, include: During the recording of video by the first electronic device, the first electronic device acquires images and first audio signals; The first electronic device determines the acoustic location of the user's voice in the image based on the image and the first audio signal; The first electronic device generates a stereo audio signal based on the sound image location and the user audio signal; the user audio signal is acquired by the second electronic device and sent to the first electronic device. The first electronic device generates video based on the image and the stereo audio signal; wherein, Determining the acoustic location of the user's voice in the image includes: The first electronic device obtains the predicted pixel position of the user in the image based on the user's sound source orientation; the user's sound source orientation is determined based on the first audio signal and the bone conduction audio signal, and the determination process includes: combining the bone conduction audio signal, performing correlation analysis on the first audio signal, enhancing the audio signals in the first audio signal that are strongly correlated with the bone conduction audio signal, and using the enhanced first audio signal to determine the sound source orientation; the bone conduction audio signal is acquired by the second electronic device and sent to the first electronic device; The first electron determines the pixel position of the face closest to the predicted pixel position from the pixel positions of at least one face, and uses this as the user's pixel position; the at least one face is a face in the image; The first electronic device determines the user's visual orientation based on the user's pixel position, and uses the visual orientation as the audio-visual orientation.
2. The method according to claim 1, characterized in that, The first electronic device generates a stereo audio signal based on the sound image location and the user audio signal, specifically including: The first electronic device generates an ambient stereo audio signal; The first electronic device generates the stereo audio signal based on the sound image location of the user's voice, the user's audio signal, and the ambient stereo audio signal.
3. The method according to claim 2, characterized in that, The first electronic device generates an ambient stereo audio signal, specifically including: The first electronic device performs adaptive blocking filtering on the first audio signal based on the user audio signal to filter out user voice information in the first audio signal; The first electronic device generates the ambient stereo audio signal based on the filtered first audio signal.
4. The method according to claim 2, characterized in that, The first electronic device generates an ambient stereo audio signal, specifically including: The first electronic device uses the first audio signal to generate a stereo first audio signal; The first electronic device performs adaptive blocking filtering on the stereo first audio signal based on the user audio signal to filter out the user voice information in the stereo first audio signal and obtain an ambient stereo audio signal.
5. The method according to claim 2, characterized in that, The first electronic device generates the stereo audio signal based on the sound image location of the user's voice, the user's audio signal, and the ambient stereo audio signal, specifically including: The first electronic device generates a user stereo audio signal based on the sound image location of the user's voice and the user's audio signal; The first electronic device enhances the user's stereo audio signal without altering the ambient stereo audio signal. The first electronic device generates the stereo audio signal based on the enhanced user stereo audio signal and the ambient stereo audio signal.
6. The method according to claim 2, characterized in that, The first electronic device generates the stereo audio signal based on the sound image location of the user's voice, the user's audio signal, and the ambient stereo audio signal, specifically including: The first electronic device generates a user stereo audio signal based on the sound image location of the user's voice and the user's audio signal; The first electronic device enhances the user stereo audio signal while suppressing the ambient stereo audio signal; The first electronic device generates the stereo audio signal based on the enhanced user stereo audio signal and the suppressed ambient stereo audio signal.
7. The method according to any one of claims 1-6, characterized in that, The user audio signal is obtained by the second electronic device performing joint noise reduction processing on the second audio signal based on the bone conduction audio signal, removing other sound information from the environment surrounding the second electronic device in the second audio signal.
8. The method according to any one of claims 1-6, characterized in that, The user audio signal is obtained by the second electronic device performing noise reduction processing on the second audio signal to remove other sound information from the environment surrounding the second electronic device.
9. An electronic device, characterized in that, The electronic device includes: one or more processors and memory; The memory is coupled to the one or more processors, and the memory is used to store computer program code, the computer program code including computer instructions, which the one or more processors invoke to cause the electronic device to execute: During video recording, images and the first audio signal are captured; Based on the image and the first audio signal, determine the acoustic image location of the user's voice in the image; A stereo audio signal is generated based on the sound image location of the user's voice and the user's audio signal; the user's audio signal is acquired by a second electronic device and sent to the electronic device. A video is generated based on the image and the stereo audio signal; wherein... Determining the acoustic location of the user's voice in the image includes: Based on the user's sound source location, the predicted pixel position of the user in the image is obtained; the user's sound source location is determined based on the first audio signal and the bone conduction audio signal, and the determination process includes: combining the bone conduction audio signal, performing correlation analysis on the first audio signal, enhancing the audio signals in the first audio signal that are strongly correlated with the bone conduction audio signal, and using the enhanced first audio signal to determine the sound source location; the bone conduction audio signal is acquired by the second electronic device and sent to the electronic device; From the pixel positions of at least one face, determine the pixel position of the face that is closest to the predicted pixel position, and use it as the user's pixel position; the at least one face is a face in the image; Based on the user's pixel position, the user's visual orientation is determined, and the visual orientation is used as the audio-visual orientation.
10. The electronic device according to claim 9, characterized in that, The one or more processors are specifically used to invoke the computer instructions to cause the electronic device to execute: Generate ambient stereo audio signals; The stereo audio signal is generated based on the sound image location of the user's voice, the user's audio signal, and the ambient stereo audio signal.
11. The electronic device according to claim 10, characterized in that, The one or more processors are specifically used to invoke the computer instructions to cause the electronic device to execute: The first audio signal is subjected to adaptive blocking filtering based on the user audio signal to filter out the user voice information in the first audio signal. The ambient stereo audio signal is generated based on the filtered first audio signal.
12. The electronic device according to claim 10, characterized in that, The one or more processors are specifically used to invoke the computer instructions to cause the electronic device to execute: Using the first audio signal, a stereo first audio signal is generated; Based on the user audio signal, the stereo first audio signal is subjected to adaptive blocking filtering to remove the user voice information from the stereo first audio signal, thereby obtaining the ambient stereo audio signal.
13. The electronic device according to claim 10, characterized in that, The one or more processors are specifically used to invoke the computer instructions to cause the electronic device to execute: Generate a user stereo audio signal based on the sound image location of the user's voice and the user's audio signal; The user stereo audio signal is enhanced without altering the ambient stereo audio signal. The stereo audio signal is generated based on the enhanced user stereo audio signal and the ambient stereo audio signal.
14. The electronic device according to claim 10, characterized in that, The one or more processors are specifically used to invoke the computer instructions to cause the electronic device to execute: Generate a user stereo audio signal based on the sound image location of the user's voice and the user's audio signal; The user stereo audio signal is enhanced while the ambient stereo audio signal is suppressed. The stereo audio signal is generated based on the enhanced user stereo audio signal and the suppressed ambient stereo audio signal.
15. A chip system applied to an electronic device, the chip system comprising one or more processors, the processors being configured to invoke computer instructions to cause the electronic device to perform the method as described in any one of claims 1-8.
16. A computer program product containing instructions, characterized in that, When the computer program product is run on an electronic device, it causes the electronic device to perform the method as described in any one of claims 1-8.
17. A computer-readable storage medium comprising instructions, characterized in that, When the instructions are executed on an electronic device, the electronic device causes the electronic device to perform the method as described in any one of claims 1-8.
Citation Information
Patent Citations
Video processing method and electronic equipment
CN110572600A
Sound effect adjusting method, device and equipment and storage medium
CN112492380A