Audio recording method, head-mounted device and storage medium

By obtaining the focal length information of the head-mounted device camera and controlling the pickup range and sensitivity of the microphone array, the problem of poor audio recording effect when shooting at a long distance is solved, and the synchronous recording of audio and video is achieved, thereby improving the recording effect.

CN120676113APending Publication Date: 2025-09-19GOERTEK INC
View PDF 25 Cites 0 Cited by

Patent Information

Application Number
CN202511122568.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing recording technology has difficulty in clearly picking up the sound at the corresponding shooting distance when the shooting distance becomes farther. The audio recording effect cannot change with the video recording distance, which affects the video shooting effect.

Method used

By obtaining the focal length information of the camera in the head-mounted device, the target parameter value is determined to control the pickup range and sensitivity of the microphone array, so that the audio recording method can zoom with the video and achieve synchronous recording of audio and video.

Benefits of technology

It achieves the matching of audio recording and video shooting scenes, improves the recording effect, ensures the synchronization of video and audio, and enhances the video shooting quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120676113A_ABST
    Figure CN120676113A_ABST
Patent Text Reader

Abstract

The invention discloses an audio recording method, head-mounted equipment and a storage medium, and relates to the technical field of head-mounted equipment. The audio recording method is applied to the head-mounted device, a microphone array and a camera are arranged in the head-mounted device, and the audio recording method comprises the following steps: acquiring focal length information of a video currently shot by the camera; a target parameter value is determined according to the focal length information, the target parameter value comprises a first parameter value used for controlling a pickup range formed by a microphone array beam, and the larger the focal length represented by the focal length information is, the smaller the span of the pickup range corresponding to the first parameter value is; and processing a signal collected by the microphone array according to the target parameter value to obtain a recorded audio. By determining the pickup range corresponding to the current focal length, zooming of the audio along with the video is realized, so that the audio is collected based on the determined pickup range, synchronization of the recorded video and the audio is ensured, and the video shooting effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of head-mounted devices, and in particular to an audio recording method, a head-mounted device, and a storage medium. Background Art

[0002] With the development of long-distance camera technology, significant progress has been made. For example, on mobile phones, the integration of optical lenses and AI algorithms has made it possible to photograph the moon. In addition, the focal length of mobile phone camera modules has continued to increase, now capable of shooting at a focal length of 200mm, and even reaching a focal length of 400mm with external accessories, allowing clear capture of objects 50 meters away. However, existing audio recording technology is relatively backward. When the distance of the camera image increases, it is difficult to clearly pick up the sound at the corresponding shooting distance. The audio recording effect does not change with the video recording distance, which affects the video shooting effect.

[0003] The above content is only used to assist in understanding the technical solution of this application and does not constitute an admission that the above content is prior art. Summary of the Invention

[0004] The main purpose of this application is to provide an audio recording method, a head-mounted device and a storage medium, aiming to solve the problem of how to achieve audio zoom with video.

[0005] To achieve the above objectives, the present application proposes an audio recording method, which is applied to a head-mounted device provided with a microphone array and a camera, and comprises: Obtain the focal length information of the video currently being shot by the camera; Determining a target parameter value according to the focal length information, wherein the target parameter value includes a first parameter value for controlling a sound pickup range of a microphone array beamforming, and a larger focal length represented by the focal length information indicates a smaller span of the sound pickup range corresponding to the first parameter value; The signal collected by the microphone array is processed according to the target parameter value to obtain recorded audio.

[0006] Optionally, the target parameter value further includes a second parameter value for controlling microphone sensitivity, and the larger the focal length represented by the focal length information, the higher the microphone sensitivity corresponding to the second parameter value.

[0007] Optionally, before the step of processing the signal collected by the microphone array according to the target parameter value to obtain recorded audio, the method further includes: determining the number of targets according to the focal length information, wherein the greater the focal length represented by the focal length information, the greater the number of targets; Controlling the target number of microphones in the microphone array to collect signals.

[0008] Optionally, the step of controlling the target number of microphones in the microphone array to collect signals includes: Get the picture currently captured by the camera; Identifying the object in the picture to obtain a target position of the object in the picture; The target number of microphones is selected from the microphone array according to the target position, wherein the distance between the selected microphones and the object is smaller than the distance between the unselected microphones and the object.

[0009] Optionally, the step of determining a target parameter value according to the focal length information includes: Determining a target focal length segment in which the focal length represented by the focal length information is located from each preset focal length segment; The parameter value corresponding to the target focal length segment in each preset parameter value is determined as the target parameter value, wherein the preset parameter value includes a parameter value for controlling the pickup range of the microphone array, and the pickup range corresponding to the preset focal length segment with a larger focal length has a smaller span.

[0010] Optionally, the sound pickup range corresponding to the first parameter value covers the shooting direction of the camera.

[0011] Optionally, the step of determining a target parameter value according to the focal length information includes: receiving a user adjustment instruction, and determining an adjustment factor according to the user adjustment instruction; The target parameter value is determined according to the adjustment factor and the focal length information, wherein, when the focal length represented by the focal length information remains unchanged, the larger the adjustment factor is, the smaller the span of the pickup range corresponding to the first parameter value is.

[0012] In addition, to achieve the above-mentioned purpose, the present application also proposes an audio recording device, which includes: A focal length determination module, used to obtain focal length information of the video currently being shot by the camera; a parameter confirmation module, configured to determine a target parameter value based on the focal length information, wherein the target parameter value includes a first parameter value for controlling a sound pickup range of a microphone array beamforming, and the larger the focal length represented by the focal length information, the smaller the span of the sound pickup range corresponding to the first parameter value; A signal processing module is used to process the signal collected by the microphone array according to the target parameter value to obtain recorded audio.

[0013] In addition, to achieve the above-mentioned purpose, the present application also proposes a head-mounted device, which is provided with a microphone array and a camera. The head-mounted device also includes: a memory, a processor, and a computer program stored in the memory and runnable on the processor, and the computer program is configured to implement the steps of the audio recording method described above.

[0014] Optionally, the microphone array includes at least microphones respectively arranged on the left and right sides of the head-mounted device.

[0015] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium. A computer program is stored on the storage medium, and when the computer program is executed by the processor, the steps of the audio recording method described above are implemented.

[0016] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps of the audio recording method described above.

[0017] One or more technical solutions proposed in this application have at least the following technical effects: The target parameter value is determined by obtaining the focal length information of the video currently being captured by the camera in the head-mounted device, wherein the target parameter value includes a first parameter value for controlling the pickup range of the microphone array beamforming. The larger the focal length represented by the focal length information, the smaller the span of the pickup range corresponding to the first parameter value. Since the smaller the span of the pickup range, the farther the sound can be picked up, the larger the focal length, the farther the shooting distance, the smaller the span of the corresponding pickup range, and the farther the sound can be picked up, thereby achieving audio zoom with video, and the video focal length can correspond to the pickup range. In addition, based on the target parameter value, the signal collected by the microphone array is processed to obtain recorded audio, which can achieve audio recording within the corresponding pickup range under the current focal length information, so that the audio recording matches the video shooting scene and improves the recording effect. Compared with the video shooting scheme under the existing recording technology, the present application achieves audio zoom with video by determining the pickup range corresponding to the current focal length, thereby capturing audio based on the determined pickup range, ensuring the synchronization of the recorded video and audio, and improving the video shooting effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0019] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0020] Figure 1 A flowchart of the first embodiment of the audio recording method of the present application is provided; Figure 2 This is a schematic diagram of the sound pickup range involved in an embodiment of the audio recording method of the present application; Figure 3 This is a schematic diagram of a head-mounted device involved in an embodiment of the audio recording method of the present application; Figure 4 A flowchart of the second embodiment of the audio recording method of the present application is provided; Figure 5 A flowchart of the third embodiment of the audio recording method of the present application is provided; Figure 6 This is a schematic diagram of the module structure of the audio recording device of this application; Figure 7 This is a schematic diagram of the device structure of the hardware operating environment involved in the audio recording method in the embodiment of the present application.

[0021] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0022] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.

[0023] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0024] Existing recording technology is relatively backward. When the distance of the photographic screen becomes farther, it is difficult to clearly pick up the sound at the corresponding photographic distance. The audio recording effect cannot change with the video recording distance, which affects the video shooting effect.

[0025] The present invention provides a solution by obtaining the focal length information of the video currently being captured by the camera in the head-mounted device to determine a target parameter value, wherein the target parameter value includes a first parameter value for controlling the pickup range of the microphone array beamforming. The larger the focal length represented by the focal length information, the smaller the span of the pickup range corresponding to the first parameter value. Since the smaller the span of the pickup range, the farther the sound can be picked up, the larger the focal length, the longer the shooting distance, the smaller the span of the corresponding pickup range, and the farther the sound can be picked up, thereby achieving audio zoom with video, and the video focal length can be matched with the pickup range. In addition, based on the target parameter value, the signal collected by the microphone array is processed to obtain recorded audio, which can achieve audio recording within the pickup range corresponding to the current focal length information, matching the audio recording with the video shooting scene, and improving the recording effect. Compared with video shooting solutions under existing recording technologies, the present invention achieves audio zoom with video by determining the pickup range corresponding to the current focal length, thereby capturing audio based on the determined pickup range, ensuring the synchronization of the recorded video and audio, and improving the video shooting effect.

[0026] It should be noted that the executor of each embodiment of the audio recording method of the present application is a head-mounted device. It is understandable that there are many types of head-mounted devices at present, such as smart glasses, VR (Virtual Reality) / AR (Augmented Reality) devices, etc., and there are many ways to implement the hardware architecture and software system of each head-mounted device. The embodiments of the present application do not limit the type of head-mounted device, the hardware architecture and the implementation of the software system used by the audio recording method. In this embodiment, a microphone array and a camera are provided in the head-mounted device. The microphone array is a set composed of at least two microphones arranged in a certain manner. In this embodiment, the number of microphones contained in the microphone array in the head-mounted device is not limited, nor is the arrangement of the microphone array in the head-mounted device limited. In this embodiment, the number and setting positions of the cameras in the head-mounted device are also not limited. Reference Figure 1 , Figure 1 This is a flowchart of the first embodiment of the audio recording method of this application. In this embodiment, the audio recording method includes steps S10 to S30: Step S10: Obtain the focal length information of the video currently shot by the camera.

[0027] A camera is an imaging device that includes optics and an image sensor and is used to capture video. Focal length is the distance from the optical center of the lens to the imaging plane, which determines the viewing angle and sharpness range of the camera's image. A larger focal length results in a narrower viewing angle, and the objects that can be clearly captured are farther away and have a smaller range. A smaller focal length results in a wider viewing angle, resulting in a larger range but more blurred distant objects. Focal length information is information that characterizes the focal length. This can be the focal length parameter value (in mm) itself, or other information that can reflect the focal length, such as focal level or focus mode. This is not a limitation in this embodiment.

[0028] Optionally, focal length information can be obtained by detecting the focal length scale of the camera lens when it is currently recording video. The focal length scale is the scale line on the lens that indicates the focal length and focus distance. The focal length information can be determined by checking the camera's corresponding API (Application Programming Interface).

[0029] Optionally, when obtaining focal length information, a focal length mode selection instruction can also be received through the user interaction interface to determine the focal length mode selected when currently shooting the video. The focal length mode can be, for example, macro mode, standard mode, or telephoto mode, etc. It can be understood that different focal length modes use different focal lengths, so the focal length mode is information that can characterize the focal length size.

[0030] It is understandable that by obtaining the camera focal length information, the distance of the current shooting target can be determined, providing a basis for subsequent adjustment of the pickup range.

[0031] Step S20: determining a target parameter value based on the focal length information, wherein the target parameter value includes a first parameter value for controlling the pickup range of the microphone array beamforming, and the larger the focal length represented by the focal length information, the smaller the span of the pickup range corresponding to the first parameter value.

[0032] The parameter values ​​required for processing the signals collected by the microphone array (hereinafter referred to as target parameter values ​​for clarity) can be determined based on the current focal length information. In this embodiment, the target parameter values ​​include parameter values ​​for controlling the pickup range of the microphone array beamforming (hereinafter referred to as first parameter values ​​for clarity). In specific implementations, the target parameter values ​​may include only the first parameter value or may also include other parameter values, which is not a limitation in this embodiment.

[0033] Microphone array beamforming is a spatial filtering technique that utilizes an array of multiple microphones and uses signal processing techniques to selectively enhance sound signals in a specific direction (i.e., the pickup range) while suppressing noise and interference from other directions. In microphone array beamforming algorithms, the pickup range can be adjusted by adjusting parameter values. Specifically, microphone array beamforming algorithms include parameters specifically used to control the pickup range of the microphone array beamforming (hereinafter referred to as pickup range control parameters for clarity). In specific implementations, various algorithms can implement microphone array beamforming, and different algorithms employ different pickup range control parameters, such as the direction parameter of the steering vector, beam direction, pickup angle, and so on. This embodiment does not restrict the microphone beamforming algorithm used, nor does it restrict the specific type of pickup range control parameter. It is understood that any parameter capable of controlling the pickup range of the microphone array beamforming algorithm can be used.

[0034] The first parameter value is the specific value of the sound pickup range control parameter determined based on the current focal length information. In this embodiment, the method for determining the first parameter value based on the focal length information is not limited. It is sufficient that the first parameter value determined based on the focal length information and the focal length represented by the focal length information satisfy the following relationship: the larger the focal length, the smaller the span of the sound pickup range corresponding to the first parameter value; conversely, the smaller the focal length, the larger the span of the sound pickup range corresponding to the first parameter value. It should be noted that the larger the focal length, the farther the camera captures and the smaller the range of the captured image. The smaller the span of the sound pickup range, the more focused the sound after beamforming processing, and the higher the signal-to-noise ratio of long-distance sound. Conversely, the smaller the focal length, the closer the camera captures and the larger the range of the captured image. The larger the span of the sound pickup range, the more diffuse the sound after beamforming processing, and the higher the signal-to-noise ratio of close-range sound, and the larger the range. Therefore, if the first parameter value determined based on the focal length information and the focal length represented by the focal length information satisfy the above relationship, the effect of audio zooming with video can be achieved.

[0035] For example, by combining Figure 2 To explain the meaning of pickup range and its span. Figure 2 Where M is a microphone array, which is used to collect audio signals. Figure 2 For example, 0°-30°, 60°-90° and 120°-150° are three different pickup ranges, and the span is 30°. For example, 80°-100° and 60°-120° are two different pickup ranges, and the span is also different. The span of the pickup range of 80°-100° is 20°, and the span of the pickup range of 60°-120° is 60°. Figure 2The area S corresponding to the pickup range of 60°-120° is marked in the figure. When the pickup range corresponding to the first parameter value is 60°-120° and the signal collected by the microphone array is beamformed according to the first parameter value, the sound signal in area S will be enhanced.

[0036] In a feasible implementation, the sound pickup range corresponding to the first parameter value can cover the shooting direction of the camera, so that the direction of the audio recording is consistent with the shooting direction of the camera. The shooting direction is the spatial direction of the camera's optical axis, which is usually consistent with the user's line of sight. The sound pickup range covers the shooting direction of the camera, that is, the shooting direction is within the sound pickup range. Figure 2 As shown, if the shooting direction is 90° (i.e., direction F in the figure), then the sound pickup range corresponding to the first parameter value is expressed as X°-Y°, where X≤90≤Y. The values ​​of X and Y vary depending on the span. In a specific embodiment, the shooting direction can also be set to the center of the sound pickup range. Then, if the span is 20°, the sound pickup range corresponding to the first parameter value is 80°-100°.

[0037] It should be noted that in Figure 2 In the , the pickup range is 0-180°, and the angle corresponding to the shooting direction is set to 90°. Figure 2 This is just an example. In a specific implementation, the value range of the sound pickup range may be, for example, 0-360°, and the angle corresponding to the shooting direction may be, for example, set to 0°. The specific settings can be made as needed.

[0038] Step S30: Process the signal collected by the microphone array according to the target parameter value to obtain recorded audio.

[0039] Processing the signals collected by the microphones according to the target parameter value may include performing beamforming processing on the signals collected by the microphone array according to the first parameter value. In a specific embodiment, recorded audio can be directly generated based on the beamformed signals, or the recorded audio can be obtained by further processing, which is not limited in this embodiment.

[0040] It should be noted that this embodiment does not limit the adopted beamforming algorithm. Therefore, the specific processing process of performing beamforming processing on the signal collected by the microphone array according to the first parameter value is also not limited.

[0041] It can be understood that by processing the collected signal with a determined target parameter value, the finally generated audio can be determined based on the target parameter value. The target parameter value includes a first parameter value for controlling the pickup range of the microphone array beamforming determined according to the current focal length information, and satisfies the relationship that the larger the focal length, the smaller the span of the pickup range corresponding to the first parameter value. Therefore, the effect of audio zooming with video can be achieved.

[0042] In a specific embodiment, more microphones can be provided in the head-mounted device, that is, a microphone array containing more microphones can be provided to increase the maximum sound pickup distance of the microphone array. Specifically, the more microphones included in the microphone array, the smaller the minimum span of the achievable beamforming pickup range, and accordingly, the farther away the sound can be recorded. In a specific embodiment, the number of matching microphones can be set according to the focal length range of the camera, so that the farthest distance that the camera can shoot matches the farthest distance that the microphone array can record, such as matching the currently mature ultra-long-distance camera; the number of microphones and the layout position of the microphones can also be set in combination with the type of head-mounted device, the hardware architecture implementation method, etc.

[0043] For example, in one feasible embodiment, the microphone array may include microphones disposed on the left and right sides of the head-mounted device to maximize the distance between the microphones and minimize the minimum span of the achievable beamforming pickup range. Figure 3 , illustrates the settings of the four microphones in the smart glasses. A, B, C, and D in the figure are the locations of the four microphones, which are used to record audio. A camera is set in the center of the glasses (i.e. point E) for shooting video.

[0044] In a feasible implementation manner, before step S30, the method further includes: Step S031, determining the number of targets according to the focal length information, wherein the larger the focal length represented by the focal length information, the greater the number of targets; In this embodiment, the number of microphones required (hereinafter referred to as the target number for clarity) can be determined based on the focal length information. That is, the target number determined based on different focal length information may be different. In this embodiment, the method for determining the target number based on the focal length information is not limited. It is sufficient that the first parameter value determined based on the focal length information and the focal length represented by the focal length information satisfy the following relationship: the greater the focal length represented by the focal length information, the greater the number of targets determined; conversely, the smaller the focal length represented by the focal length information, the fewer the number of targets determined. It should be noted that the more microphones used when collecting signals, the smaller the minimum span of the achievable beamforming pickup range. That is, to achieve a smaller pickup range, more microphones are required. Furthermore, the smaller the pickup range, the farther away the sound can be picked up. Therefore, in this embodiment, when the focal length represented by the focal length information is larger, that is, when shooting videos at longer distances, more microphones are used for signal collection, thereby enabling the capture of sounds at longer distances. Conversely, when the focal length represented by the focal length information is smaller, that is, when shooting videos at closer distances, fewer microphones are used for signal collection. This not only enables the capture of sounds at closer distances, but also reduces the complexity of the beamforming process because the number of microphone signal channels that need to be processed is reduced. It is understood that the target number ranges from 1 to N, where N is the number of microphones in the microphone array.

[0045] For example, when the focal length information is the focal length parameter value itself, multiple focal length ranges can be pre-set, such as a first focal length range, a second focal length range, and a third focal length range, where the first focal length range is smaller than the second focal length range, which is smaller than the third focal length range. The number of microphones required for each focal length range can be pre-set, with larger focal length ranges corresponding to a greater number of microphones. When the focal length is within the first focal length range, the number of microphones required for the first focal length range is determined as the target number. When the focal length is within the second focal length range, the number of microphones required for the second focal length range is determined as the target number, and so on.

[0046] As another example, when the focal length information is a focal length mode, the number of microphones required for each focal length mode can be pre-set. The larger the focal length used by the focal length mode, the more microphones are set for that focal length mode. For example, the number of microphones corresponding to macro mode is 1, the number of microphones corresponding to standard mode is 2, and the number of microphones corresponding to telephoto mode is 4. If the current camera focal length mode is detected as macro mode, the target number is determined to be 1; if the current camera focal length mode is detected as standard mode, the target number is determined to be 2; if the current camera focal length mode is detected as telephoto mode, the target number is determined to be 4.

[0047] Step S032: Control a target number of microphones in the microphone array to collect signals.

[0048] After determining the target number, the target number of microphones in the microphone array can be controlled to collect signals. For example, when there are four microphones in the microphone array and the target number is 2, two of the microphones can be controlled to collect signals. Then, in subsequent steps, the signals collected by the two microphones are processed according to the target parameter values ​​to obtain recorded audio. In a specific embodiment, the target number of microphones can be randomly determined from the microphone array; or, it can be pre-specified and fixed. For example, the four microphones in the microphone array are numbered 1, 2, 3, and 4. When the target number is 1, microphone number 1 is fixedly used. When the target number is 2, microphones numbered 1 and 2 are fixedly used, and so on. Alternatively, it can be determined in other ways.

[0049] It should be noted that when the target number value is 1, only a single microphone is used to collect signals. In the subsequent steps, when the signal collected by the microphone is processed according to the target parameter value, beamforming processing may not be performed, which is equivalent to the span of the pickup range being 360°.

[0050] In a feasible implementation, step S032 includes: Step S321, obtaining the image currently captured by the camera; Step S322, identifying the object in the picture to obtain the target position of the object in the picture; There are many ways to identify the position of the subject in the picture, which are not limited in this embodiment. For example, the target detection model implemented by AI (artificial intelligence) algorithm can be used to identify the subject in the picture and obtain the target position.

[0051] The target position refers to the position of the identified object in the image, and can be represented by pixel coordinates or other means, which are not limited in this embodiment. For example, a two-dimensional rectangular coordinate system can be established with the center of the image currently being captured by the camera as the origin. The target position is the coordinate point or area of ​​the object in the two-dimensional rectangular coordinate system, which represents the position of the object.

[0052] Step S323 : selecting a target number of microphones from the microphone array according to the target position, wherein the distances between the selected microphones and the object are smaller than the distances between the unselected microphones and the object.

[0053] In this embodiment, there is no restriction on the method of selecting a target number of microphones from the microphone array according to the target position, as long as the following condition is met: the distance between the selected microphone and the object is smaller than the distance between the unselected microphone and the object.

[0054] In one feasible embodiment, a correspondence can be pre-set between the spatial coordinate system of the microphone array and the two-dimensional rectangular coordinate system of the camera's captured image. The position information in the spatial coordinate system of the microphone array can be converted to a two-dimensional rectangular coordinate system based on this correspondence. Correspondingly, the position information in the two-dimensional rectangular coordinate system of the camera's captured image can be converted to a spatial coordinate system based on this correspondence. Therefore, based on this correspondence, the position information corresponding to the target position in the spatial coordinate system of the microphone array can be determined. Simultaneously, combined with the position information of each microphone in the microphone array in the spatial coordinate system of the microphone array, the distance between each microphone and the target position can be determined, thereby selecting the target number of microphones closest to the object being photographed for signal acquisition.

[0055] In another feasible implementation, the screen can be divided into four quadrants along the horizontal and vertical axes of the screen, and the target position can be the quadrant where the subject is located. For each microphone in the microphone array, the distance relationship between the microphone and each quadrant can be pre-set based on the actual distance between the microphone and each quadrant in the screen. For example, the microphone in the upper left corner of the head-mounted device is closest to the quadrant in the upper left corner of the screen and farthest from the quadrant in the lower right corner. After obtaining the quadrant where the subject is located, the target number of microphones closest to the quadrant is selected for signal acquisition.

[0056] It's understandable that determining the target number based on focal length information ensures that the target number of microphones required for the current focal length is called. This reduces unnecessary microphone calls at shorter focal lengths (i.e., closer distances), while ensuring sufficient microphones are called at longer focal lengths (i.e., farther distances), ensuring recording quality and enhancing the directional capture capability of distant sound sources. Furthermore, when determining the target number, the microphones to be selected are determined based on the positional relationship between the target position of the subject in the frame and each microphone. This allows the microphones to focus on the direction of the subject's sound source, achieving precise alignment of the sound source signal with the physical layout of the microphones, thereby improving audio recording clarity.

[0057] In one feasible implementation, the target parameter value may also include a parameter value for controlling microphone sensitivity (hereinafter referred to as the second parameter value for clarity). Microphone sensitivity refers to the electrical signal strength generated by a microphone at a standard sound pressure level (typically 94 dBSPL, 1 kHz frequency), reflecting the microphone's ability to capture sound signals.

[0058] In specific embodiments, there are various parameters for controlling microphone sensitivity (hereinafter referred to as sensitivity control parameters for clarity). In this embodiment, there is no limitation on the specific sensitivity control parameters; it is understood that any parameter capable of controlling microphone sensitivity will do. For example, in one possible embodiment, the sensitivity control parameter may be a gain factor. That is, the microphone sensitivity can be adjusted by adjusting the gain factor. The gain factor is the factor by which the electrical signal output by the microphone is amplified or reduced. When the gain factor increases, the strength of the electrical signal output by the microphone increases, thereby improving sensitivity. When the gain factor decreases, the strength of the electrical signal output by the microphone decreases, thereby reducing sensitivity.

[0059] The second parameter value is the specific value of the sensitivity control parameter determined based on the current focal length information. In this embodiment, the method for determining the second parameter value based on the focal length information is not limited. It only requires that the second parameter value determined based on the focal length information and the focal length represented by the focal length information satisfy the following relationship: the larger the focal length, the higher the microphone sensitivity corresponding to the second parameter value; conversely, the smaller the focal length, the lower the microphone sensitivity corresponding to the second parameter value. It should be noted that the larger the focal length, the farther the camera can shoot and the smaller the shooting range. The higher the microphone sensitivity, the stronger the microphone's response to sound, thus enabling the capture of fainter sounds, i.e., enabling the collection of sounds from farther away. Conversely, the smaller the focal length, the closer the camera can shoot and the larger the shooting range. The lower the microphone sensitivity, the weaker the microphone's response to sound, making it more suitable for capturing sounds at closer distances. Therefore, by satisfying the above relationship between the second parameter value determined based on the focal length information and the focal length represented by the focal length information, a certain audio zoom-with-video effect can be achieved.

[0060] In the case where the target parameter value includes a first parameter value and a second parameter value, the operation of processing the signal collected by the microphone according to the target parameter value may also include the operation of processing the signal collected by the microphone array according to the second parameter value. In the present embodiment, there is no restriction on the order in which the signals collected by the microphone array according to the first parameter value and the second parameter value are processed. For example, after beamforming the signal collected by the microphone array according to the first parameter value, the beamformed signal may be processed according to the second parameter value, or the signal collected by the microphone array may be processed according to the second parameter value first, and then the beamforming signal may be processed according to the first parameter value. In a specific embodiment, the recorded audio may be directly generated based on the signal processed by the first parameter value and the second parameter value, or the recorded audio may be obtained after further processing, which is not limited in the present embodiment.

[0061] Based on the above first embodiment, a second embodiment of the audio recording method of this application is proposed. In this embodiment, the same or similar contents as those of the above first embodiment can be referred to the above introduction and will not be repeated hereafter. Figure 4 , Figure 4 This is a flow chart of the second embodiment of the audio recording method of the present application. In this embodiment, the step of determining the target parameter value according to the focal length information in step S30 includes: Step A301, determining a target focal length segment in which the focal length represented by the focal length information is located from among various preset focal length segments; The preset focal length segment is obtained by dividing the focal length range of the camera, and the target focal length segment is the preset focal length segment to which the current camera focal length belongs.

[0062] For example, the preset focal length segment can be expressed as ,in, For the Preset focal length segments, and They are respectively the lower and upper focal length limits of the preset focal length segment. is the total number of segments. For example: , , .

[0063] Step A302: Determine the parameter value corresponding to the target focal length segment in each preset parameter value as the target parameter value, wherein the preset parameter value includes the parameter value for controlling the pickup range of the microphone array, and the pickup range span corresponding to the preset focal length segment with a larger focal length is smaller.

[0064] Preset parameter values ​​corresponding to each preset focal length segment can be pre-set. These preset parameter values ​​may include specific values ​​for a pickup range control parameter. It is understood that the pickup range control parameter controls the pickup range of the microphone array beamforming. Therefore, different preset focal length segments correspond to different pickup range spans. In this embodiment, the relationship between the pickup range spans corresponding to the preset focal length segments satisfies the following: Preset focal length segments with larger focal lengths have smaller pickup range spans.

[0065] The preset parameter value corresponding to the target focal length segment is determined as the target parameter value. It can be understood that the specific value of the sound pickup range control parameter in the preset parameter value corresponding to the target focal length segment is determined as the first parameter value.

[0066] In this embodiment, the camera focal length range is divided into multiple preset focal length segments, the target focal length segment where the current focal length is located is matched, and different preset parameter values ​​are pre-set for different preset focal length segments, including parameter values ​​for controlling the span of the pickup range of the microphone array. The corresponding preset parameter values ​​are determined as target parameter values ​​according to the target focal length segment, so that the span of the pickup range is automatically adapted according to the change in focal length, thereby achieving audio zoom with video.

[0067] In one possible implementation, the preset parameter values ​​may also include specific values ​​for a sensitivity control parameter. It will be understood that the sensitivity control parameter is a parameter that controls microphone sensitivity. Therefore, different preset parameter values ​​are assigned to each preset focal length segment, indicating that the microphone sensitivities corresponding to each preset focal length segment are different. In this implementation, the relationship between the microphone sensitivities corresponding to the preset focal length segments satisfies the following: Preset focal length segments with larger focal lengths have higher corresponding microphone sensitivities.

[0068] In this embodiment, the camera focal length range is divided into multiple preset focal length segments, the target focal length segment where the current focal length is located is matched, and different preset parameter values ​​are pre-set for different preset focal length segments, including parameter values ​​for controlling microphone sensitivity. The corresponding preset parameter value is determined as the target parameter value according to the target focal length segment, so that the microphone sensitivity is automatically adapted according to the change in focal length, thereby achieving audio zoom with video.

[0069] Based on the above first and / or second embodiments, a third embodiment of the audio recording method of the present application is proposed. In this embodiment, the same or similar contents as those of the above first and second embodiments can be referred to above and will not be described in detail later. Figure 5 , Figure 5 This is a flowchart of the third embodiment of the audio recording method of the present application. In this embodiment, the step of determining the target parameter value according to the focal length information in step S30 further includes: Step B301: receiving a user adjustment instruction and determining an adjustment factor according to the user adjustment instruction; User adjustment instructions refer to interactive signals in which users actively intervene in the adjustment factor through input devices (such as touch screen sliders, physical knobs, or voice control modules). The adjustment factor is used to represent the degree of user intervention in the target parameter value.

[0070] For example, when the input device is a rotary button, the head-mounted device can determine the adjustment factor based on the rotation angle of the knob; when the input device is a touch screen, the head-mounted device can display a slider on the video shooting interface to receive user adjustment instructions, and then determine the adjustment factor based on the movement position of the slider; when the input device can receive voice input, the head-mounted device can determine and execute the voice command "increase the adjustment factor" based on the current user's voice input, such as "clearer".

[0071] Step B302 : determining a target parameter value according to the adjustment factor and the focal length information, wherein, when the focal length represented by the focal length information remains unchanged, the larger the adjustment factor, the smaller the span of the sound pickup range corresponding to the first parameter value.

[0072] Optionally, the mapping relationship between the first parameter value in the target parameter value and the adjustment factor can be determined according to a preset mapping rule (such as linear mapping, nonlinear mapping, table lookup, etc.), wherein, when the focal length represented by the focal length information remains unchanged, the larger the adjustment factor, the smaller the span of the pickup range corresponding to the first parameter value.

[0073] Optionally, the target parameter value also includes a second parameter value, which can also determine the mapping relationship between the second parameter value in the target parameter value and the adjustment factor according to a preset mapping rule (such as linear mapping, nonlinear mapping, table lookup, etc.), wherein, when the focal length represented by the focal length information remains unchanged, the larger the adjustment factor, the greater the sensitivity corresponding to the second parameter value.

[0074] Optionally, a multiplication-coupled function mapping rule may be used to determine the span of the sound pickup range corresponding to the first parameter value according to the adjustment factor and the video focal length.

[0075] For example, please refer to the formula: Y= / *X ,in, is the adjustment factor, ranging from 1 to 10; For calculation parameters, it is fixed to 0.1; Xis the video focal length, with a value of 0-10. When the video focal length and other calculation parameters remain unchanged, the smaller the adjustment factor, the larger the span of the corresponding calculated pickup range. For example, when When the value is 10, the calculation formula for the pickup range span is Y= 0.01 *X , Y The value range is 0-0.1; when When the value is 1, the calculation formula for the pickup range span is Y= 0.1 *X , Y The value range is 0-1. As you can see, the smaller the adjustment factor, the greater the impact of the video focal length on the sound pickup range. When adjusting the video focal length, the user experiences a more pronounced zoom effect as the audio changes with the video focal length. Therefore, by controlling the adjustment factor, you can adjust the degree to which the video focal length affects the audio recording, achieving a personalized audio recording effect.

[0076] Similarly, a multiplication-coupled function mapping rule may also be used to determine the sensitivity corresponding to the second parameter value according to the adjustment factor and the video focal length.

[0077] For example, please refer to the formula: Y= * *X ,in, is the adjustment factor, and its value range is 0-1; For calculation parameters, it is fixed to 0.1; X is the video focal length, with a value of 0-10. When the video focal length and other calculation parameters remain unchanged, the smaller the adjustment factor, the smaller the corresponding calculated sensitivity. When the value is 1, the calculation formula of sensitivity is Y= 0.1 *X , Y The value range is 0-1; when When the value is 0.1, the sensitivity calculation formula is: Y= 0.01 * X , Y The value range is 0-0.1. As you can see, the smaller the adjustment factor, the less impact the video focal length has on sensitivity. When adjusting the video focal length, the user experiences less noticeable zoom effects on the audio as the video focal length changes. Therefore, by controlling the adjustment factor, you can adjust the degree to which the video focal length affects audio recording, achieving a personalized audio recording experience.

[0078] In this embodiment, by introducing an adjustment factor, the target parameter value can be adjusted in real time according to user instructions, thereby adjusting the degree of the zoom effect when the audio changes with the video focal length, so that the user can dynamically adjust the audio recording effect according to actual needs.

[0079] The present application also provides an audio recording device, please refer to Figure 6 , the audio recording device comprises: A focal length determination module 10 is used to obtain focal length information of the video currently being shot by the camera; a parameter confirmation module 20 for determining a target parameter value based on the focal length information, wherein the target parameter value includes a first parameter value for controlling a sound pickup range of the microphone array beamforming, and the larger the focal length represented by the focal length information, the smaller the span of the sound pickup range corresponding to the first parameter value; The signal processing module 30 is used to process the signal collected by the microphone array according to the target parameter value to obtain recorded audio.

[0080] The audio recording device provided in the embodiment of the present application adopts the audio recording method in the above embodiment. Compared with the prior art, the beneficial effects of the audio recording device provided in the present application are the same as the beneficial effects of the audio recording method provided in the above embodiment, and other technical features in the audio recording device are the same as the features disclosed in the above embodiment method, which will not be repeated here.

[0081] An embodiment of the present application provides a head-mounted device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the audio recording method of the above-mentioned embodiment 1.

[0082] Reference below Figure 7 , which shows a structural schematic diagram of a head-mounted device suitable for implementing an embodiment of the present application. Figure 7 The head-mounted device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.

[0083] like Figure 7As shown, the head-mounted device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory 1002 or programs loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the head-mounted device. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speakers, and vibrator; storage devices 1003 including, for example, magnetic tape and hard disk; and communication devices 1009. The communication device 1009 can allow the head-mounted device to communicate with other devices wirelessly or wired to exchange data. Although the figure shows a head-mounted device with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems can be implemented or have alternatively.

[0084] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are performed.

[0085] Compared with the prior art, the beneficial effects of the head-mounted device provided in the embodiment of the present application are the same as the beneficial effects of the audio recording method provided in the above embodiment, and the other technical features in the head-mounted device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.

[0086] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0087] An embodiment of the present application provides a computer-readable storage medium having computer-readable program instructions (ie, a computer program) stored thereon, and the computer-readable program instructions are used to execute the audio recording method in the above embodiment.

[0088] The computer-readable storage medium provided in the embodiments of the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0089] The computer-readable storage medium may be included in the head-mounted device, or may exist independently without being incorporated into the head-mounted device.

[0090] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the head-mounted device, the head-mounted device performs the functions defined in the method of the embodiment disclosed in this application.

[0091] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0092] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.

[0093] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.

[0094] The readable storage medium provided in the embodiment of the present application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., a computer program) for executing the above-mentioned audio recording method. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in the embodiment of the present application are the same as the beneficial effects of the audio recording method provided in the above-mentioned embodiment, and will not be repeated here.

[0095] An embodiment of the present application further provides a computer program product, including a computer program, which implements the steps of the above-mentioned audio recording method when executed by a processor.

[0096] Compared with the prior art, the beneficial effects of the computer program product provided in the embodiment of the present application are the same as the beneficial effects of the audio recording method provided in the above embodiment, and will not be repeated here.

[0097] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.

Claims

1. An audio recording method, characterized in that: The audio recording method is applied to a head-mounted device, wherein the head-mounted device is provided with a microphone array and a camera, and the audio recording method includes: Obtain the focal length information of the video currently being shot by the camera; Determining a target parameter value according to the focal length information, wherein the target parameter value includes a first parameter value for controlling a sound pickup range of a microphone array beamforming, and a larger focal length represented by the focal length information indicates a smaller span of the sound pickup range corresponding to the first parameter value; The signal collected by the microphone array is processed according to the target parameter value to obtain recorded audio.

2. The audio recording method according to claim 1, wherein: The target parameter value also includes a second parameter value for controlling microphone sensitivity, and the larger the focal length represented by the focal length information, the higher the microphone sensitivity corresponding to the second parameter value.

3. The audio recording method according to claim 1, wherein: Before the step of processing the signal collected by the microphone array according to the target parameter value to obtain recorded audio, the method further includes: determining the number of targets according to the focal length information, wherein the greater the focal length represented by the focal length information, the greater the number of targets; Controlling the target number of microphones in the microphone array to collect signals.

4. The audio recording method according to claim 3, wherein: The step of controlling the target number of microphones in the microphone array to collect signals includes: Get the picture currently captured by the camera; Identifying the object in the picture to obtain a target position of the object in the picture; The target number of microphones is selected from the microphone array according to the target position, wherein the distance between the selected microphones and the object is smaller than the distance between the unselected microphones and the object.

5. The audio recording method according to claim 1, wherein: The step of determining the target parameter value according to the focal length information comprises: Determining a target focal length segment in which the focal length represented by the focal length information is located from each preset focal length segment; The parameter value corresponding to the target focal length segment in each preset parameter value is determined as the target parameter value, wherein the preset parameter value includes a parameter value for controlling the pickup range of the microphone array, and the pickup range corresponding to the preset focal length segment with a larger focal length has a smaller span.

6. The audio recording method according to claim 1, wherein: The sound pickup range corresponding to the first parameter value covers the shooting direction of the camera.

7. The audio recording method according to any one of claims 1 to 6, wherein: The step of determining the target parameter value according to the focal length information comprises: receiving a user adjustment instruction, and determining an adjustment factor according to the user adjustment instruction; The target parameter value is determined according to the adjustment factor and the focal length information, wherein, when the focal length represented by the focal length information remains unchanged, the larger the adjustment factor is, the smaller the span of the pickup range corresponding to the first parameter value is.

8. A head-mounted device, characterized in that: The head-mounted device is provided with a microphone array and a camera, and the head-mounted device also includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the audio recording method according to any one of claims 1 to 7.

9. The head-mounted device according to claim 8, wherein: The microphone array includes at least microphones respectively arranged on the left and right sides of the head-mounted device.

10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the audio recording method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Method and device for controlling adapterization

    CN102137318A

  • Method and system for improving long-shot recording effect during videoing

    CN104244137A

  • Audio processing method and electronic equipment

    CN111050269A

  • Sound signal acquisition method and electronic equipment

    CN111641794A

  • Information processing method and device and electronic equipment

    CN111724823A