Camera acquisition data processing method and device, medium and electronic equipment

By performing tail noise removal processing on the image and audio sampling frame queues of videos recorded by action cameras, the noise problem caused by the shutter button was solved, improving video quality and user experience.

CN121864932APending Publication Date: 2026-04-14GODOX PHOTO EQUIPMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-05
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

When an action camera records video, the noise (similar to a "click") caused by the shutter button affects the video quality, resulting in a poor user experience.

Method used

By acquiring image and audio sample frames and performing tail noise removal processing, including cropping or replacing the first and last noisy frames, noise is eliminated using benchmark noise data or models to generate high-quality video files.

Benefits of technology

It effectively removes noise from camera-recorded videos, improving video quality and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121864932A_ABST
    Figure CN121864932A_ABST
Patent Text Reader

Abstract

The invention provides a camera acquisition data processing method and device, a computer readable medium and electronic equipment. The method comprises the following steps: acquiring an image sampling frame queue comprising a plurality of image sampling frames and an audio sampling frame queue comprising a plurality of audio sampling frames; image sampling frames in the image sampling frame queue are generated through image sampling and compressed encoding on the basis of original image frames collected when a target camera records a target video; the audio sampling frames in the audio sampling frame queue are generated through audio sampling and compression coding on the basis of original audio frames collected when the target camera records the target video; respectively carrying out tail noise elimination processing on the audio sampling frame queues to obtain processed audio sampling frame queues; and generating a video file according to the image sampling frame queue and the processed audio sampling frame queue. According to the invention, noise in videos recorded by cameras such as a motion camera and the like can be removed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of camera technology, and more specifically, to a method, apparatus, computer-readable medium, and electronic device for processing camera-acquired data. Background Technology

[0002] Videos recorded by cameras such as action cameras often exhibit a "click" noise, resulting in low video quality and a poor user experience. Summary of the Invention

[0003] Embodiments of this application provide a method, apparatus, computer-readable medium, and electronic device for processing camera-acquired data, which can at least to some extent remove noise from videos recorded by cameras such as action cameras.

[0004] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.

[0005] According to one aspect of the embodiments of this application, a method for processing camera-acquired data is provided. The method includes: acquiring an image sampling frame queue comprising multiple image sampling frames and an audio sampling frame queue comprising multiple audio sampling frames; the image sampling frames in the image sampling frame queue are generated based on original image frames acquired when the target camera records a target video, through image sampling and compression encoding; the audio sampling frames in the audio sampling frame queue are generated based on original audio frames acquired when the target camera records the target video, through audio sampling and compression encoding; performing tail noise removal processing on the audio sampling frame queue to obtain a processed audio sampling frame queue; and generating a video file based on the image sampling frame queue and the processed audio sampling frame queue.

[0006] According to one aspect of the embodiments of this application, a processing apparatus for camera-acquired data is provided. The apparatus includes: a queue acquisition unit, configured to acquire an image sampling frame queue including multiple image sampling frames and an audio sampling frame queue including multiple audio sampling frames; the image sampling frames in the image sampling frame queue are generated based on original image frames acquired when the target camera records a target video, through image sampling and compression encoding; the audio sampling frames in the audio sampling frame queue are generated based on original audio frames acquired when the target camera records the target video, through audio sampling and compression encoding; a noise cancellation unit, configured to perform tail noise cancellation processing on the audio sampling frame queue to obtain a processed audio sampling frame queue; and a video file generation unit, configured to generate a video file based on the image sampling frame queue and the processed audio sampling frame queue.

[0007] In some embodiments of this application, based on the foregoing scheme, before performing tail noise cancellation processing on the audio sampling frame queue, the noise cancellation unit is further configured to: perform head noise cancellation processing on the audio sampling frame queue to obtain a head-processed audio sampling frame queue; the noise cancellation unit is configured to: perform tail noise cancellation processing on the head-processed audio sampling frame queue to obtain a processed audio sampling frame queue.

[0008] In some embodiments of this application, based on the foregoing scheme, the noise cancellation unit is configured to: perform cropping processing on a number of first audio sample frames located at the beginning of the audio sample frame queue to obtain a head-processed audio sample frame queue; perform cropping processing on a number of second audio sample frames located at the end of the head-processed audio sample frame queue to obtain a processed audio sample frame queue; the processed audio sample frame queue does not contain the first audio sample frames and the second audio sample frames.

[0009] In some embodiments of this application, based on the foregoing scheme, the apparatus further includes an image sampling frame cropping unit; before generating a video file according to the image sampling frame queue and the processed audio sampling frame queue, the image sampling frame cropping unit is configured to: crop a plurality of first image sampling frames located at the beginning and a plurality of second image sampling frames located at the end of the image sampling frame queue to obtain a processed image sampling frame queue; the processed image sampling frame queue does not contain the first image sampling frames and the second image sampling frames; the video file generation unit is configured to: generate a video file according to the processed image sampling frame queue and the processed audio sampling frame queue.

[0010] In some embodiments of this application, based on the foregoing scheme, the noise cancellation unit is configured to: replace a number of first audio sample frames at the beginning of the audio sample frame queue with blank audio sample frames respectively to obtain a head-processed audio sample frame queue; and replace a number of second audio sample frames at the end of the head-processed audio sample frame queue with blank audio sample frames respectively to obtain a processed audio sample frame queue.

[0011] In some embodiments of this application, based on the foregoing scheme, the device further includes a reference noise acquisition unit; before performing head noise cancellation processing on the audio sampling frame queue, the reference noise acquisition unit is configured to: acquire pre-collected first reference noise data and second reference noise data, wherein the first reference noise data is collected after the target camera starts video recording and before the target camera ends video recording; the noise cancellation unit is configured to: perform noise cancellation processing on a plurality of first audio sampling frames located at the head of the audio sampling frame queue based on the first reference noise data to obtain processed first audio sampling frames corresponding to each first audio sampling frame; replace each first audio sampling frame in the audio sampling frame queue with the corresponding processed first audio sampling frame to obtain a head processed audio sampling frame queue; perform noise cancellation processing on a plurality of second audio sampling frames located at the tail of the head processed audio sampling frame queue based on the second reference noise data to obtain processed second audio sampling frames corresponding to each second audio sampling frame; replace each second audio sampling frame in the head processed audio sampling frame queue with the corresponding processed second audio sampling frame to obtain a processed audio sampling frame queue.

[0012] In some embodiments of this application, based on the foregoing scheme, the noise cancellation unit is configured to: perform noise cancellation processing on a plurality of first audio sample frames located at the head of the audio sample frame queue based on a predetermined model, to obtain a processed first audio sample frame corresponding to each first audio sample frame; replace each first audio sample frame in the audio sample frame queue with the corresponding processed first audio sample frame to obtain a head processed audio sample frame queue; perform noise cancellation processing on a plurality of second audio sample frames located at the tail of the head processed audio sample frame queue based on a predetermined model, to obtain a processed second audio sample frame corresponding to each second audio sample frame; replace each second audio sample frame in the head processed audio sample frame queue with the corresponding processed second audio sample frame to obtain a processed audio sample frame queue.

[0013] According to one aspect of the embodiments of this application, a computer-readable medium is provided having a computer program stored thereon, which, when executed by a processor, implements the camera data acquisition processing method as described in the above embodiments.

[0014] According to one aspect of the embodiments of this application, an electronic device is provided, including: one or more processors; and a storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the camera data acquisition processing method as described in the above embodiments.

[0015] According to one aspect of the embodiments of this application, a computer program product is provided, the computer program product including computer instructions stored in a computer-readable storage medium, a processor of a computer device reading the computer instructions from the computer-readable storage medium, and the processor executing the computer instructions, causing the computer device to perform the camera data acquisition processing method as described in the above embodiments.

[0016] In some embodiments of this application, the technical solutions involve first obtaining an image sampling frame queue containing multiple image sampling frames and an audio sampling frame queue containing multiple audio sampling frames. Then, noise removal processing is performed on the tail of the audio sampling frame queue to obtain a processed audio sampling frame queue. Finally, a video file can be generated based on the image sampling frame queue and the processed audio sampling frame queue. Therefore, this application specifically identifies and eliminates noise present at the tail of the audio sampling frame queue, thereby removing noise from the video recorded by the camera, ensuring the quality of the video recorded by the camera, and thus improving the user experience.

[0017] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings: Figure 1 A flowchart is shown for a method of processing camera-acquired data according to an embodiment of this application; Figure 2 A structural block diagram of a camera according to an embodiment of this application is shown; Figure 3A A schematic diagram of the first part of the overall process of a camera generating a video file according to an embodiment of this application is shown; Figure 3B A schematic diagram of the second part of the overall process of a camera generating a video file according to an embodiment of this application is shown; Figure 4 A schematic diagram of an image sampling frame queue and an audio sampling frame queue according to an embodiment of this application is shown; Figure 5 An embodiment according to this application is shown. Figure 1 A flowchart of the steps preceding step 130 and the details of step 130 in the embodiment; Figure 6 A schematic diagram of an image data stream and an audio data stream obtained by simultaneous audio-visual editing according to an embodiment of this application is shown; Figure 7 A schematic diagram of an audio data stream obtained by cropping only audio frames or replacing audio frames with blank audio frames according to an embodiment of this application is shown. Figure 8 A block diagram of a camera data acquisition processing apparatus according to an embodiment of this application is shown; Figure 9 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown. Detailed Implementation

[0019] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this application more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art.

[0020] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.

[0021] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0022] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0023] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0024] Mobile phones are now widely used electronic products, almost everyone owns one, and they make it very convenient to take photos and record videos. However, with the rise of short videos and retro trends, cameras in various forms, such as children's cameras, retro cameras, and action cameras, have opened up a new world in the photography market. Most of these cameras retain a shutter button similar to that of a traditional camera. You can take a photo by pressing the shutter button, and when recording video, you start recording by pressing the shutter button in video recording mode, and stop recording by pressing the shutter button again. The video and audio during recording are then stored in the camera in a predetermined format and can be played back or transferred to external storage devices such as mobile phones and computers.

[0025] However, the inventors discovered that in videos recorded using cameras such as action cameras, noise caused by button presses (similar to "clicks") often occurs, affecting the quality of video recording and resulting in a poor user experience.

[0026] The inventors also discovered that the noise caused by the shutter button is generated as follows: the mechanical vibration of the shutter button is transmitted to the microphone through the device body, producing a transient "click" sound with energy concentrated in the mid-to-high frequency range, and its waveform has a steep rising edge and a specific spectral peak.

[0027] Therefore, this application first provides a method for processing camera-acquired data, which can eliminate noise caused by buttons (such as shutter buttons) and thus improve the user experience.

[0028] Figure 1 A flowchart illustrating a method for processing camera-acquired data according to an embodiment of this application is shown. Please refer to... Figure 1 As shown, the method for processing the data acquired by the camera may include the following steps: In step 110, an image sampling frame queue including multiple image sampling frames and an audio sampling frame queue including multiple audio sampling frames are obtained; the image sampling frames in the image sampling frame queue are generated by image sampling and compression encoding based on the original image frames captured when the target camera records the target video, and the audio sampling frames in the audio sampling frame queue are generated by audio sampling and compression encoding based on the original audio frames captured when the target camera records the target video.

[0029] The target camera can be any camera that can record video in the following way: press the shutter button to start recording video, and press the shutter button again to stop recording video.

[0030] The target camera can be a children's camera, a retro camera, an action camera, an SLR camera, a pocket camera, a mirrorless camera, a panoramic camera, a rugged camera, etc.

[0031] Figure 2 A structural block diagram of a camera according to an embodiment of this application is shown. The target camera can employ... Figure 2 The structure shown. Please refer to [link / reference]. Figure 2 As shown, the camera may include buttons, a main control unit (SoC), an image acquisition unit, an audio acquisition unit, a display screen, a speaker, and a video memory. All modules except the SoC can be electrically connected to the SoC, and the buttons can be shutter buttons. The SoC may integrate a CPU, a video encoder for video encoding, and an audio encoder for audio encoding. The video encoder can be an H.264 video encoder, and the audio encoder can be an AAC encoder (Advanced Audio Coding). Of course, in other embodiments of this application, at least one of the video encoder and audio encoder can be located outside the SoC, that is, at least one of the video encoder and audio encoder can be configured as a separate hardware module.

[0032] The camera's buttons can be any function button used for recording video, such as a shutter button. The image acquisition unit can include an image sensor and an ISP (Image Signal Processor).

[0033] The image sensor can acquire raw image frames at a predetermined sampling frequency (e.g., 60 frames / s) and add a timestamp (image timestamp) to each raw image frame. The ISP is used to perform front-end image processing optimization on the raw image frames acquired by the image sensor, including noise reduction, sharpening, automatic white balance, de-mosaic, color correction and enhancement, lens correction, etc., to obtain high-quality uncompressed image frames (YUV or RGB format).

[0034] The audio acquisition unit may include a microphone and an ADC (Analog-to-Digital Converter). While the image acquisition unit is operating, the microphone picks up sound from the recording environment and converts it into an analog electrical signal. The ADC samples and quantizes the analog electrical signal according to a predetermined sampling frequency (e.g., 48kHz, meaning 48,000 sampling points per second) and sampling precision (e.g., 16-bit) to output digitized audio sample data. Multiple consecutive sampling points are encapsulated into a single PCM (Pulse Code Modulation) audio frame (e.g., 1024 consecutive sampling points as a PCM audio frame). Each PCM audio frame is also timestamped. Image sampling and audio sampling can be based on the same synchronization clock (image frames and audio frames use the same clock source). The PCM audio frame is the original audio frame. When sampling the original image frame and the original audio frame, the system timestamp of the data acquisition is recorded according to the system time.

[0035] Figure 3A A schematic diagram of the first part of the overall process of a camera generating a video file according to an embodiment of this application is shown; Figure 3B A schematic diagram of the second part of the overall process of a camera generating a video file according to an embodiment of this application is shown. Figure 3A and Figure 3B In this context, the same node represents the same position in the overall process; for example, node 1 represents image frame cropping after video encoding is completed. See also... Figure 3A and Figure 3BAs shown, the specific process includes the following steps: The image acquisition unit acquires image frames and performs image processing on them to obtain an image frame queue containing multiple image frames. Then, video encoding is performed on the image frame queue to obtain the aforementioned image sampling frame queue containing multiple image sampling frames. Simultaneously, the audio acquisition unit acquires audio frames and performs audio processing on them to obtain an audio frame queue containing multiple audio frames. The image processing and audio processing can share the same clock source, meaning their time bases can be synchronized by the same clock source. Then, audio encoding is performed on the audio frame queue to obtain the aforementioned audio sampling frame queue containing multiple audio sampling frames. During video encoding of the image frame queue and audio encoding of the audio frame queue, a presentation time stamp (PTS) is assigned to each image frame and audio frame. The display timestamps assigned to image frames and audio frames are information that reflects the order in which media data is presented. They can be based on their respective time bases. During decoding, the system will combine the PTS, system clock, and time base to calculate the specific playback time of video and audio to synchronize the two (not 100% synchronization in a strict sense, but the closest synchronization, which may involve algorithms such as frame skipping to make the two more synchronized).

[0036] In one embodiment of this application, the target video is recorded by the target camera after a predetermined delay from when the shutter button is pressed.

[0037] The scheduled delay can be set according to needs or experience, such as 0.5 seconds or 1 second.

[0038] In this embodiment of the application, by starting to record the target video only after a predetermined delay from when the shutter button is pressed, it can be ensured that noise caused by the shutter button will not be recorded.

[0039] Of course, the target camera can also start capturing raw image frames immediately when the shutter button is pressed, and start capturing raw audio frames only after a predetermined delay from when the shutter button is pressed. For example, the microphone can be activated only after a predetermined delay from when the shutter button is pressed.

[0040] In step 130, the audio sampling frame queue is subjected to tail noise cancellation processing to obtain the processed audio sampling frame queue.

[0041] Noise removal at the end of the audio sampling frame queue can be achieved by cropping the audio sampling frames at the end of the queue; noise in the audio sampling frames at the end of the queue can also be removed using a model; and noise removal can also be performed on the audio sampling frames at the end of the queue based on pre-collected benchmark noise data.

[0042] The inventors discovered that noise typically exists during the brief period at the beginning of video recording and the brief period at the end of video recording.

[0043] Figure 4 A schematic diagram of an image sampling frame queue and an audio sampling frame queue according to an embodiment of this application is shown. Please refer to... Figure 4 As shown, the image sampling frame queue and audio sampling frame queue are acquired starting at the recording start time Ts. The image sampling frame queue includes multiple image sampling frames 1, and the audio sampling frame queue includes multiple audio sampling frames. These audio sampling frames may include several first audio sampling frames 21 at the beginning, several second audio sampling frames 23 at the end, and a third audio sampling frame 22 between the beginning and end. The first audio sampling frames 21 at the beginning are consecutive audio sampling frames, the second audio sampling frames 23 at the end are also consecutive audio sampling frames, and the third audio sampling frame 22 between the beginning and end is also consecutive audio sampling frames. The first audio sampling frames 21 and the second audio sampling frames 23 may contain noisy audio sampling frames. Video recording ends at time Te.

[0044] Figure 5 An embodiment according to this application is shown. Figure 1 A flowchart detailing the steps preceding step 130 and step 130 itself is provided in this embodiment. Please refer to [link / reference needed]. Figure 5 As shown, before performing tail noise cancellation processing on the audio sampling frame queue, the method may further include the following steps: In step 120, the audio sampling frame queue is subjected to header noise cancellation processing to obtain a header-processed audio sampling frame queue.

[0045] Performing tail noise removal processing on the audio sampling frame queue to obtain a processed audio sampling frame queue may specifically include the following steps: In step 130', the tail noise cancellation process is performed on the head-processed audio sampling frame queue to obtain the processed audio sampling frame queue.

[0046] As mentioned earlier, the audio sampling frames at the beginning and end of the audio sampling frame queue will contain noise, so noise removal processing needs to be performed on the beginning and end separately.

[0047] Of course, in other embodiments of this application, the head noise cancellation process and the tail noise cancellation process can be performed simultaneously, or the tail noise cancellation process can be performed first, and then the head noise cancellation process can be performed.

[0048] In one embodiment of this application, the step of performing head noise cancellation processing on the audio sampling frame queue to obtain a head-processed audio sampling frame queue includes: truncating a plurality of first audio sampling frames located at the head of the audio sampling frame queue to obtain a head-processed audio sampling frame queue; the step of performing tail noise cancellation processing on the head-processed audio sampling frame queue to obtain a processed audio sampling frame queue includes: truncating a plurality of second audio sampling frames located at the tail of the head-processed audio sampling frame queue to obtain a processed audio sampling frame queue; the processed audio sampling frame queue does not contain the first audio sampling frames and the second audio sampling frames.

[0049] By removing several first audio sample frames from the beginning and several second audio sample frames from the end of the audio sample frame queue, a processed audio sample frame queue is obtained that does not contain any first or second audio sample frames. The first audio sample frames at the beginning can be the first audio sample frames with the earliest timestamps displayed in the audio sample frame queue; the second audio sample frames at the end can be the first audio sample frames with the latest timestamps displayed in the audio sample frame queue.

[0050] The inventors discovered that the duration of noise at the beginning of a recorded video is shorter than that at the end because the sound produced by pressing a button occurs before the electrical signal to stop recording is generated. Therefore, the number of first audio sample frames contained in the first few first audio sample frames can be less than the number of second audio sample frames contained in the last few second audio sample frames.

[0051] The number of audio sample frames at the beginning and the number at the end that need to be cut from the audio sample frame queue can be set empirically or determined in advance through experiments.

[0052] You can also set the duration of the first audio segment and the duration of the last audio segment to be cut off. Then, based on the duration of the first audio segment, determine the audio sample frames located at the beginning of the audio sample frame queue that need to be cut off from the audio sample frame queue, and based on the duration of the last audio segment, determine the audio sample frames located at the end of the audio sample frame queue that need to be cut off from the audio sample frame queue.

[0053] The duration of the first and last audio segments can be determined in advance through experiments, and the duration of the first audio segment can be shorter than the duration of the last audio segment.

[0054] In one embodiment of this application, before generating a video file based on the image sampling frame queue and the processed audio sampling frame queue, the method for processing camera-acquired data further includes: cropping a plurality of first image sampling frames located at the beginning and a plurality of second image sampling frames located at the end of the image sampling frame queue to obtain a processed image sampling frame queue; the processed image sampling frame queue does not contain the first image sampling frames and the second image sampling frames.

[0055] The first image sampling frame can be an image sampling frame that corresponds to or matches the first audio sampling frame in terms of playback time; the second image sampling frame can be an image sampling frame that corresponds to or matches the second audio sampling frame in terms of playback time. Specifically, the presentation time stamp (PTS) of the first image sampling frame can correspond to or be the same as the presentation time stamp of the first audio sampling frame; the presentation time stamp of the second image sampling frame can be a presentation time stamp that corresponds to or is the same as the presentation time stamp of the second audio sampling frame.

[0056] In this embodiment of the application, by simultaneously cropping the corresponding or matching image sampling frames and audio sampling frames, it can be ensured that the processed image sampling frame queue and the processed audio sampling frame queue remain matched in terms of time length, which facilitates subsequent processing.

[0057] Figure 6 A schematic diagram of an image data stream and an audio data stream obtained by simultaneous audio-visual slicing according to an embodiment of this application is shown. Please refer to... Figure 6 As shown, the image data stream is a queue of processed image sample frames obtained by cropping the image sample frames located at the beginning and end of the stream, and the audio data stream is a queue of processed audio sample frames obtained by cropping the audio sample frames located at the beginning and end of the stream. The image data stream may include one or more GOPs (Group of Pictures), which are a set of consecutive video frames starting with an I-frame in video coding. A GOP may include I-frames (Intra Frames) 31 and P-frames (Predictive Frames) 32. I-frames 31 are independently coded frames, and P-frames 32 are forward-predictive coded frames, encoded with reference to the preceding I-frame or P-frame. The interval between I-frames can be 15, 10, 5, or even 20 or higher; this embodiment does not limit this. Of course, a GOP may further include B-frames; when B-frames are included, a DTS (Decoding Timestamp) is also provided during the encoding process. Without B-frames, PTS and DTS are equivalent; with B-frames, decoding and display are not in the same timing sequence. Figure 6T1 and T2 in the figure represent the timestamps of the audio sampling frames that need to be cut. T1 and T2 can be set based on experience or determined experimentally.

[0058] In one embodiment of this application, the step of performing head noise cancellation processing on the audio sampling frame queue to obtain a head-processed audio sampling frame queue includes: replacing a plurality of first audio sampling frames located at the head of the audio sampling frame queue with blank audio sampling frames respectively, to obtain a head-processed audio sampling frame queue; the step of performing tail noise cancellation processing on the head-processed audio sampling frame queue to obtain a processed audio sampling frame queue includes: replacing a plurality of second audio sampling frames located at the tail of the head-processed audio sampling frame queue with blank audio sampling frames respectively, to obtain a processed audio sampling frame queue.

[0059] That is, noise cancellation processing is performed on the beginning and end of the audio sampling frame queue to obtain a processed audio sampling frame queue. This may include replacing a number of first audio sampling frames at the beginning and a number of second audio sampling frames at the end of the audio sampling frame queue with blank audio sampling frames respectively to obtain a processed audio sampling frame queue.

[0060] Blank audio sampling frames are also known as silent frames.

[0061] Figure 7 A schematic diagram illustrating an audio data stream obtained according to an embodiment of this application by either cropping only audio frames or replacing audio frames with blank audio frames. See also... Figure 7 As shown, several first audio sample frames at the beginning can be replaced with blank audio sample frames 4; several first audio sample frames at the end can also be replaced with blank audio sample frames. Of course, it is also possible not to replace the blank audio sample frames.

[0062] In one embodiment of this application, before performing head-end noise cancellation processing on the audio sampling frame queue, the method further includes: acquiring pre-collected first reference noise data and second reference noise data, wherein the first reference noise data is collected after the target camera starts video recording and before the target camera ends video recording; the head-end noise cancellation processing on the audio sampling frame queue to obtain a head-end processed audio sampling frame queue includes: performing noise cancellation processing on a plurality of first audio sampling frames located at the head of the audio sampling frame queue based on the first reference noise data to obtain the processed first audio sampling frames corresponding to each first audio sampling frame. The audio sampling frame; replacing each first audio sampling frame in the audio sampling frame queue with the corresponding processed first audio sampling frame to obtain a head processed audio sampling frame queue; the step of performing tail noise cancellation processing on the head processed audio sampling frame queue to obtain a processed audio sampling frame queue includes: performing noise cancellation processing on several second audio sampling frames located at the tail of the head processed audio sampling frame queue based on the second reference noise data to obtain a processed second audio sampling frame corresponding to each second audio sampling frame; replacing each second audio sampling frame in the head processed audio sampling frame queue with the corresponding processed second audio sampling frame to obtain a processed audio sampling frame queue.

[0063] The first reference noise data can be the noise data located at the beginning; the second reference noise data can be the noise data located at the end.

[0064] The first reference noise data can be noise data of a first duration recorded from the start of video recording by the target camera; the first reference noise data can also be noise data of a second duration recorded before the end of video recording by the target camera, that is, the duration between the start time of the recorded first reference noise data and the end time of video recording by the target camera can be the second duration.

[0065] The first and second reference noise data can be standard noise data pre-recorded in a laboratory environment.

[0066] The first duration and the second duration can be determined in advance through experiments, and the first duration can be shorter than the second duration.

[0067] Noise in the first audio sampling frame can be removed using the first reference noise data based on an acoustic algorithm; and noise in the second audio sampling frame can be removed using the second reference noise data based on an acoustic algorithm.

[0068] Since the noise generated by buttons when a camera is shooting video is often constant, by using pre-collected baseline noise data to perform noise removal processing on noisy audio sampling frames, it is possible to accurately remove noise while preserving the effective sound in the noisy audio sampling frames.

[0069] In one embodiment of this application, the step of performing head noise cancellation processing on the audio sampling frame queue to obtain a head-processed audio sampling frame queue includes: performing noise cancellation processing on a plurality of first audio sampling frames located at the head of the audio sampling frame queue based on a predetermined model to obtain a processed first audio sampling frame corresponding to each first audio sampling frame; replacing each first audio sampling frame in the audio sampling frame queue with the corresponding processed first audio sampling frame to obtain a head-processed audio sampling frame queue; the step of performing tail noise cancellation processing on the head-processed audio sampling frame queue to obtain a processed audio sampling frame queue includes: performing noise cancellation processing on a plurality of second audio sampling frames located at the tail of the head-processed audio sampling frame queue based on a predetermined model to obtain a processed second audio sampling frame corresponding to each second audio sampling frame; replacing each second audio sampling frame in the head-processed audio sampling frame queue with the corresponding processed second audio sampling frame to obtain a processed audio sampling frame queue.

[0070] It can obtain several first audio sample frames at the beginning and several second audio sample frames at the end of the audio sample frame queue according to a preset duration or number of frames.

[0071] The pre-trained model can be a pre-trained artificial intelligence model (such as a deep learning model) or other algorithmic models.

[0072] Although in the embodiments of this application, the first audio sampling frame and the second audio sampling frame are obtained first, and then the model is used to perform noise cancellation processing on the first audio sampling frame and the second audio sampling frame; in other embodiments of this application, the model can also automatically identify audio sampling frames with noise in the audio sampling frame queue and automatically perform noise cancellation processing.

[0073] In one embodiment of this application, generating a video file based on the image sampling frame queue and the processed audio sampling frame queue includes: generating a video file based on the processed image sampling frame queue and the processed audio sampling frame queue.

[0074] In one embodiment of this application, when only the audio sample frames in the audio sample frame queue are cropped, and the image sample frames in the image sample frame queue are not cropped, generating a video file based on the image sample frame queue and the processed audio sample frame queue may include: encapsulating the image sample frames in the image sample frame queue and the audio sample frames in the processed audio sample frame queue to obtain a video file. Specifically, when encapsulating the audio sample frames in the processed audio sample frame queue, the display timestamp of the audio sample frame in the processed audio sample frame queue is assigned the display timestamp corresponding to the audio sample frame when it was encoded.

[0075] Since only the audio sample frames in the audio sample frame queue are cropped, but the image sample frames in the image sample frame queue are not cropped, if the display timestamp is not controlled, the system will automatically reassign the timestamp based on the current temporal distribution of the audio sample frames. This will cause the image sample frame queue to become out of sync with the processed audio sample frame queue. In this embodiment, the display timestamp of the audio sample frames in the processed audio sample frame queue is assigned the display timestamp corresponding to the audio sample frame when it is encoded. In other words, the display timestamp of the audio sample frames in the processed audio sample frame queue is reset to the display timestamp corresponding to the audio sample frame when it is encoded. This ensures that the display timestamp of the audio sample frames in the processed audio sample frame queue can be kept synchronized with the display timestamp of the image sample frames in the image sample frame queue.

[0076] Please continue reading Figure 3A and Figure 3B As shown, the image sample frame queue obtained through video encoding and the audio sample frame queue obtained through audio encoding are respectively processed by image frame cropping and audio frame cropping to obtain processed image sample frame queues and processed audio sample frame queues. The processed image sample frame queues are then used to generate video packet queues, and the processed audio sample frame queues are used to generate audio packet queues. During image frame cropping and audio frame cropping, the display timestamps of the audio frames are updated to ensure that the display timestamps of the processed image sample frame queues and processed audio sample frame queues can be played synchronously.

[0077] The video packet queue can include multiple image compressed data packets, and the audio packet queue can include multiple audio compressed data packets.

[0078] Specifically, the main controller compresses audio frames using an audio encoder to form audio compressed data packets (AAC data packets) and generates a display timestamp for each AAC data packet (the timestamp of the first audio packet is set to a base value, which can be set to 0 or 1, and subsequent packets are incremented based on this). Both image compressed data packets and audio compressed data packets can be data blocks.

[0079] After the video packet queue and audio packet queue are encapsulated by the encapsulator, a video file can be obtained, which can be an MP4 file. During encapsulation, if the first audio packet written to the container is the 25th audio sample frame in the audio sample frame queue (the first 24 audio sample frames are truncated), then the display timestamp of this audio packet can be assigned a value of 25, so that audio and video can be synchronized starting from the 25th audio sample frame.

[0080] A container is a container with a specific format. Video and audio data (collectively referred to as media data) are encapsulated into the container according to the format (along with metadata such as timestamps and codec information), ultimately forming a video file for storage.

[0081] For example, a video stream is H.264 encoded, the GOP structure is "IPPP...", and the audio format is AAC.

[0082] In an MP4 file, each video frame (I-frame or P-frame) is encapsulated into one or more NALUs (Network Abstraction Layer Units), and then placed as a video sample into a mdat (Media Data Box). NALUs are typically composed of slices of a single frame. Meanwhile, the moov (MovieBox) sample table records the offset, size, timestamp, and whether each sample is a keyframe (I-frames are keyframes, P-frames are not).

[0083] Each AAC frame is put into mdat as an audio sample, and the corresponding information is recorded in the sample table of moov.

[0084] During playback, the player locates the position of each sample based on the index in the moov file, reads and decodes it, and then plays it synchronously according to the timestamp.

[0085] The following describes an embodiment of the apparatus described in this application, which can be used to execute the camera data acquisition processing method described in the above embodiments of this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the camera data acquisition processing method described above in this application.

[0086] Figure 8A block diagram of a camera data acquisition processing apparatus according to an embodiment of this application is shown. (Refer to...) Figure 8 As shown, a camera data processing apparatus 800 according to an embodiment of this application includes: a queue acquisition unit 810, a noise reduction unit 820, and a video file generation unit 830. The queue acquisition unit 810 acquires an image sample frame queue including multiple image sample frames and an audio sample frame queue including multiple audio sample frames. The image sample frames in the image sample frame queue are generated based on image sampling and compression encoding of original image frames acquired when the target camera records the target video. The audio sample frames in the audio sample frame queue are generated based on audio sampling and compression encoding of original audio frames acquired when the target camera records the target video. The noise reduction unit 820 performs tail noise reduction processing on the audio sample frame queue to obtain a processed audio sample frame queue. The video file generation unit 830 generates a video file based on the image sample frame queue and the processed audio sample frame queue.

[0087] In some embodiments of this application, based on the foregoing scheme, before performing tail noise cancellation processing on the audio sampling frame queue, the noise cancellation unit 820 is further configured to: perform head noise cancellation processing on the audio sampling frame queue to obtain a head-processed audio sampling frame queue; the noise cancellation unit 820 is configured to: perform tail noise cancellation processing on the head-processed audio sampling frame queue to obtain a processed audio sampling frame queue.

[0088] In some embodiments of this application, based on the foregoing scheme, the noise cancellation unit 820 is configured to: perform cropping processing on a number of first audio sample frames located at the beginning of the audio sample frame queue to obtain a head-processed audio sample frame queue; perform cropping processing on a number of second audio sample frames located at the end of the head-processed audio sample frame queue to obtain a processed audio sample frame queue; the processed audio sample frame queue does not contain the first audio sample frames and the second audio sample frames.

[0089] In some embodiments of this application, based on the foregoing scheme, the apparatus further includes an image sampling frame cropping unit; before generating a video file according to the image sampling frame queue and the processed audio sampling frame queue, the image sampling frame cropping unit is configured to: crop a plurality of first image sampling frames located at the beginning and a plurality of second image sampling frames located at the end of the image sampling frame queue to obtain a processed image sampling frame queue; the processed image sampling frame queue does not contain the first image sampling frames and the second image sampling frames; the video file generation unit 830 is configured to: generate a video file according to the processed image sampling frame queue and the processed audio sampling frame queue.

[0090] In some embodiments of this application, based on the foregoing scheme, the noise cancellation unit 820 is configured to: replace a number of first audio sample frames at the beginning of the audio sample frame queue with blank audio sample frames respectively to obtain a head-processed audio sample frame queue; and replace a number of second audio sample frames at the end of the head-processed audio sample frame queue with blank audio sample frames respectively to obtain a processed audio sample frame queue.

[0091] In some embodiments of this application, based on the foregoing scheme, the device further includes a reference noise acquisition unit; before performing head noise cancellation processing on the audio sampling frame queue, the reference noise acquisition unit is configured to: acquire pre-collected first reference noise data and second reference noise data, wherein the first reference noise data is collected after the target camera starts video recording and before the target camera ends video recording; the noise cancellation unit 820 is configured to: perform noise cancellation processing on a plurality of first audio sampling frames located at the head of the audio sampling frame queue based on the first reference noise data to obtain a processed first audio sampling frame corresponding to each first audio sampling frame; replace each first audio sampling frame in the audio sampling frame queue with the corresponding processed first audio sampling frame to obtain a head processed audio sampling frame queue; perform noise cancellation processing on a plurality of second audio sampling frames located at the tail of the head processed audio sampling frame queue based on the second reference noise data to obtain a processed second audio sampling frame corresponding to each second audio sampling frame; replace each second audio sampling frame in the head processed audio sampling frame queue with the corresponding processed second audio sampling frame to obtain a processed audio sampling frame queue.

[0092] In some embodiments of this application, based on the foregoing scheme, the noise cancellation unit 820 is configured to: perform noise cancellation processing on a plurality of first audio sample frames located at the head of the audio sample frame queue based on a predetermined model, to obtain a processed first audio sample frame corresponding to each first audio sample frame; replace each first audio sample frame in the audio sample frame queue with the corresponding processed first audio sample frame to obtain a head processed audio sample frame queue; perform noise cancellation processing on a plurality of second audio sample frames located at the tail of the head processed audio sample frame queue based on a predetermined model, to obtain a processed second audio sample frame corresponding to each second audio sample frame; replace each second audio sample frame in the head processed audio sample frame queue with the corresponding processed second audio sample frame to obtain a processed audio sample frame queue.

[0093] Figure 9 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown.

[0094] It should be noted that, Figure 9 The computer system 900 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0095] like Figure 9 As shown, the computer system 900 includes a Central Processing Unit (CPU) 901, which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 902 or programs loaded from storage portion 908 into Random Access Memory (RAM) 903, such as performing the methods described in the above embodiments. The RAM 903 also stores various programs and data required for system operation. The CPU 901, ROM 902, and RAM 903 are interconnected via a bus 904. An Input / Output (I / O) interface 905 is also connected to the bus 904.

[0096] The following components are connected to I / O interface 905: an input section 906 including a keyboard, mouse, etc.; an output section 907 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to I / O interface 905 as needed. Removable media 911, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 910 as needed so that computer programs read from them can be installed into storage section 908 as needed.

[0097] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 909, and / or installed from removable medium 911. When the computer program is executed by central processing unit (CPU) 901, it performs various functions defined in the system of this application.

[0098] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such transmitted data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0099] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0100] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.

[0101] In one aspect, this application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods described in the above embodiments.

[0102] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0103] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the methods according to the embodiments of this application.

[0104] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.

[0105] The data collection and processing plan outlined in this application must be implemented in strict accordance with the requirements of relevant national laws and regulations, obtaining the informed consent or separate consent of the data subject (or having a legal basis as stipulated by the relevant national laws and regulations), and conducting subsequent data use and processing within the scope authorized by laws and regulations and the data subject.

[0106] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A method for processing camera-acquired data, characterized in that, The method includes: A queue of image sampling frames, comprising multiple image sampling frames, and a queue of audio sampling frames, comprising multiple audio sampling frames, are obtained. The image sampling frames in the image sampling frame queue are generated by image sampling and compression encoding based on the original image frames captured when the target camera records the target video. The audio sampling frames in the audio sampling frame queue are generated by audio sampling and compression encoding based on the original audio frames captured when the target camera records the target video. The audio sampling frame queue is subjected to tail noise cancellation processing to obtain the processed audio sampling frame queue; A video file is generated based on the image sampling frame queue and the processed audio sampling frame queue.

2. The method for processing camera-acquired data according to claim 1, characterized in that, Before performing tail noise cancellation processing on the audio sample frame queue, the method further includes: The audio sampling frame queue is subjected to header noise cancellation processing to obtain a header-processed audio sampling frame queue. The step of performing tail noise removal processing on the audio sampling frame queue to obtain a processed audio sampling frame queue includes: The tail noise cancellation process is performed on the head-processed audio sampling frame queue to obtain the processed audio sampling frame queue.

3. The method for processing camera-acquired data according to claim 2, characterized in that, The step of performing header noise removal processing on the audio sampling frame queue to obtain a header-processed audio sampling frame queue includes: The first audio sample frames located at the beginning of the audio sample frame queue are truncated to obtain the audio sample frame queue after the beginning is processed. The step of performing tail noise removal processing on the pre-processed audio sample frame queue to obtain a processed audio sample frame queue includes: The second audio sample frames located at the end of the first processed audio sample frame queue are cropped to obtain a processed audio sample frame queue; the processed audio sample frame queue does not contain the first audio sample frame and the second audio sample frame.

4. The method for processing camera-acquired data according to claim 3, characterized in that, Before generating a video file based on the image sampling frame queue and the processed audio sampling frame queue, the method further includes: The first image sample frames at the beginning and the second image sample frames at the end of the image sample frame queue are cropped to obtain a processed image sample frame queue; the processed image sample frame queue does not contain the first image sample frames and the second image sample frames. The step of generating a video file based on the image sampling frame queue and the processed audio sampling frame queue includes: A video file is generated based on the processed image sampling frame queue and the processed audio sampling frame queue.

5. The method for processing camera-acquired data according to claim 2, characterized in that, The step of performing header noise removal processing on the audio sampling frame queue to obtain a header-processed audio sampling frame queue includes: Replace the first audio sample frames at the beginning of the audio sample frame queue with blank audio sample frames to obtain the audio sample frame queue after the beginning is processed. The step of performing tail noise removal processing on the pre-processed audio sample frame queue to obtain a processed audio sample frame queue includes: The second audio sample frames located at the tail of the processed audio sample frame queue are replaced with blank audio sample frames to obtain the processed audio sample frame queue.

6. The method for processing camera-acquired data according to claim 2, characterized in that, Before performing header noise cancellation on the audio sample frame queue, the method further includes: Acquire pre-collected first reference noise data and second reference noise data, wherein the first reference noise data is collected after the target camera starts video recording and the second reference noise data is collected before the target camera ends video recording; The step of performing header noise removal processing on the audio sampling frame queue to obtain a header-processed audio sampling frame queue includes: Based on the first reference noise data, noise cancellation processing is performed on several first audio sampling frames located at the head of the audio sampling frame queue to obtain the processed first audio sampling frame corresponding to each first audio sampling frame. Replace each of the first audio sample frames in the audio sample frame queue with the corresponding processed first audio sample frame to obtain the first processed audio sample frame queue. The step of performing tail noise removal processing on the pre-processed audio sample frame queue to obtain a processed audio sample frame queue includes: Based on the second reference noise data, noise removal processing is performed on several second audio sampling frames located at the tail of the first processed audio sampling frame queue to obtain the processed second audio sampling frame corresponding to each second audio sampling frame. Replace each of the second audio sample frames in the first processed audio sample frame queue with the corresponding processed second audio sample frame to obtain the processed audio sample frame queue.

7. The method for processing camera-acquired data according to claim 2, characterized in that, The step of performing header noise removal processing on the audio sampling frame queue to obtain a header-processed audio sampling frame queue includes: Based on a predetermined model, noise cancellation processing is performed on several first audio sampling frames located at the head of the audio sampling frame queue to obtain the processed first audio sampling frame corresponding to each first audio sampling frame. Replace each of the first audio sample frames in the audio sample frame queue with the corresponding processed first audio sample frame to obtain the first processed audio sample frame queue. The step of performing tail noise removal processing on the pre-processed audio sample frame queue to obtain a processed audio sample frame queue includes: Based on a predetermined model, noise cancellation processing is performed on several second audio sampling frames located at the tail of the first processed audio sampling frame queue to obtain the processed second audio sampling frame corresponding to each second audio sampling frame. Replace each of the second audio sample frames in the first processed audio sample frame queue with the corresponding processed second audio sample frame to obtain the processed audio sample frame queue.

8. A processing device for camera-acquired data, characterized in that, The device includes: The queue acquisition unit is used to acquire an image sampling frame queue including multiple image sampling frames and an audio sampling frame queue including multiple audio sampling frames; the image sampling frames in the image sampling frame queue are generated by image sampling and compression encoding based on the original image frames captured when the target camera records the target video; the audio sampling frames in the audio sampling frame queue are generated by audio sampling and compression encoding based on the original audio frames captured when the target camera records the target video. A noise cancellation unit is used to perform tail noise cancellation processing on the audio sampling frame queue to obtain a processed audio sampling frame queue. The video file generation unit is used to generate a video file based on the image sampling frame queue and the processed audio sampling frame queue.

9. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the camera data acquisition processing method as described in any one of claims 1 to 7.

10. An electronic device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the method for processing camera-acquired data as described in any one of claims 1 to 7.