Stereo video recording methods, devices, storage media and electronic equipment

By identifying audio loss during stereo video recording and performing audio-visual synchronization processing, the problem of audio-visual asynchrony caused by audio loss is solved, thus improving the recording effect.

CN114339352BActive Publication Date: 2025-10-31GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111674893.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-10-31
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

During stereo video recording, missing audio leads to audio-visual asynchrony, affecting the recording quality.

Method used

By acquiring the audio recording status of stereo video, identifying audio loss states, and performing audio-visual synchronization processing during the recording process, including audio signal supplementation and image frame extraction, audio-visual synchronization is ensured.

Benefits of technology

During stereo video recording, it effectively avoids poor video recording quality caused by missing audio, improves recording disaster recovery capabilities, and ensures audio-visual synchronization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114339352B_ABST
    Figure CN114339352B_ABST
Patent Text Reader

Abstract

This application discloses a stereo video recording method, apparatus, storage medium, and electronic device. The method includes: acquiring the audio recording status corresponding to the stereo video; if the audio recording status is an audio missing status, then performing audio-visual synchronization processing on the stereo video. Using this application embodiment can ensure the audio-visual synchronization effect of the stereo video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a stereo video recording method, apparatus, storage medium and electronic device. Background Technology

[0002] With the development of computer technology, people's demand for high-quality audiovisual videos is constantly increasing. To ensure that videos have good audiovisual effects, stereo video can be recorded. Stereo video has a sense of the location and distribution of each sound source, which can improve the clarity, intelligibility and immersiveness of the sound, and is therefore highly favored by people. Summary of the Invention

[0003] This application provides a stereo video recording method, apparatus, storage medium, and electronic device, the technical solution of which is as follows:

[0004] In a first aspect, embodiments of this application provide a stereo video recording method, the method comprising:

[0005] Get the audio recording status corresponding to the stereo video;

[0006] If the audio recording status is in the audio missing state, then the stereo video is processed for audio-visual synchronization.

[0007] Secondly, embodiments of this application provide a stereo video recording device, the device comprising:

[0008] The status acquisition module is used to acquire the audio recording status corresponding to the stereo video.

[0009] The audio-visual synchronization module is used to perform audio-visual synchronization processing on the stereo video if the audio recording state is an audio missing state.

[0010] Thirdly, embodiments of this application provide a computer storage medium storing a plurality of instructions adapted for loading by a processor and executing the above-described method steps.

[0011] Fourthly, embodiments of this application provide an electronic device that may include: a processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and to execute the above-described method steps.

[0012] The beneficial effects of the technical solutions provided in some embodiments of this application include at least the following:

[0013] In one or more embodiments of this application, the terminal obtains the corresponding audio recording status during the stereo video recording process; if the audio recording status is an audio missing status, the terminal performs audio-visual synchronization processing on the stereo video; this can ensure audio-visual synchronization of stereo video in the event of audio missing, avoid poor video recording effect caused by audio missing, and improve the disaster recovery processing capability during stereo video recording. Attached Figure Description

[0014] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 This is a flowchart illustrating a stereo video recording method provided in an embodiment of this application;

[0016] Figure 2 This is a schematic diagram of a stereo video recording scenario provided in the embodiments of this application;

[0017] Figure 3 This is a schematic diagram illustrating a scenario where audio recording is missing, related to the stereo video recording method provided in this application embodiment;

[0018] Figure 4 This is a schematic diagram of a stereo video recording method provided in the embodiments of this application, which relates to a stereo video encoding scenario;

[0019] Figure 5 This is a flowchart illustrating another stereo video recording method provided in an embodiment of this application;

[0020] Figure 6 This is a schematic diagram of a scene involving audio-visual synchronization related to the stereo video recording method provided in the embodiments of this application;

[0021] Figure 7 This is a schematic diagram of the structure of a stereo video recording device provided in an embodiment of this application;

[0022] Figure 8 This is a schematic diagram of the structure of a status acquisition module provided in an embodiment of this application;

[0023] Figure 9 This is a schematic diagram of the structure of a state determination unit provided in an embodiment of this application;

[0024] Figure 10This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0025] Figure 11 This is a schematic diagram of the structure of the operating system and user space provided in the embodiments of this application;

[0026] Figure 12 yes Figure 11 Architecture diagram of the Android operating system in China;

[0027] Figure 13 yes Figure 11 Architecture diagram of the iOS operating system. Detailed Implementation

[0028] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0029] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. In the description of this application, it should be noted that, unless otherwise expressly specified and limited, "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances. Furthermore, in the description of this application, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist; for example, A and / or B can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.

[0030] The present application will now be described in detail with reference to specific embodiments.

[0031] In one embodiment, such as Figure 1As shown, a stereo video recording method is proposed. This method can be implemented using a computer program and can run on a stereo video recording device based on the von Neumann architecture. The computer program can be integrated into an application or run as a standalone utility application. The stereo video recording device can be a terminal device, including but not limited to: personal computers, tablets, handheld devices, in-vehicle devices, wearable devices, computing devices, or other processing devices connected to a wireless modem. In different networks, the terminal device can be called by different names, such as: user equipment, access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent or user equipment, cellular phone, cordless phone, terminal device in 5G networks or future evolved networks, etc.

[0032] Specifically, the stereo video recording method includes:

[0033] S101: Get the audio recording status corresponding to the stereo video.

[0034] The stereo video can be understood as a video with stereo sound. It is understandable that the recording of stereo video usually requires recording the audio corresponding to the video based on at least two channels, which can also be understood as multi-channel video recording.

[0035] In one or more embodiments, stereo video can typically be based on "binaural audio," a recording technique that highly reproduces the realistic listening experience. Stereo video recording can be based on stereo recording methods to record the corresponding audio portion of the stereo video, such as binaural recording. This recording method simulates the entire process of human hearing, thereby capturing and reproducing stereo sound without distortion. A typical binaural recording scheme involves first creating a simulated human head, and then placing two omnidirectional microphones inside the ear canals of the simulated head, in positions similar to the eardrums of a human ear. During playback, headphones alone can almost perfectly reproduce the spatial sense (360 degrees) of the recording environment, giving the listener an immersive experience. Currently, this stereo recording technology is used in the recording of certain specialized stereo record videos.

[0036] Understandably, while simulated human head recording has excellent stereo fidelity, it has certain limitations and is not suitable for portable recording scenarios in people's daily lives. For example, it is difficult to apply to users who need to record stereo video using portable terminals.

[0037] In one or more embodiments, the terminal records stereo video, which can be achieved using headphones that work with the terminal. These headphones can be wireless or wired. For example, using wireless headphones, a user can use the terminal and wear wireless headphones (such as true wireless TWS earbuds). Sound is picked up using at least two microphones on the wireless headphones. By utilizing the time difference, volume difference, and timbre difference between the user's two ears to determine the audio location, direct sound, and reflected sound, the spatial sense and layering of the corresponding audio portion of the video are captured and recorded. By utilizing the characteristics of the human ear, the user themselves act as the audio carrier for stereo video recording, thereby achieving the effect of recording stereo video based on the user's two ears.

[0038] In a specific implementation scenario, such as Figure 2 As shown, Figure 2 This is a schematic diagram of a stereo video recording scenario. Users can open a target application (such as a camera app) with stereo video recording capabilities on their terminal to record in stereo. To achieve a stereo effect during recording, users wear wireless headphones that work with the terminal. A communication connection (such as Bluetooth or point-to-point) is pre-established between the wireless headphones and the terminal. The terminal's camera records a video image stream containing the user (consisting of at least one image frame), while the wireless headphones record the user's audio video stream (consisting of at least one audio frame). This video image stream and audio video stream together constitute a stereo video. In stereo video, there is typically a one-to-one correspondence between image frames and audio frames at the same time. Synchronization of audio and video is achieved when at least each image frame corresponds to an audio frame during stereo video recording.

[0039] It is understandable that during the recording of stereo video, there may be a situation where a certain audio segment corresponding to the stereo video is missing. For example, the terminal may record the stereo video image stream normally, but due to objective factors, the audio corresponding to a certain image frame in the image stream is not recorded, resulting in the audio corresponding to a certain image frame being missing in the stereo video.

[0040] For example, during recording, after the user starts recording the stereo video stream, they might put on wireless headphones to record with both ears a period of time t after the video stream starts recording. In this case, the audio during "this period t" is missing. When encoding the video stream and the audio stream, there is no audio data during this period t. The audio data recorded after time t will be shifted forward during the encoding process to align with the video stream. Thus, due to the missing audio during "this period t", the final encoded stereo video will be out of sync with the audio and video.

[0041] For example, if a user is recording stereo audio using wireless headphones but is not using the headphones (e.g., removing the headphones, turning them off, or placing them in their case), the wireless headphones or the terminal will sense that the user is not wearing them and will stop using the wireless headphones' audio recording system (e.g., the audio recording component) and switch to the terminal's audio recording system to continue recording. Figure 3 As shown, Figure 3 This is a schematic diagram illustrating a scenario where audio recording is missing. Figure 3 In this case, after the terminal enables stereo recording, it simultaneously records the video image stream and the video audio stream, such as... Figure 3 As shown, when the user is Figure 3 The "00:04.00" time indicates that the wireless headphones were not in use (e.g., the wireless headphones were removed). The terminal senses that the user is not wearing the headphones and switches from the wireless headphones' audio recording system (e.g., an audio recording component) to the terminal's own audio recording system (e.g., loading the terminal's Hardware Abstraction Layer (HAL) for audio recording). Figure 3 The audio recording continues at "00:04.30" as shown. There is a certain switching time difference in the entire switching process (e.g., Figure 3 The time difference shown is 300ms. During the terminal stereo video recording process, the recording thread synchronously calls the audio reading function to record each frame of audio signal. However, switching from the wireless headset's audio recording system (such as the audio recording component) to the terminal's audio recording system will cause the recording thread to be blocked. During this blocking period (e.g., a 300ms switching time difference), there is no audio data, but the video image recording thread continues to record the image stream data. The stereo video encoding process will shift the audio data that will be switched to the terminal's audio recording system later forward. Figure 4 As shown, Figure 4 This is a schematic diagram of a stereo video encoding scenario involved in this application. At "00:04.30", the terminal switches to the audio recording system to continue recording audio data. The audio data recorded after "00:04.30" will be shifted forward during encoding to align with the image frame corresponding to "00:04.00" on the timeline for encoding. This will ultimately cause the recorded stereo video to be out of sync with the audio and video, such as the aforementioned 300ms audio-video desynchronization phenomenon.

[0042] Understandably, electronic devices can acquire the audio recording status corresponding to the stereo video during the recording process, and the audio recording status includes at least an audio missing status and an audio normal status.

[0043] The audio missing state can be understood as the recording state of stereo video when at least part of the image stream corresponding to the audio is missing during the stereo recording process.

[0044] The audio normal state corresponds to the audio missing state. The audio normal state can be understood as the recording state of stereo video when all image streams are missing during the stereo recording process.

[0045] In one feasible implementation, the terminal can monitor the user's wearing status of the headphones used with the terminal during the stereo video recording process (from the start to the end of recording). If the user does not use the headphones during recording, there will usually be audio loss, and the terminal will determine the audio recording status of the stereo video as an audio loss state. If the user uses the headphones normally during recording, there will usually be no audio loss, and the terminal will determine the audio recording status of the stereo video as an audio normal state. Furthermore, monitoring the user's wearing status of the headphones can be achieved through sensors carried on the headphones to detect whether the user is wearing them.

[0046] In one feasible implementation, during stereo video recording, the terminal records the recording time of each audio signal frame to generate a recording timestamp while recording at least one frame of audio signal corresponding to the stereo video. The terminal can obtain the recording timestamps of each audio signal frame corresponding to the stereo video and determine the audio recording status for the stereo video based on the recording timestamps of each audio signal. Specifically, after obtaining the recording timestamps of each audio signal frame, each recording timestamp is compared with the previous recording timestamp (the previous frame of audio signal corresponding to the recording timestamp) to determine the recording interval. This can be understood as calculating the time difference between the recording timestamps of two consecutive audio signals and using this time difference as the recording interval. For example, obtaining the recording timestamps corresponding to the i-th and (i-1)-th audio signals, calculating the time difference between the recording timestamp of the i-th audio signal and the (i-1)-th audio signal, and using this time difference as the recording interval, typically using the recording interval of the i-th audio signal. Here, i is a natural number greater than 1.

[0047] Understandably, we can then determine whether the recording interval matches the time threshold. If the recording interval equals the time threshold, the two are considered to match, and the audio recording status corresponding to the stereo video is determined to be normal. If the recording interval is greater than the time threshold, the two are considered to be mismatched. In this case, there is usually a missing audio, and the audio recording status corresponding to the stereo video is determined to be missing audio.

[0048] Understandably, stereo video recording is typically based on a set signal sampling rate. The time threshold can be determined based on the signal sampling rate, such as by using the reciprocal of the signal sampling rate as the time threshold. When recording each frame of audio signal i (i is an integer), the timestamp corresponding to the current audio signal i and the timestamp of the previous frame "audio signal i-1" are obtained to determine the recording interval. Based on the recording interval, the audio recording status of the stereo video is determined. It can be understood that the audio recording status of stereo video is a continuously updated process as the recording time increases.

[0049] Understandably, the time threshold can also be an empirical value, set according to the actual application of stereo video recording.

[0050] In one feasible implementation, the audio recording status corresponding to the stereo video can be determined based on the audio recording system currently recording the stereo video. The audio recording system can be a hardware abstraction layer or audio recording component responsible for audio recording within the corresponding device. By monitoring changes in the audio recording system recording the stereo video, the recording status of the stereo video's audio signal can be determined.

[0051] In one or more embodiments, the terminal records stereo video, which can be achieved based on a sound pickup device such as headphones used in conjunction with the terminal. The sound pickup device can be wireless or wired headphones. For example, using wireless headphones, the user can use the terminal and wear wireless headphones (such as true wireless TWS headphones). Sound is picked up based on the microphones of at least two channels of the wireless headphones. By utilizing the time difference, volume difference, and timbre difference between the user's two ears to determine the audio location of the recorded video, as well as direct sound and reflected sound, the spatial sense and layering of the corresponding audio part of the video are reflected and recorded. The sound pickup device has an audio recording system.

[0052] Understandably, during the recording of stereo video, the audio recording system on the microphone device is usually assumed to be the default audio recording system; illustratively, such as... Figure 3 As shown, when the user is Figure 3 As shown at "00:04.00", if the wireless headphones are not in use (e.g., removed), the terminal senses that the user is not wearing a microphone (wireless headphones). At this time, in order to record stereo video normally, the terminal will switch from the microphone's audio recording system (e.g., audio recording component) to the terminal's own audio recording system. The microphone's audio recording system is usually the default audio recording system. Switching from the default recording system to the terminal's audio recording system will result in a system switching delay, during which audio signal loss will occur.

[0053] Understandably, the terminal can obtain a default audio recording system for the stereo video. In some embodiments, when the terminal is used in conjunction with a microphone, the audio recording system of the microphone is used as the default audio recording system for the stereo video. In some embodiments, the terminal's audio recording system may also be the default audio recording system, and when the user wears the microphone, the system will switch from the terminal's audio recording system to the microphone's audio recording system.

[0054] Understandably, during the recording process, the terminal can detect the target audio recording system corresponding to the stereo video in real time or periodically. This audio recording system is used to record the audio signal corresponding to the stereo video; the purpose is to detect whether the target audio recording system is the default audio recording system.

[0055] Understandably, if the target audio recording system does not match the default audio recording system, and the system is usually switched from the default audio recording system to another audio recording system, the terminal can determine that the stereo video is in a state of missing audio.

[0056] Understandably, if the target audio recording system matches the default audio recording system, the terminal can determine that the stereo video is in a normal audio state.

[0057] S102: If the audio recording state is an audio missing state, then the stereo video is processed for audio-visual synchronization.

[0058] Understandably, audio-visual synchronization processing for stereo video involves synchronizing the audio and video during the recording process of the stereo video, specifically the data before encoding, rather than performing the synchronization process after the stereo video has been encoded and then playing it back.

[0059] In one or more embodiments, the audio recording of stereo video is a continuous recording process, during which the audio signal recording may be interrupted due to objective factors. In this application, when the encoding of the stereo video is not completed (after encoding the video audio stream and video image stream, the final encoded stereo video is generated upon completion of encoding), audio-visual synchronization processing is performed on the video audio stream and / or video image stream corresponding to the (incompletely encoded) stereo video. After the audio-visual synchronization processing is completed and the recording of the video audio stream and video image stream is finished, the video audio stream and video image stream corresponding to the stereo video are encoded to generate the encoded stereo video.

[0060] In one or more embodiments, after determining the audio missing state, the terminal can determine the missing position of the audio stream corresponding to the stereo video (the missing position can be understood as the missing audio signal position corresponding to the missing time), and fill the audio data at the missing position with a signal, such as filling in a pre-set reference audio signal, such as the reference audio signal corresponding to a certain background music.

[0061] In one or more embodiments, after determining the audio missing state, the terminal can determine the audio missing position of the video image stream corresponding to the stereo video, and perform frame extraction processing on at least one frame image corresponding to the audio missing position to align the video image stream and the video audio stream, ensuring audio-visual synchronization. For example, if the audio missing position of the video image stream is between the m-th frame image and the n-th frame image indicated by the missing time, then the corresponding frame image image corresponding to "the m-th frame image and the n-th frame image" is subjected to frame extraction processing.

[0062] In one feasible implementation, a frame-skipping threshold can be set to ensure smooth playback of the image after frame-skipping processing, reducing the user's visual sensitivity to the image. The terminal can determine the total number of frames corresponding to the audio missing time in the video image stream. When the total number of frames is less than or equal to the frame-skipping threshold, image frame skipping is performed on the image frames indicated by the total number of frames. When the total number of frames is greater than the frame-skipping threshold, if image frame skipping is performed on the image frames corresponding to the total number of frames, the user's visual sensitivity to the missing image will be higher; the terminal can then supplement the audio stream with audio signals.

[0063] Furthermore, the terminal can simultaneously perform image frame extraction based on a target frame value that is less than the extraction threshold, and simultaneously fill in the remaining missing audio data corresponding to the "frame difference between the total number of frames and the target number of frames" with audio signals. This ensures that the user's visual sensitivity to missing images is not high, improving the intelligence of audio-visual synchronization processing.

[0064] In this embodiment, the terminal obtains the corresponding audio recording status during the stereo video recording process; if the audio recording status is an audio missing status, the terminal performs audio-visual synchronization processing on the stereo video; this can ensure audio-visual synchronization of stereo video in the event of audio missing, avoid poor video recording effect caused by audio missing, and improve the disaster recovery processing capability during stereo video recording.

[0065] Please see Figure 5 , Figure 5 This is a flowchart illustrating another embodiment of the stereo video recording method proposed in this application. Specifically:

[0066] S201: Obtain the recording timestamp of at least one frame of audio signal corresponding to the stereo video;

[0067] In one or more embodiments, when the terminal begins stereo video recording, the operating system running on the terminal can create a multimedia recording object (such as a MediaRecorder object). The operating system controls the encapsulated AudioSource class to create an AudioRecorder object (AudioRecorder is an audio recording track used to record sound). Each frame of the recorded audio signal is read in a loop through a callback function (such as the dataCallback function). On one hand, the callback function (such as the dataCallback function) can obtain the timestamp and data volume of the recorded audio signal through AudioFlinger by calling the timestamp acquisition function (getTimestamp function). AudioFlinger can be understood as the executor of the audio strategy, responsible for managing the input / output stream devices and processing and transmitting the audio stream data. On the other hand, the read audio data, i.e., at least one frame of the currently recorded audio signal, can be written into a buffer (such as a buffer). For example, the queueInputBuffer function can be called to write the data corresponding to the audio signal into the buffer (such as a buffer).

[0068] Understandably, during the audio recording process of stereo video, the terminal's operating system creates a timestamp variable in AudioRecord, reads the timestamp value of each audio signal from the timestamp variable, and saves it. This allows the terminal to subsequently obtain the recording timestamp of at least one frame of audio signal corresponding to the stereo video through the timestamp variable to determine whether audio loss has occurred. In other words, the terminal obtains at least one frame of audio signal recorded for the stereo video and saves the recording timestamp corresponding to each of the at least one frame of audio signal based on the timestamp variable. This allows the terminal to obtain the recording timestamp of the current audio signal corresponding to the stereo video through the timestamp variable during each timestamp loop reading process (e.g., obtaining the timestamp of the recorded audio signal based on the getTimestamp function in each getTimestamp loop).

[0069] S202: Calculate the recording interval time between two consecutive frames of the audio signal based on the recording timestamp of each audio signal;

[0070] Understandably, the reading of the recording timestamp of each frame of audio signal can be based on the callback cycle of the corresponding callback function. For example, based on the callback cycle corresponding to the audio sampling rate, the function execution flow of each round of calling the callback function can be traversed. In the getTimestamp loop involved in the function execution flow of each callback function, the recording timestamp of the audio signal read by the callback function is saved through the timestamp variable. Based on the recording timestamp of the audio signal recorded by the timestamp variable, the recording status of the stereo video can be further determined. Specifically, the recording interval time of the current audio signal relative to the previous frame of audio signal can be calculated through the recording timestamp of the audio signal. The audio recording status includes audio missing status and audio normal status.

[0071] The recording interval can be understood as the difference between the recording time of the audio signal and the recording time of the previous frame of audio signal. It can also be understood as the time difference between two consecutive audio signals calculated by obtaining the recording timestamps of the two consecutive audio signals. The time difference is also the recording interval.

[0072] The audio missing state can be understood as the recording state of stereo video when at least some image frames are missing during the stereo recording process.

[0073] The normal audio state corresponds to the missing audio state. The normal audio state can be understood as the recording state of stereo video when the audio signal is missing and all image frames are absent during the stereo recording process.

[0074] S203: If the recording interval is greater than the time threshold, then the audio recording status of the stereo video is determined to be an audio missing status;

[0075] S204: If the recording interval is less than or equal to the time threshold, then the audio recording status for the stereo video is determined to be the normal audio status.

[0076] Understandably, stereo video recording is typically based on a set signal sampling rate. The time threshold can be determined based on the signal sampling rate, such as by using the reciprocal of the signal sampling rate as the time threshold. When recording each frame of audio signal i (where i is an integer), the recording interval is determined by obtaining the timestamp corresponding to the current audio signal i and the timestamp of the previous frame "audio signal i-1". The audio recording status of the stereo video is then determined based on the recording interval. It can be understood that the audio recording status of stereo video is a continuously updated process that progresses with the recording time.

[0077] Understandably, the time threshold can also be an empirical value, set according to the actual application of stereo video recording. For example, an empirical value could be set to 200ms, 110ms, etc.

[0078] S205: If the audio recording state is an audio missing state, then obtain the audio missing time corresponding to the stereo video;

[0079] Understandably, the terminal can determine the reference signal data corresponding to the missing audio time and perform audio-visual synchronization processing on the stereo video based on the reference signal data. The reference signal data, categorized by data type, can be either audio data or image data; that is, the terminal can perform audio-visual synchronization processing by adding corresponding audio signal data to the stereo video's audio-video stream to supplement the missing audio signal portion and ensure audio-visual synchronization; or, it can delete corresponding image signal data from the stereo video's image stream to remove the image frames corresponding to the missing audio signal portion and ensure audio-visual synchronization; or, it can add corresponding audio signal data to the stereo video's audio-video stream and delete corresponding image signal data from the stereo video's image stream to ensure audio-visual synchronization.

[0080] S206: Obtain the recording sampling rate corresponding to the stereo video, and determine the amount of missing data based on the audio missing time and the recording sampling rate.

[0081] Understandably, determining the reference signal data corresponding to the audio missing time can be achieved by: obtaining the recording sampling rate corresponding to the stereo video, determining the amount of missing data based on the audio missing time and the recording sampling rate, and determining the reference signal data corresponding to the amount of missing data.

[0082] The amount of missing data can be obtained by multiplying the recording interval time by the recording sampling rate. The recording sampling rate is a pre-set signal sampling rate for stereo video, which can be understood as the amount of analog signal sampled by the recording device (such as a terminal) per unit time. It can be understood as the number of signal frames with missing data being the amount of missing data.

[0083] S207: Obtain the first audio data corresponding to the amount of missing data, and perform audio completion processing on the stereo video based on the first audio data;

[0084] Understandably, the reference signal data corresponding to the determined amount of missing data can be audio data. Therefore, the terminal simply needs to obtain the reference signal data corresponding to the amount of missing data to be supplemented. In some implementations, the reference signal data can be audio signal data where all signal characteristic values ​​are target values ​​(e.g., 0). By supplementing the reference signal data corresponding to the target values, the missing portion is effectively muted for the user's auditory perception, achieving audio-visual synchronization.

[0085] In one or more embodiments, the terminal can generate first audio data corresponding to the amount of missing data required for the audio missing time based on the target audio features, and add the first audio data to the audio missing position corresponding to the stereo video. For example, if there is a missing audio signal between the audio signal of the i-th frame and the (i-1)-th frame, the data can be added to the position between the audio signal of the i-th frame and the (i-1)-th frame.

[0086] The target audio feature can be understood as a pre-set audio attribute feature used to add audio signals. The terminal generates the first audio parameter according to the pre-set target audio feature. Taking the audio missing time as an example, the terminal can generate the first audio data corresponding to the missing data amount of the corresponding audio missing time.

[0087] In one or more embodiments, the target audio feature can be a target value set for the audio feature. For example, if the target value is set to 0, then this segment of the first audio data will usually be presented as silent.

[0088] In one or more embodiments, the terminal can select appropriate first audio data based on the audio background type corresponding to the current stereo video for audio filling. Specifically, the terminal can determine the audio background type corresponding to the stereo video, generate corresponding first audio data based on the audio background type, and add the first audio data to the audio missing position corresponding to the stereo video. This can be understood as adding the first audio data to the audio missing position of the video audio stream corresponding to the stereo video. The audio missing position can be determined based on the signal missing position corresponding to the aforementioned audio missing state.

[0089] Understandably, the terminal can pre-establish at least one audio mapping relationship between a reference audio background type and reference audio data. This mapping relationship can be represented in the form of a mapping table, mapping combination, mapping array, etc. After determining the audio background type corresponding to the stereo video, the terminal obtains the audio data corresponding to the audio background type as the first audio data based on the audio mapping relationship. In one or more embodiments, the obtained first audio data can be instrumental music, background music, accompaniment music, etc. Using instrumental music as the first audio data for audio supplementation allows for a smooth transition between the supplemented audio data and the original audio stream, improving the audio-visual synchronization effect of the stereo video.

[0090] In one feasible implementation, before recording stereo video, the terminal can obtain the audio background type set by the user for the stereo video, such as pop, folk, rock, etc.

[0091] In one feasible implementation, the audio background type can be determined by the terminal performing image recognition processing on the stereo video to determine the audio background type corresponding to the stereo video.

[0092] Understandably, the terminal can pre-train a type recognition model. Specifically, the terminal can recognize stereo video based on the pre-trained type recognition model. The trained type recognition model performs image recognition on the image portions of the stereo video. For example, by recognizing the stereo video through the type recognition model, the audio background type corresponding to the image portion of the recorded stereo video can be determined. The stereo video (or simply a segment of the recorded stereo video image stream) is input into the type recognition model, and the audio background type is output.

[0093] A large amount of sample video data can be obtained in advance, and reference type labels can be marked on the sample video data. Then, the sample video data can be input into the initial type recognition model for training, and a trained type recognition model can be obtained.

[0094] Specifically, in practical applications, the type recognition model can be implemented by fitting one or more of the following deep learning object analysis algorithms: Convolutional Neural Network (CNN), Deep Neural Network (DNN), Recurrent Neural Networks (RNN), Embedding, Gradient Boosting Decision Tree (GBDT), and Logistic Regression (LR). Furthermore, an error backpropagation algorithm can be introduced to optimize the existing neural network model, which can improve the recognition accuracy of the type recognition model based on the neural network model.

[0095] In one feasible implementation, the terminal can perform audio recognition processing on the stereo video to determine the audio background type corresponding to the stereo video.

[0096] Understandably, the terminal can pre-train a type recognition model (the type recognition model can be the same as described above, possessing the ability to recognize images and / or audio). Specifically, the terminal can identify the audio stream portion of a stereo video based on the pre-trained type recognition model, and perform image recognition on the audio stream portion of the stereo video using the trained type recognition model. For example, by identifying the stereo video using the type recognition model, the audio background type corresponding to the audio portion of the recorded stereo video can be determined. By inputting the stereo video (or simply a segment of the recorded stereo video audio stream) into the type recognition model, the audio background type is output.

[0097] In a specific implementation scenario, such as Figure 6 As shown, Figure 6 This is a schematic diagram of a scene involving audio-visual synchronization, as described in this application. Figure 6 In this case, after the terminal enables stereo recording, it simultaneously records the video image stream and the video audio stream, such as... Figure 6 As shown, when the user is Figure 6 The "00:04.00" time indicates that the wireless headphones were not in use (e.g., the wireless headphones were removed). The terminal senses that the user is not wearing the headphones and switches from the wireless headphones' audio recording system (e.g., an audio recording component) to the terminal's own audio recording system (e.g., loading the terminal's (audio) hardware abstraction layer HAL for audio recording). Figure 6 The audio recording continues at "00:04.30" as shown. There is a certain switching time difference in the entire switching process (e.g., Figure 6 The time difference shown is 300ms). The terminal determines the first audio data corresponding to the audio missing time (e.g., the audio missing time is 300ms) and fills the missing position in the video audio stream of the stereo video (that is, the position of the missing 300ms audio stream signal) with the first audio data to complete the audio completion process.

[0098] Understandably, in each getTimestamp loop, the timestamp value of the current audio signal is compared with the timestamp value of the previous audio signal to determine the recording interval. If the recording interval is greater than the time threshold, a jump occurs, and the audio recording status is in the audio missing state. By determining the first audio data corresponding to the amount of missing data required for the audio missing time, the "first audio data" is returned to MediaRecorder (multimedia recording object) through the getInputFramesLost interface for audio completion processing, thereby completing the audio data and ensuring audio-visual synchronization.

[0099] S208: Obtain the first image data corresponding to the amount of missing data, and perform image frame extraction processing on the stereo video based on the first image data;

[0100] The first image data can be understood as the image data in the stereo video that needs to be extracted, which is usually a portion of the image frame data in the video image stream of the stereo video (such as during or while recording). This portion of the image data to be extracted corresponds to the aforementioned determined audio missing time, such as the first image data corresponding to the audio missing time.

[0101] In one or more embodiments, after determining the audio missing time, the terminal can determine the audio missing position of the video image stream corresponding to the stereo video based on the audio missing time (e.g., audio missing duration), and perform frame extraction processing on at least one frame corresponding to the audio missing position to align the video image stream and the video audio stream, ensuring audio-visual synchronization. For example, if the audio missing position of the video image stream is from frame m to frame n, then the frames corresponding to "frame m to frame n" are subjected to frame extraction processing.

[0102] S209: Determine the first data quantity and the second data quantity based on the missing data quantity, and obtain the second audio data corresponding to the first data quantity and the second image data corresponding to the second data quantity.

[0103] S210: Perform audio completion processing on the stereo video based on the second audio data and perform image frame extraction processing on the stereo video based on the second image data.

[0104] Understandably, considering the effect of audio-visual synchronization when some audio signals are missing, processing both the video stream and audio stream of stereo video can shorten the number of image frames for the missing audio signal portion, so that the video can quickly transition to the audio portion when playing the encoded stereo video, thus improving the intelligence of audio-visual synchronization.

[0105] The second audio data can be understood as the audio signal data corresponding to the audio completion processing part during the simultaneous image and audio processing;

[0106] The second image data can be understood as the image signal data corresponding to the audio completion processing part during the simultaneous image and audio processing;

[0107] In one feasible implementation, a frame-skipping threshold can be set to ensure smooth playback of the image after frame-skipping processing, minimizing the image sensitivity to the user's naked eye. The terminal can determine the amount of missing data (e.g., the number of missing audio signal frames) corresponding to the audio missing time in the video image stream; use a target number of frames less than the frame-skipping threshold as the second data amount to obtain the second image data corresponding to the second data amount; and generate the second audio data corresponding to the difference based on the difference between the shortened time after frame-skipping and the audio missing time corresponding to the second image data; for example, the difference between the missing data amount and the second data amount can be calculated, and a first data amount can be determined based on the difference to obtain the second audio data of the first data amount.

[0108] Understandably, the stereo video is then processed by audio completion based on the second audio data and by image frame extraction based on the second image data. This reduces the user's visual sensitivity to missing images while ensuring the audio-visual synchronization of the stereo video. In one or more embodiments, the image frame extraction method can be an interval frame extraction method, such as extracting one image frame at intervals of a target value, or a continuous frame extraction method.

[0109] In one feasible implementation, the terminal can also set a frame extraction ratio (e.g., 0.5), that is, for a segment of image frames (let's say frame v) with missing audio, the number of extracted frames is calculated according to the frame extraction ratio (c) and the number of extracted frames (v frame), and the number of extracted frames is determined, which is also the second image data indicated by the second data volume; then, based on the difference between the shortening time after frame extraction and the audio missing time corresponding to the second image data, the second audio data corresponding to the difference is generated. For example, the difference between the missing data volume and the second data volume can be calculated, and a first data volume can be determined based on the difference to obtain the second audio data of the first data volume.

[0110] In this embodiment, the terminal obtains the corresponding audio recording status during stereo video recording. If the audio recording status is an audio missing status, the terminal performs audio-visual synchronization processing on the stereo video. This ensures audio-visual synchronization in the event of audio loss in the stereo video, avoiding poor video recording quality caused by audio loss and improving the disaster recovery capability during stereo video recording. Furthermore, audio-visual synchronization can be achieved in multiple ways, enriching the audio-visual synchronization processing methods and enhancing the intelligence of video recording.

[0111] The following will combine Figure 7 This application provides a detailed description of the stereo video recording device provided in its embodiments. It should be noted that... Figure 7 The stereo video recording apparatus shown is used to perform the present application. Figures 1-6The methods shown in the embodiments are for illustrative purposes only, illustrating the parts relevant to the embodiments of this application. For specific technical details not disclosed, please refer to this application. Figures 1-6 The example shown.

[0112] Please see Figure 7 This diagram illustrates the structure of a stereo video recording device according to an embodiment of this application. The stereo video recording device 1 can be implemented as all or part of a user terminal through software, hardware, or a combination of both. According to some embodiments, the stereo video recording device 1 includes a status acquisition module 11 and an audio-visual synchronization module 12, specifically used for:

[0113] Status acquisition module 11 is used to acquire the audio recording status corresponding to the stereo video;

[0114] The audio-visual synchronization module 12 is used to perform audio-visual synchronization processing on the stereo video if the audio recording state is an audio missing state.

[0115] Optional, such as Figure 8 As shown, the status acquisition module 11 includes:

[0116] The timestamp acquisition unit 111 is used to acquire the recording timestamp of at least one frame of audio signal corresponding to the stereo video;

[0117] The status determination unit 112 is used to determine the audio recording status of the stereo video based on the recording timestamp of each audio signal.

[0118] Optionally, the stereo video recording device 1 is specifically used for:

[0119] The recording timestamp of at least one frame of audio signal corresponding to the stereo video is read using the timestamp variable.

[0120] Optional, such as Figure 9 As shown, the state determination unit 112 includes:

[0121] The time determination subunit 1121 is used to determine the recording interval time for the audio signal based on the recording timestamp of the at least one frame of audio signal;

[0122] The state determination subunit 1122 is used to determine the audio recording state for the stereo video based on the recording interval time.

[0123] Optionally, the state determination subunit 1122 is specifically used for:

[0124] If the recording interval is greater than the time threshold, then the audio recording status for the stereo video is determined to be an audio missing status.

[0125] If the recording interval is less than or equal to the time threshold, then the audio recording status for the stereo video is determined to be normal audio status.

[0126] Optionally, the audio-visual synchronization module 12 is specifically used for:

[0127] Obtain the audio missing time corresponding to the stereo video;

[0128] Based on the audio missing time, reference signal data is obtained, and the stereo video is processed for audio-visual synchronization based on the reference signal data.

[0129] Optionally, the audio-visual synchronization module 12 is specifically used to: obtain the recording sampling rate corresponding to the stereo video, and determine the amount of missing data based on the audio missing time and the recording sampling rate;

[0130] Determine the reference signal data corresponding to the amount of missing data.

[0131] Optionally, the audio-visual synchronization module 12 is specifically used to: acquire the first audio data corresponding to the amount of missing data; or,

[0132] Obtain the first image data corresponding to the amount of missing data; or,

[0133] Based on the amount of missing data, a first data amount and a second data amount are determined, and second audio data corresponding to the first data amount and second image data corresponding to the second data amount are obtained.

[0134] Optionally, the audio-visual synchronization module 12 is specifically used to: obtain a frame-skipping threshold, determine a second data amount based on the frame-skipping threshold and the amount of missing data, wherein the second data amount is less than or equal to the frame-skipping threshold;

[0135] The first data quantity is determined based on the difference between the missing data quantity and the second data quantity.

[0136] Optionally, the audio-visual synchronization module 12 is specifically used for:

[0137] If the reference signal data is the first audio data, then the stereo video is subjected to audio completion processing based on the first audio data;

[0138] If the reference signal data is first image data, then the stereo video is subjected to image frame extraction processing based on the first image data;

[0139] If the reference signal data is second audio data and second image data, then the stereo video is subjected to audio completion processing based on the second audio data and image frame extraction processing based on the second image data.

[0140] Optionally, the audio-visual synchronization module 12 is specifically used to: generate first audio data corresponding to the amount of missing data based on the target audio features; or,

[0141] Determine the audio background type corresponding to the stereo video, and generate the first audio data corresponding to the audio missing time based on the audio background type.

[0142] Optionally, the status acquisition module 11 is specifically used to: acquire the default audio recording system for the stereo video;

[0143] A target audio recording system corresponding to the stereo video is detected, and the audio recording system is used to record the audio signal corresponding to the stereo video.

[0144] If the audio recording system is not compatible with the default audio recording system, then the stereo video is determined to be in an audio missing state.

[0145] If the audio recording system matches the default audio recording system, then the stereo video is determined to be in a normal audio state.

[0146] It should be noted that the stereo video recording device provided in the above embodiments is only illustrated by the division of the above functional modules when executing the stereo video recording method. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the stereo video recording device and the stereo video recording method embodiments provided in the above embodiments belong to the same concept, and the implementation process is detailed in the method embodiments, which will not be repeated here.

[0147] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0148] In this embodiment, the terminal obtains the corresponding audio recording status during stereo video recording. If the audio recording status is an audio missing status, the terminal performs audio-visual synchronization processing on the stereo video. This ensures audio-visual synchronization in the event of audio loss in the stereo video, avoiding poor video recording quality caused by audio loss and improving the disaster recovery capability during stereo video recording. Furthermore, audio-visual synchronization can be achieved in multiple ways, enriching the audio-visual synchronization processing methods and enhancing the intelligence of video recording.

[0149] This application also provides a computer storage medium that can store multiple instructions, which are adapted to be loaded and executed by a processor as described above. Figures 1-6The stereo video recording method described in the illustrated embodiment can be found in the following documentation for its specific execution process. Figures 1-6 The specific details of the illustrated embodiments will not be elaborated here.

[0150] This application also provides a computer program product storing at least one instruction, which is loaded and executed by the processor as described above. Figures 1-6 The stereo video recording method described in the illustrated embodiment can be found in the following documentation for its specific execution process. Figures 1-6 The specific details of the illustrated embodiments will not be elaborated here.

[0151] Please refer to Figure 10 This diagram illustrates a structural block diagram of an electronic device provided in an exemplary embodiment of this application. The electronic device in this application may include one or more components such as a processor 110, a memory 120, an input device 130, an output device 140, and a bus 150. The processor 110, memory 120, input device 130, and output device 140 may be connected via the bus 150.

[0152] Processor 110 may include one or more processing cores. Processor 110 connects to various parts of the electronic device via various interfaces and lines, and performs various functions and processes data of electronic device 100 by running or executing instructions, programs, code sets, or instruction sets stored in memory 120, and by calling data stored in memory 120. Optionally, processor 110 may be implemented using at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). Processor 110 may integrate one or more of the following: central processing unit (CPU), graphics processing unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into processor 110 and may be implemented separately using a communication chip.

[0153] The memory 120 may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory 120 may include a non-transitory computer-readable storage medium. The memory 120 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as touch functionality, sound playback functionality, image playback functionality, etc.), instructions for implementing the various method embodiments described below, etc. The operating system may be the Android system, including systems deeply developed based on the Android system, the iOS system developed by Apple Inc., including systems deeply developed based on the iOS system, or other systems. The data storage area may also store data created by the electronic device during use, such as phonebook data, audio and video data, chat log data, etc.

[0154] See Figure 11 As shown, the memory 120 can be divided into operating system space and user space. The operating system runs in the operating system space, while native and third-party applications run in the user space. To ensure that different third-party applications can achieve good running performance, the operating system allocates corresponding system resources for each application. However, different application scenarios within the same third-party application have different requirements for system resources. For example, in local resource loading scenarios, third-party applications have high requirements for disk read speed; in animation rendering scenarios, third-party applications have high requirements for GPU performance. Since the operating system and third-party applications are independent of each other, the operating system often cannot promptly perceive the current application scenario of a third-party application, resulting in the operating system's inability to adapt system resources accordingly to the specific application scenario of the third-party application.

[0155] In order for the operating system to distinguish the specific application scenarios of third-party applications, it is necessary to establish data communication between the third-party applications and the operating system. This would allow the operating system to obtain the current scenario information of the third-party applications at any time, and then perform targeted system resource adaptation based on the current scenario.

[0156] Taking the Android operating system as an example, the programs and data stored in memory 120 are as follows: Figure 12As shown, the memory 120 can store the Linux kernel layer 320, the system runtime library layer 340, the application framework layer 360, and the application layer 380. The Linux kernel layer 320, system runtime library layer 340, and application framework layer 360 belong to the operating system space, while the application layer 380 belongs to the user space. The Linux kernel layer 320 provides low-level drivers for various hardware components of the electronic device, such as display drivers, audio drivers, camera drivers, Bluetooth drivers, Wi-Fi drivers, and power management. The system runtime library layer 340 provides support for key features of the Android system through several C / C++ libraries. For example, the SQLite library provides database support, the OpenGL / ES library provides 3D graphics support, and the Webkit library provides browser kernel support. The system runtime library layer 340 also provides the Android runtime library, which mainly provides core libraries that allow developers to write Android applications using the Java language. The Application Framework Layer 360 provides various APIs that may be used when building applications. Developers can also use these APIs to build their own applications, such as activity management, window management, view management, notification management, content provider, package management, call management, resource management, and location management. At least one application runs in the Application Layer 380. These applications can be native applications that come with the operating system, such as contacts, SMS, clock, and camera apps; or third-party applications developed by third-party developers, such as games, instant messaging, and photo editing apps.

[0157] Taking the operating system as an example (iOS), the programs and data stored in memory 120 are as follows: Figure 12As shown, the iOS system includes: Core OS layer 420, Core Services layer 440, Media layer 460, and Cocoa Touch layer 480. Core OS layer 420 includes the operating system kernel, drivers, and low-level program frameworks. These low-level program frameworks provide hardware-level functionality for use by the program frameworks located in Core Services layer 440. Core Services layer 440 provides system services and / or program frameworks required by applications, such as Foundation framework, account framework, advertising framework, data storage framework, network connectivity framework, geolocation framework, motion framework, etc. Media layer 460 provides applications with audiovisual interfaces, such as interfaces related to graphics and images, audio technology, video technology, and AirPlay (wireless playback of audio and video transmission technologies). Cocoa Touch layer 480 provides various commonly used interface-related frameworks for application development and is responsible for user touch interaction on electronic devices. Examples include local notification services, remote push services, advertising frameworks, game tool frameworks, message user interface (UI) frameworks, UIKit frameworks, map frameworks, and so on.

[0158] exist Figure 13 The framework shown includes, but is not limited to, the base framework in the core service layer 440 and the UIKit framework in the touchable layer 480. The base framework provides many basic object classes and data types, offering the most basic system services to all applications, and is independent of the UI. The UIKit framework, on the other hand, provides a basic UI class library for creating touch-based user interfaces. iOS applications can use the UIKit framework to provide their UI, thus providing the application's infrastructure for building user interfaces, drawing, handling user interaction events, responding to gestures, and so on.

[0159] The methods and principles for implementing data communication between third-party applications and the operating system in the iOS system can be referenced from the Android system, and will not be elaborated here.

[0160] The input device 130 is used to receive input instructions or data, and includes, but is not limited to, a keyboard, mouse, camera, microphone, or touch device. The output device 140 is used to output instructions or data, and includes, but is not limited to, a display device and a speaker. In one example, the input device 130 and the output device 140 can be combined into a touch screen, which is used to receive touch operations from the user using a finger, stylus, or any suitable object on or near it, and to display the user interface of various applications. The touch screen is usually located on the front panel of the electronic device. The touch screen can be designed as a full-screen, curved screen, or irregularly shaped screen. The touch screen can also be designed as a combination of a full-screen and a curved screen, or a combination of an irregularly shaped screen and a curved screen; this embodiment of the application does not limit this.

[0161] In addition, those skilled in the art will understand that the structure of the electronic device shown in the above figures does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements. For example, the electronic device may also include radio frequency circuits, input units, sensors, audio circuits, wireless fidelity (WiFi) modules, power supplies, Bluetooth modules, etc., which will not be described in detail here.

[0162] In the embodiments of this application, the executing entity for each step can be the electronic device described above. Optionally, the executing entity for each step is the operating system of the electronic device. The operating system can be Android, iOS, or other operating systems; this embodiment of the application does not limit this.

[0163] The electronic device in this embodiment may also be equipped with a display device, which can be various devices capable of display functions, such as: cathode ray tube display (CR), light-emitting diode display (LED), electronic ink screen, liquid crystal display (LCD), plasma display panel (PDP), etc. Users can use the display device on the electronic device 101 to view displayed text, images, videos, and other information. The electronic device may be a smartphone, tablet computer, gaming device, AR (Augmented Reality) device, automobile, data storage device, audio playback device, video playback device, laptop, desktop computing device, wearable device such as electronic watch, electronic glasses, electronic helmet, electronic bracelet, electronic necklace, electronic clothing, etc.

[0164] exist Figure 10 In the illustrated electronic device, which can be a terminal, the processor 110 can be used to call the application stored in the memory 120 and specifically perform the following operations:

[0165] Get the audio recording status corresponding to the stereo video;

[0166] If the audio recording status is in the audio missing state, then the stereo video is processed for audio-visual synchronization.

[0167] In one embodiment, when the processor 110 performs the step of acquiring the audio recording status corresponding to the stereo video, it specifically performs the following operations:

[0168] Obtain the recording timestamp of at least one frame of audio signal corresponding to the stereo video;

[0169] Based on the recording timestamps of each audio signal, the audio recording status for the stereo video is determined.

[0170] In one embodiment, before the processor 110 performs the step of acquiring the recording timestamp of at least one frame of audio signal corresponding to the stereo video, it further includes:

[0171] The recording timestamp of at least one frame of audio signal corresponding to the stereo video is read using the timestamp variable.

[0172] In one embodiment, when the processor 110 determines the audio recording status for the stereo video based on the recording timestamps of each of the audio signals, it specifically performs the following operations:

[0173] Based on the recording timestamps of each audio signal, calculate the recording interval time between two consecutive frames of the audio signal;

[0174] Based on the recording interval, the audio recording status for the stereo video is determined.

[0175] In one embodiment, when the processor 110 determines the audio recording status corresponding to the stereo video based on the recording interval time, it specifically performs the following operations:

[0176] If the recording interval is greater than the time threshold, then the audio recording status for the stereo video is determined to be an audio missing status.

[0177] If the recording interval is less than or equal to the time threshold, then the audio recording status for the stereo video is determined to be normal audio status.

[0178] In one embodiment, when the processor 110 performs the audio-visual synchronization processing on the stereo video, it specifically performs the following operations:

[0179] Obtain the audio missing time corresponding to the stereo video;

[0180] Based on the audio missing time, reference signal data is obtained, and the stereo video is processed for audio-visual synchronization based on the reference signal data.

[0181] In one embodiment, when the processor 110 performs the operation of acquiring reference signal data based on the audio missing time, it specifically performs the following operations:

[0182] Obtain the recording sampling rate corresponding to the stereo video, and determine the amount of missing data based on the audio missing time and the recording sampling rate;

[0183] Determine the reference signal data corresponding to the amount of missing data.

[0184] In one embodiment, when the processor 110 executes the reference signal data corresponding to the amount of missing data, it specifically performs the following operations:

[0185] Obtain the first audio data corresponding to the amount of missing data; or,

[0186] Obtain the first image data corresponding to the amount of missing data; or,

[0187] Based on the amount of missing data, a first data amount and a second data amount are determined, and second audio data corresponding to the first data amount and second image data corresponding to the second data amount are obtained.

[0188] In one embodiment, when the processor 110 performs the operation of determining the first data amount and the second data amount based on the amount of missing data, it specifically performs the following operations:

[0189] Obtain a frame extraction threshold, and determine a second data quantity based on the frame extraction threshold and the amount of missing data, wherein the second data quantity is less than or equal to the frame extraction threshold.

[0190] The first data quantity is determined based on the difference between the missing data quantity and the second data quantity.

[0191] In one embodiment, when the processor 110 performs audio-visual synchronization processing on the stereo video based on the reference signal data, it specifically performs the following operations:

[0192] If the reference signal data is the first audio data, then the stereo video is subjected to audio completion processing based on the first audio data;

[0193] If the reference signal data is first image data, then the stereo video is subjected to image frame extraction processing based on the first image data;

[0194] If the reference signal data is second audio data and second image data, then the stereo video is subjected to audio completion processing based on the second audio data and image frame extraction processing based on the second image data.

[0195] In one embodiment, when the processor 110 performs the operation of acquiring the first audio data corresponding to the amount of missing data, it specifically performs the following operations:

[0196] Generate first audio data corresponding to the amount of missing data based on the target audio features; or...

[0197] Determine the audio background type corresponding to the stereo video, and generate the first audio data corresponding to the audio missing time based on the audio background type.

[0198] In one embodiment, when the processor 110 performs the step of acquiring the audio recording status corresponding to the stereo video, it specifically performs the following operations:

[0199] Obtain the default audio recording system for the stereo video;

[0200] A target audio recording system corresponding to the stereo video is detected, and the audio recording system is used to record the audio signal corresponding to the stereo video.

[0201] If the audio recording system is not compatible with the default audio recording system, then the stereo video is determined to be in an audio missing state.

[0202] If the audio recording system matches the default audio recording system, then the stereo video is determined to be in a normal audio state.

[0203] In this embodiment, the terminal obtains the corresponding audio recording status during stereo video recording. If the audio recording status is an audio missing status, the terminal performs audio-visual synchronization processing on the stereo video. This ensures audio-visual synchronization in the event of audio loss in the stereo video, avoiding poor video recording quality caused by audio loss and improving the disaster recovery capability during stereo video recording. Furthermore, audio-visual synchronization can be achieved in multiple ways, enriching the audio-visual synchronization processing methods and enhancing the intelligence of video recording.

[0204] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory, or random access memory, etc.

[0205] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.

Claims

1. A stereo video recording method, characterized in that, The method includes: Get the audio recording status corresponding to the stereo video; If the audio recording status is audio missing, then the stereo video is processed for audio-visual synchronization. The step of performing audio-visual synchronization processing on the stereo video includes: The process involves: obtaining the audio missing time corresponding to the stereo video; determining the amount of missing data based on the audio missing time; determining a first data amount and a second data amount based on the amount of missing data; obtaining the second audio data corresponding to the first data amount and the second image data corresponding to the second data amount; performing audio completion processing on the stereo video based on the second audio data; and performing image frame extraction processing on the stereo video based on the second image data.

2. The method according to claim 1, characterized in that, The step of obtaining the audio recording status corresponding to the stereo video includes: Obtain the recording timestamp of at least one frame of audio signal corresponding to the stereo video; Based on the recording timestamps of each audio signal, the audio recording status for the stereo video is determined.

3. The method according to claim 2, characterized in that, The process of obtaining the recording timestamp of at least one frame of audio signal corresponding to the stereo video includes: The recording timestamp of at least one frame of audio signal corresponding to the stereo video is read using the timestamp variable.

4. The method according to claim 2, characterized in that, Determining the audio recording status for the stereo video based on the recording timestamps of each audio signal includes: Based on the recording timestamps of each audio signal, calculate the recording interval time between two consecutive frames of the audio signal; Based on the recording interval, the audio recording status for the stereo video is determined.

5. The method according to claim 4, characterized in that, Determining the audio recording status for the stereo video based on the recording interval includes: If the recording interval is greater than the time threshold, then the audio recording status for the stereo video is determined to be an audio missing status. If the recording interval is less than or equal to the time threshold, then the audio recording status for the stereo video is determined to be normal audio status.

6. The method according to claim 1, characterized in that, Determining the first data volume and the second data volume based on the missing data volume includes: Obtain a frame extraction threshold, and determine a second data quantity based on the frame extraction threshold and the amount of missing data, wherein the second data quantity is less than the frame extraction threshold. The first data quantity is determined based on the difference between the missing data quantity and the second data quantity.

7. The method according to claim 1, characterized in that, The step of obtaining the audio recording status corresponding to the stereo video includes: Obtain the default audio recording system for the stereo video; The target audio recording system corresponding to the stereo video is detected, and the target audio recording system is used to record the audio signal corresponding to the stereo video. If the target audio recording system does not match the default audio recording system, then the stereo video is determined to be in an audio missing state. If the target audio recording system matches the default audio recording system, then the stereo video is determined to be in a normal audio state.

8. A stereo video recording device, characterized in that, The device includes: The status acquisition module is used to acquire the audio recording status corresponding to the stereo video. The audio-visual synchronization module is used to perform audio-visual synchronization processing on the stereo video if the audio recording state is an audio missing state. Specifically, the audio-visual synchronization module is used to obtain the audio missing time corresponding to the stereo video; determine the amount of missing data based on the audio missing time; determine a first data amount and a second data amount based on the amount of missing data; obtain the second audio data corresponding to the first data amount and the second image data corresponding to the second data amount; perform audio completion processing on the stereo video based on the second audio data; and perform image frame extraction processing on the stereo video based on the second image data.

9. A computer storage medium, characterized in that, The computer storage medium stores a plurality of instructions, which are adapted to be loaded by a processor and executed as method steps as claimed in any one of claims 1 to 7.

10. An electronic device, characterized in that, include: A processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and executed the method steps as claimed in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Human-computer interaction type software screen recording method

    CN111105816A