Media file recording method, media file analysis method and related devices
By using the audio group and solid color frame images of the test film source during media file recording, the audio and video synchronization is automatically analyzed, and the problems of low efficiency and low accuracy in the existing technology are solved, and efficient and accurate audio and video synchronization analysis is achieved.
Patent Information
- Application Number
- CN202410076850.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-18
- Publication Date
- 2025-07-25
AI Technical Summary
In the prior art, the method of testing media files audio and video synchronization is inefficient and has low accuracy, and manual drag is required for audio and video synchronization analysis.
The test film source is used to include audio groups and solid color frame images arranged in equal distances. After recording the media file, the audio and video synchronization is automatically analyzed based on the energy value matching of the audio signal and video frame images without manual dragging.
It improves the efficiency and accuracy of audio and video synchronous analysis, and reduces the workload and accuracy fluctuations in manual operations.
Smart Images

Figure CN120378682A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of multimedia technology, and in particular to a method for recording a media file, a method for analyzing a media file, and related devices. Among them, the related devices include a media file recording device, a media file analysis device, and related devices, an electronic device, a storage medium, and a computer program product. Background Art
[0002] During the process of watching a video, it is necessary to have audio-visual synchronization, also known as lip-sync, so that there will be no situation of "words not matching the mouth", and in the case of subtitles in the video picture, there will also be no situation of the sound not matching the subtitles. Taking the scenario of video recording and broadcasting as an example, for example, in the education field, media files such as recorded teaching files are recorded and broadcast. After obtaining the media file through recording and broadcasting, in order to ensure the audio-visual synchronization of the recorded media file, it is necessary to test the audio-visual synchronization of the media file. However, the current method for testing the audio-visual synchronization of media files has the problem of low efficiency. Summary of the Invention
[0003] Based on this, it is necessary to provide a method for recording a media file, a method for analyzing a media file, and related devices that can improve the test efficiency of audio-visual synchronization of media files for the above technical problems.
[0004] In a first aspect, this application provides a method for recording a media file, and the method includes:
[0005] Obtain a test source video, where the test source video includes a test audio source and a test video source. Among them, the test audio source includes a plurality of identical audio groups arranged at equal distances, and each audio group includes at least one audio signal; the test video source includes a plurality of first solid-color frame images arranged at equal distances, and the other frame images of the test video source are second solid-color frame images. The first solid-color frame images are different from the second solid-color frame images, and the first solid-color frame images correspond to the audio signals one by one;
[0006] Play the test source video during the recording process, and obtain the recorded media file after the recording ends.
[0007] Based on the media file recording method of the embodiments of the present application described above, a test video source is played during the recording process, that is, the information of the test video source is recorded into the media file, and the test video source includes a test audio source and a test video source. The test video source includes first solid color frame images arranged at equal intervals, and other frame images are second solid color frame images. Moreover, the first solid color frame images correspond one-to-one with the audio signals in the audio groups arranged at equal intervals in the test audio source. Thus, after the media file is recorded accordingly, based on the one-to-one correspondence between the first solid color frame images and the audio signals, and the characteristics of the solid color frame images, the analysis of the audio-visual synchronization of the recorded media file can be realized, without the need for manual dragging for the analysis of audio-visual synchronization, and the efficiency is high.
[0008] In some embodiments, the audio group includes a first audio signal and a second audio signal spaced apart by a first preset distance, and the energy values of the first audio signal and the second audio signal are different from each other and have a preset relationship.
[0009] Thus, by setting two first audio signals and second audio signals with different energy values and having a preset relationship for each audio group, it is helpful to, after the media file is recorded accordingly, based on the one-to-one correspondence between the first solid color frame images and the audio signals, find one of the audio signals, and then combine the preset relationship and the first preset distance to check whether there is another audio signal, so as to verify the accuracy of the audio signal found based on the one-to-one correspondence, which is helpful to improve the accuracy of audio-visual synchronization analysis.
[0010] In some embodiments, the audio group further includes a third audio signal spaced apart from the second audio signal by a second preset distance, and the energy values of the first audio signal, the second audio signal, and the third audio signal are different from each other and have a preset relationship.
[0011] Thus, by setting three first audio signals, the second audio signal, and the third audio signal with different energy values and having a preset relationship for each audio group, it is helpful to, after the media file is recorded accordingly, based on the one-to-one correspondence between the first solid color frame images and the audio signals, find one of the audio signals, and then combine the preset relationship and the first preset distance and the second preset distance to check whether there are the other two audio signals, so as to verify the accuracy of the audio signal found based on the one-to-one correspondence, which is helpful to further improve the accuracy of audio-visual synchronization analysis.
[0012] In some embodiments, the audio signals in the audio group are instantaneous pulse signals.
[0013] Thus, by setting the audio signals in the audio group as instantaneous pulse signals, not only can a relatively high energy peak be generated for detection, but also other audios in other recorded files will not be affected, which can help improve the test efficiency while not affecting the performance of the media files obtained by recording.
[0014] In some embodiments, the maximum energy value among the energy values of the audio signals in the audio group is 1 db.
[0015] Thus, by setting the maximum energy value among the energy values of the audio signals in the audio group to 1 db, the energy value of the audio signal in the test audio source will not be too large to affect the performance of the media file, so as to further maintain the performance of the media file.
[0016] In some embodiments, the first solid-color frame image is a white image;
[0017] Since the Y component of a white image is 255 under YUV, which is the highest value of the Y component, it helps to actually find the first solid-color frame image based on the Y component of 255 in the subsequent process of analyzing the media file, and helps to improve the efficiency of the subsequent audio-visual synchronization analysis of the media file.
[0018] In some embodiments, the equal-distance interval between the audio groups includes 10 seconds.
[0019] In some embodiments, the first preset distance is the same as the second preset distance.
[0020] In some embodiments, the first preset interval distance includes 1 second.
[0021] In some embodiments, the energy value of the second audio signal is greater than the energy values of the first audio signal and the third audio signal, or the energy values of the first audio signal, the second audio signal, and the third audio signal increase in sequence; or the energy values of the first audio signal, the second audio signal, and the third audio signal decrease in sequence.
[0022] In some embodiments, the energy values of the first audio signal, the second audio signal, and the third audio signal are 0.5M, 1 db, and 0.75M respectively.
[0023] In a second aspect, the present application provides a media file analysis method, and the method includes:
[0024] Decoding the media file to be analyzed into an audio sequence and a video frame sequence;
[0025] Dividing the audio sequence into several audio segments according to a preset audio duration, and calculating the audio energy value of the audio segments;
[0026] Calculate the video energy value of the video frames in the video frame sequence;
[0027] When the video energy value is the same as the preset video energy value, within a preset time range of the video timestamp of the video frame where the video energy value is the same as the preset video energy value, search for an audio energy value that matches the preset audio energy value;
[0028] Based on the audio timestamp corresponding to the matched audio energy value and the video timestamp, determine the audio-visual synchronization situation of the media file.
[0029] Based on the media file analysis method of the embodiment of the present application as described above, after decoding the media file to be analyzed into an audio sequence and a video frame sequence, the audio sequence is divided into several audio segments to calculate the audio energy value of each audio segment, and the video energy value of the video frames in the video frame sequence is calculated. When a video energy value that is the same as the preset video energy value is found, an audio energy value that matches the video energy value is further searched for. Then, in combination with the video timestamp corresponding to the video energy value and the audio timestamp corresponding to the audio energy value, the audio-visual synchronization situation of the media file is determined. During the analysis process, there is no need for manual dragging, and the audio-visual synchronization can be analyzed by combining the audio energy value of the audio signal and the video energy value of the video signal, with high efficiency.
[0030] In some embodiments, the determining the audio-visual synchronization situation of the media file based on the audio timestamp corresponding to the matched audio energy value and the video timestamp includes:
[0031] In the preset time direction of the audio timestamp of the matched audio energy value, search for whether there is an audio energy value that is separated from the matched audio energy value by a preset distance and has a preset relationship;
[0032] If it exists, based on the audio timestamp corresponding to the matched audio energy value and the video timestamp, determine the audio-visual synchronization situation of the media file.
[0033] Thus, further search for whether there is an audio energy value that is separated from the matched audio energy value by a preset distance and has a preset relationship in the preset time direction of the audio timestamp of the matched audio energy value, and only if it exists, based on the audio timestamp corresponding to the matched audio energy value and the video timestamp, determine the audio-visual synchronization situation of the media file. Thus, it is possible to determine that the found matched audio energy value is the correct audio energy value based on the found audio energy value with a preset relationship, so as to further improve the accuracy of the audio-visual synchronization analysis of the media file.
[0034] In some embodiments, searching for whether there is an audio energy value that is at a preset distance from the matched audio energy value and has a preset relationship in the preset time direction of the audio timestamp of the matched audio energy value includes at least one of the following:
[0035] Searching for whether there is an audio energy value that is at a first preset distance from the matched audio energy value and has a first preset relationship in the first preset time direction of the audio timestamp of the matched audio energy value;
[0036] Searching for whether there is an audio energy value that is at a second preset distance from the matched audio energy value and has a second preset relationship in the second preset time direction of the audio timestamp of the matched audio energy value;
[0037] Searching for whether there is an audio energy value that is at a third preset distance from the matched audio energy value and has a third preset relationship in the first preset time direction of the audio timestamp of the matched audio energy value, where the third preset distance is greater than the first preset distance;
[0038] Searching for whether there is an audio energy value that is at a fourth preset distance from the matched audio energy value and has a fourth preset relationship in the second preset time direction of the audio timestamp of the matched audio energy value, where the fourth preset distance is greater than the second preset distance.
[0039] In some embodiments, it includes at least one of the following:
[0040] The duration of the audio segment is 20 milliseconds;
[0041] The video energy value of the video frame is the Y component in the YUV color space of the video frame;
[0042] The preset video energy value is 255.
[0043] In a third aspect, the present application also provides a media file recording device. The device includes:
[0044] A test source acquisition module, configured to acquire a test source, where the test source includes a test audio source and a test video source. Among them, the test audio source includes several identical audio groups arranged at equal distances, and each audio group includes at least one audio signal; the test video source includes several first solid-color frame images arranged at equal distances, and the other frame images of the test video source are second solid-color frame images. The first solid-color frame images are different from the second solid-color frame images, and the first solid-color frame images correspond to the audio signals one by one;
[0045] A recording module, configured to play the test video source during the recording process and obtain the recorded media file after the recording ends.
[0046] In a fourth aspect, the present application further provides a media file analysis device. The device includes:
[0047] A decoding module, configured to decode the media file to be analyzed into an audio sequence and a video frame sequence;
[0048] An audio energy calculation module, configured to divide the audio sequence into several audio segments with a preset audio duration and calculate the audio energy values of the audio segments;
[0049] A video energy calculation module, configured to calculate the video energy values of the video frames in the video frame sequence;
[0050] A matching search module, configured to, when the video energy value is the same as a preset video energy value, search for an audio energy value that matches a preset audio energy value within a preset time range of the video timestamps of the video frames with the same video energy value as the preset video energy value;
[0051] A synchronization determination module, configured to determine the audio-visual synchronization situation of the media file based on the audio timestamps corresponding to the matched audio energy values and the video timestamps.
[0052] In a fifth aspect, the present application further provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method in any of the above embodiments are implemented.
[0053] In a sixth aspect, the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method in any of the above embodiments are implemented.
[0054] In a seventh aspect, the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the method in any of the above embodiments are implemented. Description of the Drawings
[0055] Figure 1 It is an application environment diagram of the media file recording method and the media file analysis method in an embodiment of the present application;
[0056] Figure 2 It is a flowchart of the media file recording method in an embodiment;
[0057] Figure 3Schematic diagram of the principle of the test piece source in one embodiment;
[0058] Figure 4 Schematic diagram of the principle of the test piece source in another embodiment;
[0059] Figure 5 Schematic diagram of the principle of the test piece source in another embodiment;
[0060] Figure 6 Schematic diagram of the principle of the test piece source in another embodiment;
[0061] Figure 7 Schematic flowchart of the media file analysis method in another embodiment;
[0062] Figure 8 Schematic flowchart for determining the audio-visual synchronization of a media file based on the audio timestamps and video timestamps corresponding to the matched audio energy values in one embodiment;
[0063] Figure 9 Schematic diagrams of the audio-visual synchronization of media files in some specific examples;
[0064] Figure 10 Block diagram of the structure of a media file recording device in one embodiment;
[0065] Figure 11 Block diagram of the structure of a media file analysis device in one embodiment;
[0066] Figure 12 Internal structure diagram of a computer device in one embodiment. Detailed implementation manners
[0067] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.
[0068] The embodiments of the technical solutions of the present application will be described in detail below with reference to the accompanying drawings. The following embodiments are only used to illustrate the technical solutions of the present application more clearly, and thus are only examples and cannot be used to limit the protection scope of the present application.
[0069] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion.
[0070] In the description of the embodiments of the present application, technical terms such as "first" and "second" are only used to distinguish different objects, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity, specific order, or primary-secondary relationship of the indicated technical features. In the description of the embodiments of the present application, the meaning of "a plurality" is more than two, unless otherwise specifically defined.
[0071] Referring to "embodiment" herein means that a specific feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the present application. The phrase appearing in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0072] In the description of the embodiments of the present application, the term "and / or" is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0073] In the description of the embodiments of the present application, the term "a plurality" refers to more than two (including two). Similarly, "a plurality of groups" refers to more than two groups (including two groups), and "a plurality of pieces" refers to more than two pieces (including two pieces).
[0074] Currently, for the scenario of testing the audio-visual synchronization of media files, in order to accurately test the audio-visual synchronization of media files, the existing implementation is to prepare a customized source video. Through this source video, the audio-visual synchronization of media files is specifically tested. Among them, the image of this source video has scales, and the scales are symmetrically arranged on the left and right. There is a cursor below the scales, and the cursor can move left and right. During the playback of this source video, a prompt sound will be emitted when the cursor passes the scale, for example, a beep, and there will be no sound at other positions. Based on this source video, during the process of recording a media file, the source video is played through a source video playback device and undergoes video acquisition / audio acquisition through a recording and broadcasting system to record a media file (such as an MP4 (a multimedia file format using MPEG-4) file). It can be seen that the cursor and the prompt sound of this source video will be recorded into this media file at the same time. Then, the media file is played through a media file playback device, and by dragging the cursor, it is manually observed where the sound is heard when the cursor is at a certain position. If a beep sound is emitted when the cursor is in the middle, it indicates that the audio and video are synchronized. If a beep sound is emitted when the cursor is on the left, it means that the sound lags behind the video. For example, first, the lips are seen to move, and then the words are heard. If a beep sound is emitted when the cursor is on the right, it means that the sound is ahead of the video. For example, the sound is heard first, and then the lips are seen to move.
[0075] In this method, the recorded mp4 file needs to be carefully dragged manually using a media file player, which is laborious and inefficient. Moreover, based on manual dragging, there are usually large fluctuations in accuracy, resulting in low accuracy.
[0076] Based on this, through research, it is found that it is possible to achieve the analysis of audio-visual synchronization of media files without manual dragging. By using a test source file containing a test audio source and a test video source, the one-to-one correspondence between the test audio source and the test video source, and the audio source includes several identical audio groups arranged at equal distances, and the test video source contains several first solid-color frame images arranged at equal distances. Since the image energy value of the solid-color frame image is constant, and after the audio signal of the audio group of the audio source is determined, its audio energy value is also constant. Therefore, after recording a media file based on this test source file, by combining the energy value of the audio signal and the energy value of the video signal, the analysis of the audio-visual synchronization of the media file can be realized without manual viewing and dragging, so as to improve the efficiency of the audio-visual synchronization analysis of the media file.
[0077] The embodiments of the present application provide a media file recording method and a media file analysis method, which can be applied to an application environment as Figure 1 shown. Among them, the source file playback device 10 is used to play the test source file, and the media file recording device 20 records during the process of the source file playback device 10 playing the test source file to obtain the recorded media file. After obtaining the recorded media file, the media file analysis device 30 analyzes the recorded media file to determine the audio-visual synchronization situation of the recorded media file. Among them, the source file playback device 10 and the media file recording device 20 can be different devices or the same device. The media file recording device 20 and the media file analysis device 30 can be different devices or the same device. The data storage system can store the data required or to be processed during the media file recording and media file analysis processes, such as the test source file, the recorded media file, the results obtained by analyzing the media file, etc. The data storage system can be integrated on the server, or placed in the cloud or other network servers, or can be set on one or more of the source file playback device 10, the media file recording device 20, and the media file analysis device 30. Among them, the source file playback device 10, the media file recording device 20, and the media file analysis device 30 can be, but are not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server can be implemented by an independent server or a server cluster composed of multiple servers.
[0078] In some embodiments, such asFigure 2 As shown, a method for recording a media file is provided. Taking the media file recording device 20 in Figure 1 as an example, the method includes the following steps:
[0079] Step S201: Obtain a test source, where the test source includes a test audio source and a test video source. Among them, the test audio source includes several identical audio groups arranged at equal distances, and each audio group includes at least one audio signal; the test video source includes several first solid-color frame images arranged at equal distances, and other frame images of the test video source are second solid-color frame images. The first solid-color frame images are different from the second solid-color frame images, and the first solid-color frame images correspond one-to-one with the audio signals.
[0080] Among them, the test audio source is an audio source that contains audio signals. The test audio source contains several audio groups arranged at equal distances.
[0081] Refer to Figure 3 As shown, the test audio source contains several identical audio groups A arranged at equal distances, such as the audio groups A at times T1, T2, and T3, and the interval between T1 and T2 is the same as the interval between T2 and T3. In the test video source, first solid-color frame images are only set at positions corresponding to the audio groups of the test audio source, while second solid-color frame images are set at other positions. As Figure 3 shown, first solid-color frame images are only set at the spaced-apart times T1, T2, T3... while second solid-color frame images are set at other times.
[0082] Among them, the solid-color frame image refers to a solid-color image, and the pixel values of each pixel point in the solid-color image are exactly the same.
[0083] Step S202: Play the test source during the recording process and obtain the recorded media file after the recording ends.
[0084] Based on the media file recording method of the embodiment of the present application as described above, the test source is played during the recording process, that is, the information of the test source is recorded into the media file. The test source includes a test audio source and a test video source. The test video source includes first solid-color frame images arranged at equal distances, and other frame images are second solid-color frame images. The first solid-color frame images correspond one-to-one with the audio signals in the audio groups arranged at equal distances in the test audio source. Therefore, after the media file is recorded accordingly, based on the one-to-one correspondence between the first solid-color frame images and the audio signals, and the characteristics of the solid-color frame images, the audio-visual synchronization analysis of the recorded media file can be realized, without manual dragging for audio-visual synchronization analysis, and the efficiency is high.
[0085] Among them, in the above test audio source, the number of audio signals included in each audio group is not limited. In some embodiments, each audio group may include only one audio signal. In some embodiments, each audio group may also include more than two audio signals.
[0086] In some embodiments, each audio group may include a first audio signal and a second audio signal spaced apart by a first preset distance, and the energy values of the first audio signal and the second audio signal are different from each other and have a preset relationship.
[0087] Among them, the magnitude relationship between the energy values of the first audio signal and the second audio signal, and the corresponding relationship with the first solid color frame image are not limited. Several of the following methods will be used as examples for illustration.
[0088] Reference Figure 4 As shown, each audio group may include two audio signals: a first audio signal s1 and a second audio signal s2, and the first audio signal s1 and the second audio signal s2 are spaced apart by a first preset distance, and the audio energy value of the first audio signal s1 is less than the audio energy value of the second audio signal s2. It should be understood that in other embodiments, it may also be that the audio energy value of the first audio signal s1 is greater than the audio energy value of the second audio signal s2. The embodiments of the present application do not make specific limitations in this regard.
[0089] Figure 4 As shown, the second audio signal s2 in the audio group corresponds to the first solid color frame image. For example, in the test video source, at times t12, t22, t32, etc. when the second audio signal s2 appears, the first solid color frame image appears, and at other times, it is the second solid color frame image. It should be understood that in other embodiments, it may also be that the first audio signal s1 in the audio group corresponds to the first solid color frame image. For example, in the test video source, at times t11, t21, t31, etc. when the first audio signal s1 appears, the first solid color frame image appears, and at other times, it is the second solid color frame image.
[0090] Thus, by setting a first audio signal and a second audio signal with different energy values and a preset relationship for each audio group, it is helpful to, after recording and obtaining a media file accordingly, based on the one-to-one correspondence between the first solid color frame image and the audio signal, find one of the audio signals, and then combine the preset relationship and the first preset distance to check whether there is another audio signal, so as to verify the accuracy of the audio signal found based on the one-to-one correspondence, which helps to improve the accuracy of audio-visual synchronization analysis.
[0091] In some other embodiments, there may be three audio signals in each audio group, that is, each audio group may include a first audio signal and a second audio signal spaced apart by a first preset distance, and a third audio signal spaced apart from the second audio signal by a second preset distance. The energy values of the first audio signal, the second audio signal, and the third audio signal are different from each other and have a preset relationship.
[0092] Among them, the magnitude relationship of the energy values of the first audio signal, the second audio signal, and the third audio signal, and the corresponding relationship with the first solid-color frame image are not limited. Several of the following methods will be used as examples for illustration.
[0093] Reference Figure 5 As shown, each audio group may contain three audio signals: a first audio signal S1, a second audio signal S2, and a third audio signal S3. The first audio signal S1 and the second audio signal S2 are spaced apart by a first preset distance, and the second audio signal S2 and the third audio signal S3 are spaced apart by a second preset distance. Among them, the first preset distance and the second preset distance may be the same or different. In the embodiments of the present application, the case where the first preset distance and the second preset distance are the same will be used as an example for illustration. Among them, in some embodiments, the first preset distance and the second preset distance may include 1 second. In this case, the equal-distance interval between audio groups may be set to include 10 seconds.
[0094] Thus, by setting three first audio signals, the second audio signal, and the third audio signal with different energy values and having a preset relationship for each audio group, it helps to, after recording a media file accordingly, based on the one-to-one correspondence between the first solid-color frame image and the audio signal, find one of the audio signals, and then combine the preset relationship, the first preset distance, and the second preset distance to check whether there are the other two audio signals, so as to verify the accuracy of the audio signal found based on the one-to-one correspondence, which helps to further improve the accuracy of audio-visual synchronization analysis.
[0095] Among them, the audio energy values of the first audio signal S1, the second audio signal S2, and the third audio signal S3 are not limited, as long as they can have a certain preset relationship.
[0096] Such as Figure 5As shown, the audio energy value of the first audio signal S1 is less than that of the second audio signal S2, and the audio energy value of the second audio signal S2 is greater than that of the third audio signal S3. Among them, the manner in which the audio energy value of the first audio signal S1 is less than that of the second audio signal S2 is not limited. For example, the audio energy value of the first audio signal S1 is the difference between the audio energy value of the second audio signal S2 and a certain first predetermined value, or the audio energy value of the first audio signal S1 is a first predetermined ratio of the audio energy value of the second audio signal S2, and the first predetermined ratio is less than 1. The manner in which the audio energy value of the second audio signal S2 is greater than that of the third audio signal S3 is not limited. For example, the audio energy value of the third audio signal S3 is the difference between the audio energy value of the second audio signal S2 and a certain second predetermined value, or the audio energy value of the third audio signal S3 is a second predetermined ratio of the audio energy value of the second audio signal S2, and the second predetermined ratio is less than 1. Among them, the first predetermined value and the second predetermined value may be the same or different, and the first predetermined ratio and the second predetermined ratio may be the same or different.
[0097] In other embodiments, the relationship of the audio energy values of the first audio signal S1, the second audio signal S2, and the third audio signal S3 can also be other settings. For example, the audio energy values of the first audio signal S1, the second audio signal S2, and the third audio signal S3 increase or decrease in sequence.
[0098] Taking the audio energy values of the first audio signal S1, the second audio signal S2, and the third audio signal S3 increasing in sequence as an example, it can be that the audio energy value of the second audio signal S2 is greater than that of the first audio signal S1 by a certain preset value or a certain preset ratio, and the audio energy value of the third audio signal S3 is greater than that of the second audio signal S2 by another preset value or another preset ratio, and these two preset values or preset ratios may be the same or different.
[0099] Taking the audio energy values of the first audio signal S1, the second audio signal S2, and the third audio signal S3 decreasing in sequence as an example, it can be that the audio energy value of the first audio signal S1 is greater than that of the second audio signal S2 by a certain preset value or a certain preset ratio, and the audio energy value of the second audio signal S2 is greater than that of the third audio signal S3 by another preset value or another preset ratio, and these two preset values or preset ratios may be the same or different.
[0100] Figure 5As shown, the second audio signal S2 in the audio group corresponds to the first solid color frame image. For example, in a test video source, at times t12, t22, t32, etc. when the second audio signal S2 appears, the first solid color frame image appears, and at other times, it is the second solid color frame image. It should be understood that in other embodiments, it may also be that the first audio signal S1 in the audio group corresponds to the first solid color frame image. For example, in a test video source, at times t11, t21, t31, etc. when the first audio signal S1 appears, the first solid color frame image appears, and at other times, it is the second solid color frame image; it may also be that the third audio signal S3 in the audio group corresponds to the first solid color frame image. For example, in a test video source, at times t13, t23, t33, etc. when the third audio signal S3 appears, the first solid color frame image appears, and at other times, it is the second solid color frame image.
[0101] It should be understood that in the description of the above embodiments, the one-to-one correspondence between the first solid color frame image and the audio signal is illustrated by taking the time stamps of the first solid color frame image and the audio signal being exactly the same as an example. In actual settings, the time stamps of the first solid color frame image and the audio signal may not be exactly the same, but may be separated by a specified duration. The embodiments of the present application do not make specific limitations on this.
[0102] In some embodiments, in the above embodiments, the audio signals in the audio group are instantaneous pulse signals.
[0103] An instantaneous pulse signal is an electrically-signal that changes instantaneously. Its characteristic is that within an extremely short time, the voltage or current shows a very rapid change, so that an extremely high energy peak can be generated within an extremely short time (such as at the nanosecond or microsecond level). Thus, by setting the audio signals in the audio group as instantaneous pulse signals, not only can a relatively high energy peak be generated to be detected, but it will not affect the other audio of other recorded files, which can help improve the test efficiency while not affecting the performance of the media files obtained by recording.
[0104] In some embodiments, the maximum energy value among the energy values of the audio signals in the audio group is 1 db.
[0105] Thus, by setting the maximum energy value among the energy values of the audio signals in the audio group to 1 db, the energy value of the audio signals in the test audio source will not be too large to affect the performance of the media files, so as to further maintain the performance of the media files.
[0106] Among them, when the maximum energy value among the energy values of the audio signals in the audio group is set to 1 db, the energy values of the other audio signals in the audio group can be set to a predetermined proportion of 1 db. Such as Figure 5 、 6As shown, the audio energy value of the second audio signal S2 is 1 db. The audio energy value of the first audio signal S1 can be half of the audio energy value of the second audio signal S2, that is, 0.5 db, while the audio energy value of the third audio signal S3 can be three - quarters of the audio energy value of the second audio signal S2, that is, 0.75 db. It should be understood that in other embodiments, different settings can also be made. For example, the audio energy value of the first audio signal S1 is 0.75 db, the audio energy value of the third audio signal S2 is 0.5 db, or set to other values.
[0107] Among them, in the above - mentioned embodiments, the solid - color type of the first solid - color frame image is not limited, as long as it can be easily distinguished from other frame images. In the embodiments of the present application, the first solid - color frame image can be a white image, or an image with a Y component of 255 in YUV (a color - coding space). The type of the second solid - color frame image is also not limited, as long as it does not affect the recording and display of other images. In the embodiments of the present application, the second solid - color frame image can be a black image or an image with a Y component of 0 in YUV (a color - coding space).
[0108] Since the white image has a Y component of 255 in YUV, which is the highest value of the Y component, it helps to actually find the first solid - color frame image based on the Y component of 255 during the subsequent analysis of the media file, and helps to improve the efficiency of the audio - video synchronization analysis of the subsequent media file.
[0109] In some embodiments, as Figure 7 shown, a media - file analysis method is provided. Taking the media - file analysis device 30 in Figure 1 as an example for illustration, it includes the following steps:
[0110] Step S701: Decode the media file to be analyzed into an audio sequence and a video - frame sequence.
[0111] Among them, the media file to be analyzed refers to the file whose audio - video synchronization situation needs to be analyzed. In the embodiments of the present application, the media file to be analyzed can be a media file recorded based on the media - file recording method in the above - mentioned embodiments.
[0112] Among them, the method of decoding the media file to be analyzed into an audio sequence and a video - frame sequence is not limited, as long as it can decode the media file to be analyzed into an audio sequence and a video - frame sequence.
[0113] Step S702: Divide the audio sequence into several audio segments with a preset audio duration, and calculate the audio energy value of the audio segments.
[0114] The specific value of the preset audio duration is not limited. Based on the example of the test video source described above, the preset audio duration can be less than the duration between two adjacent first solid-color frame images in the test video source.
[0115] In some embodiments, the preset audio duration can also be determined in combination with the resolution and frame smoothness of the video. Taking a 1080P30 video as an example, the playback duration of one frame image is approximately 33 milliseconds. Therefore, the preset audio duration can be set to less than 33 milliseconds, such as 20 milliseconds, or can be set to equal 33 milliseconds, or can be greater than 33 milliseconds. The embodiments of the present application do not limit this.
[0116] The method for calculating the audio energy value of the audio segment is not limited, as long as it can calculate the energy value of the audio signal.
[0117] Step S703: Calculate the video energy value of the video frames in the video frame sequence.
[0118] When calculating the video energy value of the video frame, it can be calculated for each video frame in the video frame sequence. The method for calculating the video energy value of the video frame is not limited. In the embodiments of the present application, the Y component in the YUV color space of the video frame can be determined as the video energy value of the video frame.
[0119] Step S704: When the video energy value is the same as the preset video energy value, within the preset time range of the video timestamp of the video frame whose video energy value is the same as the preset video energy value, search for the audio energy value that matches the preset audio energy value.
[0120] Among them, the preset video energy value can be the energy value of the first solid-color frame image in the test video source described above. Taking the first solid-color frame as a white image, when the Y component in the YUV color space of the video frame is determined as the video energy value of the video frame, the preset video energy value can be 255.
[0121] When the calculated video energy value is the same as the preset video energy value, it indicates that the video frame is the first solid-color frame image. Therefore, within the preset time range of the video timestamp of this video frame, search for the audio energy value that matches the preset audio energy value. If a matching audio energy value can be found, it indicates that the audio signal corresponding to this video frame has been found, and then enter step S705.
[0122] Step S705: Based on the audio timestamp corresponding to the matching audio energy value and the video timestamp, determine the audio-visual synchronization situation of the media file.
[0123] When determining the audio-visual synchronization of a media file based on the audio timestamps and video timestamps corresponding to the matched audio energy values, it can be determined in combination with the one-to-one correspondence between the audio signal and the first solid-color frame image set in the test source video. Taking the case where the timestamps of the audio signal and the first solid-color frame image are exactly the same as an example, it can be determined that the audio and video of the media file are synchronized when the audio timestamp is the same as the video timestamp, and it can be determined that the audio of the media file is earlier than the video when the audio timestamp is earlier than the video timestamp.
[0124] Based on the media file analysis method of the embodiment of the present application as described above, after decoding the media file to be analyzed into an audio sequence and a video frame sequence, the audio sequence is divided into several audio segments to calculate the audio energy values of each audio segment, and the video energy values of the video frames in the video frame sequence are calculated. In the case of finding a video energy value that is the same as the preset video energy value, further search for the audio energy value that matches the video energy value, and then combine the video timestamp corresponding to the video energy value and the audio timestamp corresponding to the audio energy value to determine the audio-visual synchronization of the media file. There is no need for manual dragging during the analysis process, and the audio-visual synchronization can be analyzed by combining the audio energy value of the audio signal and the video energy value of the video signal, with high efficiency.
[0125] In some embodiments, referring to Figure 8 as shown, determining the audio-visual synchronization of the media file based on the audio timestamp corresponding to the matched audio energy value and the video timestamp in step S705 above may include:
[0126] Step S7051: Search in the preset time direction of the audio timestamp of the matched audio energy value to find out whether there is an audio energy value that is at a preset distance from the matched audio energy value and has a preset relationship.
[0127] Among them, the preset time direction of the audio timestamp of the matched audio energy value can be determined based on the correspondence between the audio signal and the first solid-color frame image in the audio group in the test source video as described above.
[0128] Step S7052: In the case where it exists, determine the audio-visual synchronization of the media file based on the audio timestamp corresponding to the matched audio energy value and the video timestamp.
[0129] Therefore, further search in the preset time direction of the audio timestamp of the matched audio energy value to find out whether there is an audio energy value that is at a preset distance from the matched audio energy value and has a preset relationship. And only when such an audio energy value exists, determine the audio-visual synchronization situation of the media file based on the audio timestamp corresponding to the matched audio energy value and the video timestamp. Thus, it is possible to determine that the found matched audio energy value is the correct one based on the found audio energy value with the preset relationship, so as to further improve the accuracy of the audio-visual synchronization analysis of the media file.
[0130] Among them, the method of searching in the preset time direction of the audio timestamp of the matched audio energy value to find out whether there is an audio energy value that is at a preset distance from the matched audio energy value and has a preset relationship in step S7051 is not limited. The following takes several methods as examples for illustration.
[0131] In some embodiments, it is possible to search in the first preset time direction of the audio timestamp of the matched audio energy value to find out whether there is an audio energy value that is at a first preset distance from the matched audio energy value and has a first preset relationship.
[0132] Take Figure 4 the shown test video source as an example. Since the second audio signal in the audio group corresponds to the first solid-color frame image, the video timestamp of the video frame whose determined video energy value is the same as the preset video energy value may be timestamp t12, t22 or t32 in an ideal state. Therefore, the audio timestamp of the matched audio energy value is also ideally t12, t22 or t32, that is, the same as the video timestamp. Thus, further search on the left side of t12, t22 or t32 (i.e., the preset time direction is the earlier time direction of the audio timestamp) to find out whether there is an audio energy value that is at a preset distance from the matched audio energy value and has a preset relationship.
[0133] In some embodiments, it is possible to search in the second preset time direction of the audio timestamp of the matched audio energy value to find out whether there is an audio energy value that is at a second preset distance from the matched audio energy value and has a second preset relationship.
[0134] Take the example that the first audio signal in the audio group of the test audio source of the test video source corresponds to the first solid-color frame image. After determining the audio timestamp of the matched audio energy value, it is possible to further search on the right side of this audio timestamp (i.e., the preset time direction is the later time direction of the audio timestamp) to find out whether there is an audio energy value that is at a preset distance from the matched audio energy value and has a preset relationship.
[0135] In some embodiments, it is possible to search in the first preset time direction of the audio timestamp of the matched audio energy value to find out whether there is an audio energy value that is at a first preset distance from the matched audio energy value and has a first preset relationship, and to search in the first preset time direction of the audio timestamp of the matched audio energy value to find out whether there is an audio energy value that is at a third preset distance from the matched audio energy value and has a third preset relationship, where the third preset distance is greater than the first preset distance.
[0136] In this case, the audio group of the test audio source of the test source video can include three audio signals, and the third audio signal corresponds to the first solid color frame image. After determining the audio timestamp of the matched audio energy value, it is possible to further search on the left side of this audio timestamp (i.e., the preset time direction is the earlier time direction of the audio timestamp) to find out whether there is an audio energy value that is at a first preset distance from the matched audio energy value and has a first preset relationship, and whether there is an audio energy value that is at a third preset distance from the matched audio energy value and has a third preset relationship. Among them, the degree to which the third preset distance is greater than the first preset distance can be set according to actual technical needs. For example, the third preset distance can be set to twice the first preset distance.
[0137] In some embodiments, it is possible to search in the second preset time direction of the audio timestamp of the matched audio energy value to find out whether there is an audio energy value that is at a second preset distance from the matched audio energy value and has a second preset relationship, and to search in the second preset time direction of the audio timestamp of the matched audio energy value to find out whether there is an audio energy value that is at a fourth preset distance from the matched audio energy value and has a fourth preset relationship, where the fourth preset distance is greater than the second preset distance.
[0138] In this case, the audio group of the test audio source of the test source video can include three audio signals, and the first audio signal corresponds to the first solid color frame image. After determining the audio timestamp of the matched audio energy value, it is possible to further search on the right side of this audio timestamp (i.e., the preset time direction is the later time direction of the audio timestamp) to find out whether there is an audio energy value that is at a second preset distance from the matched audio energy value and has a second preset relationship, and whether there is an audio energy value that is at a fourth preset distance from the matched audio energy value and has a fourth preset relationship. Among them, the degree to which the fourth preset distance is greater than the second preset distance can be set according to actual technical needs. For example, the fourth preset distance can be set to twice the fourth preset distance.
[0139] In some embodiments, it is possible to search in the first preset time direction of the audio timestamp of the matched audio energy value to find out whether there is an audio energy value that is separated from the matched audio energy value by a first preset distance and has a first preset relationship, and to search in the second preset time direction of the audio timestamp of the matched audio energy value to find out whether there is an audio energy value that is separated from the matched audio energy value by a second preset distance and has a second preset relationship.
[0140] Taking Figure 5 the test video source shown as an example, since the second audio signal in the audio group corresponds to the first solid-color frame image, the video timestamp of the video frame whose determined video energy value is the same as the preset video energy value may be timestamp t12, t22 or t32 in an ideal state. Therefore, the audio timestamp of the matched audio energy value is also ideally t12, t22 or t32, that is, the same as the video timestamp. Thus, further search is carried out on the left side of t12, t22 or t32 (i.e., the preset time direction is the earlier time direction of the audio timestamp), and on the right side of t12, t22 or t32 (i.e., the preset time direction is the later time direction of the audio timestamp) to find out whether there is an audio energy value that is separated from the matched audio energy value by a preset distance and has a preset relationship. Combining Figure 6 as shown, the matched audio energy value is 1 db. If an audio energy value of 0.5 db is found at a first preset distance on the left side and an audio energy value of 0.75 db is found at a second preset distance on the right side, it can be determined that there is an audio energy value that is separated from the matched audio energy value by a preset distance and has a preset relationship.
[0141] Based on the above-described embodiments, the following will be exemplified by combining the media file recording method and examples of media methods in some specific examples.
[0142] In this example, before executing the media file recording method, a test video source as shown in Figure 6 can be made first. The test video source includes a test audio source and a test video source.
[0143] Among them, the test audio source includes several identical audio groups arranged at equal distances, and the distance between each audio group is 10 seconds. Each audio group includes three audio signals, and the audio energy values of the three audio signals are 0.5 db, 1 db and 0.75 db respectively. The distance between any two adjacent audio signals of these three audio signals is 1 second. At this time, the distance between each audio group is 10 seconds, which can be the distance between the 1 db audio signals of each audio group is 10 seconds. Among them, the middle 1 db audio signal is corresponded to the first solid-color frame image in the test video source, so as to facilitate finding the left and right (lagging or leading).
[0144] The test video source includes a number of first solid-color frame images arranged at equal intervals, and the first solid-color frame images are white images, and the white images correspond to the 1db audio signal in the audio group.
[0145] After obtaining the test video source, play the test video source to the recording device (i.e., the media file recording device). During the recording process, the recording device records the test video source at the same time, and the recorded media file can be obtained after the recording ends.
[0146] After obtaining the recorded media file, the audio-visual synchronization situation of the recorded media file can be analyzed. When analyzing, first decode the media file to be analyzed into an audio sequence and a video frame sequence.
[0147] For the audio sequence, divide the audio sequence into several audio segments with a duration of 20 seconds, and calculate the audio energy value of each audio segment.
[0148] For the video frame sequence, calculate the video energy value of the video frames in the video frame sequence. Specifically, use the Y component in the YUV color space of the video frame as the video energy value of the video frame.
[0149] Find the video frame with the video energy value being the video energy peak value (i.e., the preset video energy value, that is, the Y component in the YUV color space is 255). Around the video timestamp of this video frame, search for the audio segment that causes the corresponding audio energy value to be 1db. If an audio segment with an audio energy value of 1db is found, further search whether there is an audio segment with an audio energy value of 0.5db on the left side of this audio segment and whether there is an audio segment with an audio energy value of 0.75db on the right side.
[0150] If any one does not exist or both do not exist, it means that this audio segment is not the audio signal in the test audio source corresponding to the video frame, or this video frame is not the first solid-color frame image in the test video source, and it may be other solid-color frame images recorded during the media file recording process.
[0151] If both exist, it means that the audio signal in the test audio source corresponding to the video frame has been found, and the subsequent analysis process of the audio-visual synchronization situation can be entered.
[0152] During the analysis process of the audio-visual synchronization situation, the audio-visual synchronization situation of the media file can be determined based on the audio timestamp and the video timestamp corresponding to the matched audio energy value. As Figure 9 shown, there may be three situations in the determined audio-visual synchronization situation of the media file:
[0153] In the first situation, the video timestamp is greater than the audio timestamp, that is, the audio lags behind the video. As Figure 9 shown, the audio lags behind the video by a duration of T1;
[0154] In the second case, the video timestamp is equal to the audio timestamp, that is, the audio is synchronized with the video, that is, the audio and video are synchronized. As Figure 9 shown, the video timestamp of the first solid-color frame image is the same as the audio timestamp T0 of the corresponding 1 db audio signal;
[0155] In the third case, the video timestamp is less than the audio timestamp, that is, the audio is ahead of the video. As Figure 9 shown, the audio is ahead of the video by a duration of T2.
[0156] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.
[0157] Based on the same inventive concept, the embodiments of the present application also provide a media file recording device for implementing the above-mentioned media file recording method. The implementation solutions provided by this device to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the media file recording device provided below can refer to the limitations on the media file recording method in the above text, and will not be repeated here.
[0158] In one embodiment, as Figure 10 shown, a media file recording device is provided, including: a test clip source acquisition module 901 and a recording module 902, where:
[0159] The test clip source acquisition module 901 is used to acquire a test clip source, and the test clip source includes a test audio source and a test video source. Among them, the test audio source includes a number of identical audio groups arranged at equal distances, and each audio group includes at least one audio signal; the test video source includes a number of first solid-color frame images arranged at equal distances, and the other frame images of the test video source are second solid-color frame images. The first solid-color frame images are different from the second solid-color frame images, and the first solid-color frame images correspond to the audio signals one by one;
[0160] The recording module 902 is used to play the test clip source during the recording process and obtain the recorded media file after the recording ends.
[0161] In some embodiments, the audio group includes a first audio signal and a second audio signal that are spaced apart by a first preset distance, and the energy values of the first audio signal and the second audio signal are different from each other and have a preset relationship.
[0162] In some embodiments, the audio group further includes a third audio signal that is spaced apart from the second audio signal by a second preset distance, and the energy values of the first audio signal, the second audio signal, and the third audio signal are different from each other and have a preset relationship.
[0163] In some embodiments, the audio signals in the audio group are instantaneous pulse signals.
[0164] In some embodiments, the maximum energy value among the energy values of the audio signals in the audio group is 1 db.
[0165] In some embodiments, it includes any one of the following:
[0166] The first solid-color frame image is a white image;
[0167] The second solid-color frame image is a black image;
[0168] The equal-distance interval between the audio groups includes 10 seconds;
[0169] The first preset distance is the same as the second preset distance;
[0170] The first preset interval distance includes 1 second;
[0171] The energy value of the second audio signal is greater than the energy values of the first audio signal and the third audio signal, or the energy values of the first audio signal, the second audio signal, and the third audio signal increase in sequence; or the energy values of the first audio signal, the second audio signal, and the third audio signal decrease in sequence;
[0172] The energy values of the first audio signal, the second audio signal, and the third audio signal are 0.5 M, 1 db, and 0.75 M respectively.
[0173] Each module in the above media file recording device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0174] Based on the same inventive concept, an embodiment of the present application further provides a media file analysis device for implementing the media file analysis method involved above. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the media file analysis device provided below can refer to the limitations on the media file analysis method in the above text, and will not be repeated here.
[0175] In one embodiment, as Figure 11 shown, a media file analysis device is provided, including: a decoding module 1001, an audio energy calculation module 1002, a video energy calculation module 1003, a matching search module 1004, and a synchronization determination module 1005, where:
[0176] The decoding module 1001 is configured to decode the media file to be analyzed into an audio sequence and a video frame sequence;
[0177] The audio energy calculation module 1002 is configured to divide the audio sequence into several audio segments with a preset audio duration, and calculate the audio energy values of the audio segments;
[0178] The video energy calculation module 1003 is configured to calculate the video energy values of the video frames in the video frame sequence;
[0179] The matching search module 1004 is configured to, when the video energy value is the same as a preset video energy value, search for an audio energy value that matches a preset audio energy value within a preset time range of the video timestamp of the video frame whose video energy value is the same as the preset video energy value;
[0180] The synchronization determination module 1005 is configured to determine the audio-visual synchronization situation of the media file based on the audio timestamp corresponding to the matched audio energy value and the video timestamp.
[0181] In one embodiment, a computer device is provided. This computer device can be a terminal, and its internal structure diagram can be as Figure 12As shown in the figure. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be achieved through WIFI, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a media file recording method and / or a media file analysis method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad set on the computer device housing, or an external keyboard, touchpad, or mouse, etc.
[0182] Those skilled in the art can understand that Figure 12 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0183] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, it implements the steps of the media file recording method and / or the media file analysis method in any of the above embodiments.
[0184] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, it implements the steps of the media file recording method and / or the media file analysis method in any of the above embodiments.
[0185] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, it implements the steps of the media file recording method and / or the media file analysis method in any of the above embodiments.
[0186] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0187] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memories can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0188] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0189] The above-described embodiments merely represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A method for recording a media file, characterized in that, The method includes: Obtaining a test source, where the test source includes a test audio source and a test video source. Among them, the test audio source includes a plurality of identical audio groups arranged at equal distances, and each audio group includes at least one audio signal; the test video source includes a plurality of first solid-color frame images arranged at equal distances, and the other frame images of the test video source are second solid-color frame images. The first solid-color frame images are different from the second solid-color frame images, and the first solid-color frame images correspond one-to-one with the audio signals; Playing the test source during the recording process and obtaining the recorded media file after the recording ends.
2. The method according to claim 1, characterized in that The audio group includes a first audio signal and a second audio signal spaced apart by a first preset distance, and the energy values of the first audio signal and the second audio signal are different from each other and have a preset relationship.
3. The method according to claim 2, wherein The audio group further includes a third audio signal spaced apart from the second audio signal by a second preset distance, and the energy values of the first audio signal, the second audio signal, and the third audio signal are different from each other and have a preset relationship.
4. The method according to any one of claims 1 to 3, characterized in that The audio signals in the audio group are instantaneous pulse signals.
5. The method according to any one of claims 1 to 4, characterized in that, Including any one of the following: The first solid-color frame image is a white image; The second solid-color frame image is a black image; The first preset distance is the same as the second preset distance; The energy value of the second audio signal is greater than the energy values of the first audio signal and the third audio signal, or the energy values of the first audio signal, the second audio signal, and the third audio signal increase in sequence; or the energy values of the first audio signal, the second audio signal, and the third audio signal decrease in sequence.
6. A method for analyzing media files, characterized in that, The method includes: Decoding the media file to be analyzed into an audio sequence and a video frame sequence; Dividing the audio sequence into a plurality of audio segments at a preset audio duration and calculating the audio energy values of the audio segments; Calculating the video energy values of the video frames in the video frame sequence; When the video energy value is the same as the preset video energy value, within a preset time range of the video timestamp of the video frame with the same video energy value as the preset video energy value, searching for an audio energy value that matches the preset audio energy value; Based on the audio timestamp corresponding to the matched audio energy value and the video timestamp, determining the audio-visual synchronization situation of the media file.
7. The method according to claim 6, wherein The determining the audio-visual synchronization situation of the media file based on the audio timestamp corresponding to the matched audio energy value and the video timestamp includes: Searching in the preset time direction of the audio timestamp of the matched audio energy value to find whether there is an audio energy value that is spaced apart from the matched audio energy value by a preset distance and has a preset relationship; If it exists, determining the audio-visual synchronization situation of the media file based on the audio timestamp corresponding to the matched audio energy value and the video timestamp.
8. The method according to claim 7, characterized in that, Searching for whether there exists an audio energy value that is at a preset distance from the matched audio energy value and has a preset relationship in the preset time direction of the audio timestamp of the matched audio energy value includes at least one of the following items: Searching for whether there exists an audio energy value that is at a first preset distance from the matched audio energy value and has a first preset relationship in the first preset time direction of the audio timestamp of the matched audio energy value; Searching for whether there exists an audio energy value that is at a second preset distance from the matched audio energy value and has a second preset relationship in the second preset time direction of the audio timestamp of the matched audio energy value; Searching for whether there exists an audio energy value that is at a third preset distance from the matched audio energy value and has a third preset relationship in the first preset time direction of the audio timestamp of the matched audio energy value, where the third preset distance is greater than the first preset distance; Searching for whether there exists an audio energy value that is at a fourth preset distance from the matched audio energy value and has a fourth preset relationship in the second preset time direction of the audio timestamp of the matched audio energy value, where the fourth preset distance is greater than the second preset distance.
9. The method according to any one of claims 6 to 8, characterized in that: The video energy value of the video frame is the Y component in the YUV color space of the video frame.
10. A media file recording device, characterized in that, The apparatus includes: A test source acquisition module, configured to acquire a test source, where the test source includes a test audio source and a test video source. Among them, the test audio source includes a plurality of identical audio groups arranged at equal distances, and each audio group includes at least one audio signal; the test video source includes a plurality of first solid-color frame images arranged at equal distances, and the other frame images of the test video source are second solid-color frame images. The first solid-color frame images are different from the second solid-color frame images, and the first solid-color frame images correspond to the audio signals one by one; A recording module, configured to play the test source during recording and obtain the recorded media file after recording ends.
11. A media file analysis device, characterized in that, The apparatus includes: A decoding module, configured to decode the media file to be analyzed into an audio sequence and a video frame sequence; An audio energy calculation module, configured to divide the audio sequence into a plurality of audio segments with a preset audio duration and calculate the audio energy values of the audio segments; A video energy calculation module, configured to calculate the video energy values of the video frames in the video frame sequence; A matching search module, configured to, when the video energy value is the same as the preset video energy value, search for an audio energy value that matches the preset audio energy value within a preset time range of the video timestamp of the video frame whose video energy value is the same as the preset video energy value; A synchronization determination module, configured to determine the audio-visual synchronization situation of the media file based on the audio timestamp corresponding to the matched audio energy value and the video timestamp.
12. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 9.