Audio file marking method, device, electronic device and storage medium
By performing time code alignment and abnormal audio identification on multi-channel audio files and automatically labeling abnormal audio using a preset model, the problem of low efficiency in troubleshooting audio anomalies in film and television shooting and meeting records is solved, and efficient audio file processing is achieved.
Patent Information
- Application Number
- CN202511000547.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-07-21
AI Technical Summary
In scenarios such as filming and conference recording that require high-frequency audio collection, wirelessly received audio data is easily interfered with, resulting in abnormal audio. Existing technologies require manual investigation and labeling of abnormal audio, which is labor-intensive and inefficient.
By obtaining multi-channel audio files with time code alignment, using the preset sound recognition model to identify abnormal audio segments, and writing the metadata of the marked intervals into the audio files, the abnormal audio can be automatically labeled and located.
It improves the efficiency of audio file anomaly identification, reduces manual intervention, accurately marks abnormal audio clips, and improves the efficiency of audio file organization and editing.
Smart Images

Figure CN120496531B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of audio processing technology, and in particular to a method, device, electronic device, and storage medium for marking audio files. Background Art
[0002] In scenarios such as film and television shooting and meeting recording that require high-frequency audio collection, wirelessly received audio data is easily interfered with and anomalies may occur. Manual troubleshooting of the audio data is usually required to locate and label the abnormal audio, and to find recording files that can be used to replace the abnormal audio. This results in a large workload and low efficiency in troubleshooting audio data anomalies. Summary of the Invention
[0003] The embodiments of the present application provide a method, device, electronic device, and storage medium for marking audio files, which can improve the efficiency of troubleshooting audio data anomalies.
[0004] In a first aspect, this embodiment provides a method for marking an audio file, comprising:
[0005] Acquire an audio file to be marked; the audio file to be marked includes an audio file of at least two channels with time codes aligned;
[0006] Identify the audio file of any channel in the audio file to be marked according to a preset sound recognition model to obtain a recognition result;
[0007] If the recognition result indicates that an abnormal audio segment exists in the audio file of any channel of the audio file to be marked, determining a marking interval corresponding to the abnormal audio segment;
[0008] The metadata corresponding to the marking interval is written into the audio file to be marked to obtain a target audio file, and the target audio file is used to locate the abnormal audio segment.
[0009] In some embodiments, obtaining the audio file to be marked includes:
[0010] Acquire a first audio file and at least one second audio file; the first audio file and the second audio file are from different memories;
[0011] The first time code of the first audio file is aligned with the second time code of at least one of the second audio files to obtain the audio file to be marked.
[0012] In some embodiments, aligning the first time code of the first audio file with the second time code of at least one second audio file to obtain the audio file to be marked includes:
[0013] Splitting at least one second audio file based on the first time code and at least one second time code to obtain a split second audio file;
[0014] The audio file to be marked is generated according to the first audio file and at least one of the cut second audio files.
[0015] In some embodiments, the segmenting of at least one second audio file based on the first time code and at least one second time code to obtain a segmented second audio file includes:
[0016] determining a first start time and a first end time of the first time code, and a second start time and a second end time of the second time code;
[0017] When the first start time is later than the second start time, cutting the start portion of the second audio file based on the first start time to obtain the cut second audio file;
[0018] In a case where the first end time is earlier than the second end time, the end portion of the second audio file is cut based on the first end time to obtain the cut second audio file.
[0019] In some embodiments, the method further comprises:
[0020] Obtain a target timeline file for the audio file to be marked; the target timeline file at least includes the start time and end time of the audio file to be marked;
[0021] The timeline tag corresponding to the marked interval is written into the target timeline file to obtain a marked timeline file; the marked timeline file is used to locate the time corresponding to the abnormal audio segment.
[0022] In some embodiments, the sound recognition model includes a sound extraction network and a sound recognition network; and recognizing an audio file of any channel of the audio file to be marked according to the preset sound recognition model to obtain a recognition result includes:
[0023] Performing sound feature extraction on a first audio segment through the sound extraction network to obtain sound feature data of the first audio segment; the first audio segment is any audio segment of the at least one audio segment of the audio file to be marked;
[0024] Outputting, by the sound recognition network, a first probability parameter that the sound feature data belongs to human voice feature data;
[0025] If the value of the first probability parameter is less than or equal to a first threshold, the recognition result is determined to be that the first audio segment is the abnormal audio segment.
[0026] In some embodiments, the abnormal audio segment includes a first abnormal segment and a second abnormal segment, the first abnormal segment is an audio segment containing noise; the second abnormal segment is an audio segment without human voice; the method further includes:
[0027] If the value of the first probability parameter is less than or equal to the first threshold and greater than a second threshold, determining that the recognition result is that the first audio segment is the first abnormal segment; wherein the first threshold is greater than or equal to the second threshold;
[0028] If the value of the first probability parameter is less than or equal to the second threshold, the recognition result is determined to be that the first audio segment is the second abnormal segment.
[0029] In a second aspect, this embodiment further provides an audio file marking device, comprising:
[0030] A first acquisition module is configured to acquire an audio file to be marked; the audio file to be marked includes an audio file of at least two channels with time codes aligned;
[0031] A first recognition module is used to recognize an audio file of any channel in the audio file to be marked according to a preset sound recognition model to obtain a recognition result;
[0032] a first determining module, configured to, if the recognition result indicates that an abnormal audio segment exists in the audio file of any one of the channels, determine a marked interval corresponding to the abnormal audio segment;
[0033] The first writing module is used to write metadata corresponding to the marking interval into the audio file to be marked to obtain a target audio file, and the target audio file is used to locate the abnormal audio segment.
[0034] In a third aspect, this embodiment further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the audio file marking method is implemented.
[0035] In a fourth aspect, this embodiment further provides a computer-readable storage medium storing a computer program, which is loaded by a processor to execute the steps in the audio file marking method.
[0036] The audio file marking method provided by this embodiment can effectively improve the efficiency of abnormal audio recognition of audio files by performing abnormal audio recognition on multiple time-synchronized audio files, avoiding performing abnormal audio recognition on individual audio files with different times one by one; recognize the audio file of any channel in the audio file to be marked by using a preset sound recognition model, avoiding the large workload caused by manual abnormal audio recognition of audio files, and further improving the efficiency of abnormal audio recognition of audio files; and automatically mark the abnormal audio segments of the audio file to be marked by writing the metadata of the marking interval corresponding to the abnormal audio segment into the audio file to be marked, thereby more effectively improving the efficiency of abnormal audio recognition of audio files and the efficiency of editing the audio file to be marked. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0038] Figure 1 A flowchart of a method for marking an audio file provided in an embodiment of the present application;
[0039] Figure 2 Another flowchart of a method for marking an audio file provided in an embodiment of the present application;
[0040] Figure 3 A schematic diagram of the structure of an audio file marking device provided in an embodiment of the present application;
[0041] Figure 4 A schematic diagram of the hardware structure of an audio file marking device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0042] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0043] In the description of the present application, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present application. In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the said features. In the description of the present application, "multiple" means two or more, unless otherwise clearly and specifically defined.
[0044] “A and / or B” includes the following three combinations: A only, B only, and a combination of A and B.
[0045] The use of "suitable for" or "configured to" in this application is intended to be open and inclusive language, and does not exclude devices that are adapted or configured to perform additional tasks or steps. In addition, the use of "based on" is intended to be open and inclusive, as a process, step, calculation, or other action that is "based on" one or more stated conditions or values may, in practice, be based on additional conditions or values beyond those stated.
[0046] In this application, the word "exemplary" is used to mean "serving as an example, illustration, or illustration." Any embodiment described in this application as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments. The following description is given to enable any person skilled in the art to implement and use the present application. In the following description, details are listed for the purpose of explanation. It should be understood that one of ordinary skill in the art can recognize that the present application can be implemented without using these specific details. In other instances, well-known structures and processes are not elaborated in detail to avoid obscuring the description of the present application with unnecessary details. Therefore, the present application is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed in this application.
[0047] In scenarios such as film and television shooting and meeting recording that require high-frequency audio collection, the amount of collected audio is large and the duration is long. The collected audio is generally stored in multiple storage devices (for example, memory cards). The workload is large in the process of manual post-processing and editing of the collected audio, and the problem of not being able to find the audio material may occur easily.
[0048] Figure 1A flowchart of a method for marking an audio file provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, the audio file marking method provided in the embodiment of the present application may include but is not limited to the following steps and combinations of the following steps.
[0049] Step 101: Acquire an audio file to be marked; the audio file to be marked includes an audio file of at least two channels with aligned time codes.
[0050] The audio file to be marked and the time code can be determined based on actual conditions and are not limited here. As an example, the audio file to be marked can be a single-audio multi-channel file. The time code can be a Society of Motion Picture and Television Engineers (SMPTE) time code.
[0051] Exemplarily, obtaining the audio file to be marked may include obtaining at least two audio files; each of the at least two audio files is from a different storage; a time code may be embedded in each of the at least two audio files; and the time code of each of the at least two audio files is aligned to obtain the audio file to be marked, and the at least two audio files are respectively used as audio files of at least two channels of the audio file to be marked.
[0052] In some embodiments, the audio file marking method can be applied to a wireless recording system, which includes multiple independent recording devices, each of which includes a memory. The wireless recording system synchronizes the SMPTE time code through a wireless protocol during recording to ensure the time code consistency of the audio files in each memory, which facilitates subsequent alignment processing.
[0053] In some embodiments, obtaining an audio file to be marked includes:
[0054] Obtain a first audio file and at least one second audio file; the first audio file and the second audio file are from different memories;
[0055] A first time code of a first audio file is aligned with a second time code of at least one second audio file to obtain an audio file to be marked.
[0056] Exemplarily, acquiring the first audio file may include receiving at least one audio file sent from a first memory; and determining the first audio file from the at least one audio file based on a selection instruction, wherein the selection instruction may carry identification information of the first audio file, and the identification information may be a first time code of the first audio file. Acquiring at least one second audio file may include receiving a second audio file sent from each second memory in the at least one second memory, wherein each second audio file may be a single-channel audio file.
[0057] In some embodiments, the first audio file may be embedded with a first time code; the second audio file may be embedded with a second time code; and the time interval of the first time code may be less than or equal to the time interval of the second time code. The first memory may be a first memory card, the second memory may be a second memory card, and the first memory card and the second memory card may be different memory cards.
[0058] As an example, the first memory card can be a primary memory card, used to store the sounds recorded in the primary recording scene during the day's recording process; the second memory card can be a secondary memory card, used to store the sounds recorded in all recording scenes during the day's recording process. In some embodiments, the multiple independent recording devices may include a primary recording device and multiple secondary recording devices, wherein the primary recording device includes a primary memory card and the secondary recording devices include secondary memory cards.
[0059] It should be noted that the first memory card may store audio multiple times for different primary recording scenes during a given day's recording process, with each recording scene storing an audio file corresponding to at least one audio file. The second memory card may store audio only once during a given day's recording process, with each recording scene corresponding to an audio file. Different second memory cards may correspond to different second audio files, and different second memory cards may be worn by different users whose audio is to be collected. It is understood that the number of audio files recorded by the first memory card is greater than or equal to the number of audio files recorded by the second memory card, and this is not limited here.
[0060] Exemplarily, aligning a first time code of a first audio file with a second time code of at least one second audio file to obtain an audio file to be marked may include: determining a first alignment time of the second audio file in the second time code based on a first start time of the first time code using audio and video processing technology; performing alignment based on the first start time and the first alignment time to obtain an aligned second audio file; and / or determining a second alignment time of the second audio file in the second time code based on a first end time of the first time code; performing alignment based on the first end time and the second alignment time to obtain an aligned second audio file; and mixing the first audio file and the at least one aligned second audio file to obtain an audio file to be marked including at least two channels with aligned time codes. The first audio file may be a primary channel audio file, and the second audio file may be a secondary channel audio file.
[0061] Here, by aligning the alignment time of the second audio file with the start time of the first audio file, and / or aligning the alignment time of the second audio file with the end time of the first audio file, an audio file to be marked is generated according to the first audio time and the aligned second audio file. This allows for time synchronization of multiple audio files to obtain multi-channel audio files with aligned time codes, thereby avoiding the problem of difficulty in searching and matching audio materials of the same time, and effectively improving the efficiency of organizing and editing the collected audio.
[0062] In some embodiments, aligning a first time code of a first audio file with a second time code of at least one second audio file to obtain an audio file to be marked includes:
[0063] Splitting the at least one second audio file based on the first time code and the at least one second time code to obtain a split second audio file;
[0064] An audio file to be marked is generated according to the first audio file and at least one cut second audio file.
[0065] Exemplarily, based on the first time code and at least one second time code, at least one second audio file is cut, and the cut second audio file can be obtained by performing overlap matching on the first time interval of the first time code and the second time interval of the second time code to obtain an overlap matching result; if the overlap matching result indicates that the second time interval includes a third time interval that does not overlap with the first time interval, the second audio file is cut according to the third time interval to obtain a cut second audio file.
[0066] Exemplarily, generating the audio file to be marked based on the first audio file and the at least one cut second audio file can be based on mixing the first audio file and the at least one cut second audio file to obtain the audio file to be marked including at least two channels with time code alignment.
[0067] Here, by cutting at least one second audio file, generating an audio file to be marked based on the first audio file and at least one cut second audio file, the time codes of the first audio file and the second audio file can be matched more accurately, avoiding the problem of difficulty in querying and matching audio materials of the same time, and effectively improving the efficiency of organizing and editing the collected audio.
[0068] In some embodiments, segmenting at least one second audio file based on the first time code and the at least one second time code to obtain a segmented second audio file includes:
[0069] determining a first start time and a first end time of the first time code, and a second start time and a second end time of the second time code;
[0070] When the first start time is later than the second start time, cutting the start portion of the second audio file based on the first start time to obtain a cut second audio file;
[0071] When the first end time is earlier than the second end time, the end portion of the second audio file is cut based on the first end time to obtain a cut second audio file.
[0072] Exemplarily, determining the first start time and the first end time of the first time code may be by determining the first start time and the first end time of the first time code based on the first time interval of the first time code; determining the second start time and the second end time of the second time code may be by determining the second start time and the second end time of the second time code based on the second time interval of the second time code.
[0073] Exemplarily, when the first start time is later than the second start time, the starting part of the second audio file is cut based on the first start time, and the cut second audio file may be: if the first start time is later than the second start time, the time interval between the first start time and the second start time is determined to be the third time interval, the target audio part corresponding to the third time interval is determined to be the starting part of the second audio file, the target audio part is cut to obtain the cut second audio file.
[0074] Exemplarily, when the first end time is earlier than the second end time, the end part of the second audio file is cut based on the first end time, and the cut second audio file may be: if the first end time is earlier than the second end time, the time interval between the first end time and the second end time is determined to be the third time interval, the target audio part corresponding to the third time interval is determined to be the end part of the second audio file, the target audio part is cut, and the cut second audio file is obtained.
[0075] Here, when the start time of the second audio file is earlier than the start time of the first audio file, the start part of the second audio file is cut; when the end time of the second audio file is later than the end time of the first audio file, the end part of the second audio file is cut. This can more accurately match the time codes of the first audio file and the second audio file, avoiding the problem of difficulty in searching and matching audio materials of the same time, and effectively improving the efficiency of organizing and editing the collected audio.
[0076] Step 102: Recognize the audio file of any channel in the audio file to be marked according to a preset sound recognition model to obtain a recognition result.
[0077] Exemplarily, the audio file to be marked includes a first-channel audio file, which is an audio file of any channel of the audio file of at least two channels of the audio file to be marked. The sound recognition model can be determined according to actual conditions and is not limited here. As an example, the sound recognition model can be an artificial intelligence (AI) sound recognition model, a voiceprint recognition model based on deep learning, or a noise classification model, which is used to distinguish between human voice audio, noise audio, or human voice-missing audio (for example, audio without human voice) in the first-channel audio file. It should be noted that the embodiment of the present application can identify each audio file of at least two channels of audio files with time codes aligned, and the first-channel audio file can be the first audio file or the second audio file.
[0078] Exemplarily, the recognition result may indicate whether there is an abnormal audio segment in the first channel audio file. It is understandable that any audio segment of the first channel audio file in the to-be-marked audio file may be identified according to a preset sound recognition model to obtain an identification result of whether each audio segment is an abnormal audio segment.
[0079] In some embodiments, the sound recognition model includes a sound extraction network and a sound recognition network; an audio file of any channel in the to-be-labeled audio file is recognized according to the preset sound recognition model to obtain a recognition result, including:
[0080] Extracting sound features from the first audio segment using a sound extraction network to obtain sound feature data of the first audio segment; the first audio segment is any audio segment in at least one audio segment of the audio file to be marked;
[0081] Outputting, through a voice recognition network, a first probability parameter that the sound feature data belongs to human voice feature data;
[0082] If the value of the first probability parameter is less than or equal to the first threshold, the recognition result is determined to be that the first audio segment is an abnormal audio segment.
[0083] Exemplarily, the sound extraction network can be determined according to actual conditions, which is not limited here. As an example, the sound extraction network can be a feature extraction network that extracts sound features. The sound feature data includes a sound feature matrix, which includes at least one sound feature. The first probability parameter output by the sound recognition network that the sound feature data belongs to the human voice feature data can be obtained by mapping the sound feature matrix through the sound recognition network to obtain a first probability matrix of the sound feature matrix. The first probability matrix includes the first probability parameter of each sound feature in the sound feature matrix belonging to the human voice feature. The first threshold is a pre-set probability critical value for determining whether an audio segment is an abnormal audio segment, which is not limited here.
[0084] Exemplarily, the value of the first probability parameter is compared with a preset first threshold to obtain a comparison result; if the comparison result shows that the value of the first probability parameter is greater than the preset first threshold, the recognition result is determined to be that the first audio segment is normal human voice audio; if the comparison result shows that the value of the first probability parameter is less than or equal to the preset first threshold, the recognition result is determined to be that the first audio segment is an abnormal audio segment.
[0085] In some embodiments, the first audio clip can be any audio clip of a first-channel audio file or any audio clip of a second-channel audio file; the second-channel audio file is an audio file other than the first-channel audio file in the at least two channels of the audio file to be marked. If the first audio clip of the first-channel audio file is an abnormal audio clip, the second audio clip of the second-channel audio file with the same time code can be called and inserted into the position of the first audio clip.
[0086] Here, the sound feature data of the first audio segment in the audio file to be marked is extracted through the sound extraction network, and the first probability parameter that the sound feature data belongs to human voice feature data is output through the sound recognition network. Then, the value of the first probability parameter is compared with the first threshold to determine whether the first audio segment is an abnormal audio segment, thereby avoiding the large workload caused by manual abnormal audio recognition of audio files and further improving the efficiency of abnormal audio recognition of audio files.
[0087] In some embodiments, the abnormal audio segment includes a first abnormal segment and a second abnormal segment, the first abnormal segment is an audio segment containing noise; the second abnormal segment is an audio segment without human voice; the method further includes:
[0088] If the value of the first probability parameter is less than or equal to a first threshold and greater than a second threshold, then the recognition result is determined to be that the first audio segment is a first abnormal segment; wherein the first threshold is greater than or equal to the second threshold;
[0089] If the value of the first probability parameter is less than or equal to the second threshold, the recognition result is determined to be that the first audio segment is a second abnormal segment.
[0090] Exemplarily, the comparison result includes a first comparison result and a second comparison result; the value of the first probability parameter is compared with a preset first threshold to obtain a first comparison result; if the first comparison result indicates that the value of the first probability parameter is greater than the preset first threshold, the recognition result is determined to be that the first audio segment is human voice audio without abnormalities; if the first comparison result indicates that the value of the first probability parameter is less than or equal to the preset first threshold, the value of the first probability parameter is compared with the preset second threshold to obtain a second comparison result; if the second comparison result indicates that the value of the first probability parameter is greater than the preset second threshold, the recognition result is determined to be that the first audio segment is a first abnormal segment; if the second comparison result indicates that the value of the first probability parameter is less than or equal to the preset second threshold, the recognition result is determined to be that the first audio segment is a second abnormal segment.
[0091] Here, by comparing the value of the first probability parameter with the first threshold and / or the second threshold, it is determined whether the first audio segment is an abnormal audio segment and the abnormal type of the abnormal audio segment, which can more accurately mark the audio file as abnormal, and further improve the efficiency of abnormal audio recognition in audio files.
[0092] Step 103: When the recognition result indicates that an abnormal audio segment exists in the audio file of any channel, a marking interval corresponding to the abnormal audio segment is determined.
[0093] Exemplarily, the time interval in which the abnormal audio segment is located is determined as the marking interval. In some embodiments, the starting point of the marking interval can be determined as the marking point (for example, Q point), where the marking point can be the entry mark of the marking interval.
[0094] Step 104: write metadata corresponding to the marked interval into the audio file to be marked to obtain a target audio file, which is used to locate the abnormal audio segment.
[0095] For example, the metadata may include the start time and duration of the marked interval, as well as an identifier of the channel to which the abnormal audio segment corresponding to the marked interval belongs. In some embodiments, the metadata also includes at least information about the abnormality type of the abnormal audio segment (e.g., an audio segment containing noise or an audio segment without human voice).
[0096] Exemplarily, the metadata corresponding to the marked interval is written into the audio file to be marked, and the marked audio file to be marked can be obtained by writing the metadata corresponding to the marked interval into the marked interval of the abnormal audio segment in the audio file to be marked to obtain the target audio file.
[0097] In some embodiments, the method further comprises:
[0098] Obtain a target timeline file for the audio file to be marked; the target timeline file at least includes the start time and end time of the audio file to be marked;
[0099] The timeline tag corresponding to the marked interval is written into the target timeline file to obtain a marked timeline file; the marked timeline file is used to locate the time corresponding to the abnormal audio segment.
[0100] For example, the timeline tag may include a marker track, which includes at least marker information such as the start time and duration of the marker interval, so that the target timeline file can locate the abnormal audio segment by the start time and duration.
[0101] It should be noted that some editing software can directly edit the marked audio files to be marked, while other editing software cannot directly edit the audio files to be marked, and require the target timeline file to assist in locating abnormal audio clips in the audio files to be marked.
[0102] Here, the timeline tag corresponding to the marked interval of the abnormal audio segment is written into the target timeline file to obtain the marked timeline file, which can assist the audio file to be marked to locate the time corresponding to the abnormal audio segment, effectively improving the efficiency of abnormal audio recognition of audio files and the efficiency of editing the audio files to be marked.
[0103] The audio file marking method provided by this embodiment can effectively improve the efficiency of abnormal audio recognition of audio files by performing abnormal audio recognition on multiple time-synchronized audio files, avoiding performing abnormal audio recognition on individual audio files with different times one by one; recognize the audio file of any channel in the audio file to be marked by using a preset sound recognition model, avoiding the large workload caused by manual abnormal audio recognition of audio files, and further improving the efficiency of abnormal audio recognition of audio files; and automatically mark the abnormal audio segments of the audio file to be marked by writing the metadata of the marking interval corresponding to the abnormal audio segment into the audio file to be marked, thereby more effectively improving the efficiency of abnormal audio recognition of audio files and the efficiency of editing the audio file to be marked.
[0104] The following describes the audio file marking method provided by the embodiment of the present application.
[0105] Figure 2 Another flowchart of a method for marking an audio file provided in an embodiment of the present application is shown in FIG. Figure 2 As shown, the method is applied to a wireless recording system, which can obtain at least one audio file of a main memory card, an audio file of a first secondary memory card, and an audio file of a second secondary memory card; in response to a user's selection instruction, select a first audio file from the at least one audio file of the main memory card; identify the time codes of the first audio file, the second audio file of the first secondary memory card, and the second audio file of the second secondary memory card; align and cut the second audio file of the first secondary memory card and the second audio file of the second secondary memory card respectively according to the time codes, and obtain a single-audio multi-channel file based on the first audio file and the cut audio files; identify abnormal audio segments of the first audio file and the cut audio files; on the one hand, write metadata corresponding to the marked interval of the abnormal audio segment into the single-audio multi-channel file to obtain a marked single-audio multi-channel file; on the other hand, write timeline tags corresponding to the marked interval of the abnormal audio segment into the timeline file to obtain a marked timeline file.
[0106] The embodiment of the present application can realize automatic alignment of audio of multiple memory cards, and automatic mixing of audio of multiple memory cards. In the film and television variety show shooting scene, the wireless recording system includes multiple independent recording devices, each of which has a memory card, and the main recording device and the auxiliary recording device are connected via a wireless protocol. Through audio and video processing technology, the single-channel files of multiple auxiliary cards are aligned with the time code of the main card file, and the second audio file is cut according to the time code interval that does not overlap with the first audio file, and the first audio file of the main memory card and the second audio file cut by the auxiliary memory card are mixed into a single multi-channel file.
[0107] The present invention also implements an automatic marking mechanism for abnormal audio file intervals on both the primary and secondary storage cards. In the case of human voice, this method uses an AI sound recognition model to identify whether an audio clip in an audio file is a human voice or noise, outputs the abnormal audio clip and the marked interval where the abnormal audio clip is located, and writes the marked interval into the metadata "Q point".
[0108] Through the method provided in the embodiment of the present application, editors can quickly locate the time interval where the abnormal audio clip is located, thereby improving editing efficiency; through the AI intelligent sound recognition model, it can be identified that there may be abnormal intervals in the audio files of the current primary storage card or secondary storage card, and written into the final mixed audio file in the form of metadata; finally, the timeline file of the corresponding editing software can be exported, and the timeline file contains the marking information of the marked interval where the abnormal audio clip is located.
[0109] The method provided by the embodiment of the present application combines audio files from multiple memory cards into one, improving the efficiency of collaboration between sound engineers and editors. Both primary and secondary channel audio files are embedded with SMPTE time codes. This method requires only the user to select the audio file from the primary memory card, cut the secondary memory card's audio file based on the primary memory card's audio file's time code start and / or end positions, then mix the audio files stored on the memory cards of multiple recording devices, ultimately outputting a single multi-channel audio file. This reduces communication time and workload between editors and sound engineers.
[0110] The method provided in the embodiments of the present application can improve the fault tolerance of a wireless recording system. In film and television variety show shooting scenarios, the wireless recording system includes multiple independent recording devices, each with a memory card, and the main and secondary devices are connected via a wireless protocol. Due to the wireless transmission, the risk of transmission anomalies is reduced.
[0111] Figure 3 A structural diagram of an audio file marking device provided in an embodiment of the present application is shown as follows: Figure 3 As shown, the audio file marking device 300 includes:
[0112] The first acquisition module 301 is configured to acquire an audio file to be marked; the audio file to be marked includes an audio file of at least two channels with time codes aligned;
[0113] The first recognition module 302 is configured to recognize an audio file of any channel in the to-be-marked audio file according to a preset sound recognition model to obtain a recognition result;
[0114] The first determining module 303 is configured to determine a marked interval corresponding to the abnormal audio segment if the recognition result indicates that the audio file of any channel has an abnormal audio segment;
[0115] The first writing module 304 is configured to write metadata corresponding to the marked interval into the audio file to be marked to obtain a target audio file, which is used to locate the abnormal audio segment.
[0116] In some embodiments, the first acquisition module 301 is further used to acquire a first audio file and at least one second audio file; the first audio file and the second audio file are from different storage devices; and the first time code of the first audio file is aligned with the second time code of the at least one second audio file to obtain the audio file to be marked.
[0117] In some embodiments, the first acquisition module 301 is further used to split at least one second audio file based on the first time code and at least one second time code to obtain a split second audio file; and generate an audio file to be marked based on the first audio file and the at least one split second audio file.
[0118] In some embodiments, the first acquisition module 301 is further configured to determine a first start time and a first end time of the first time code, and a second start time and a second end time of the second time code; when the first start time is later than the second start time, cutting the start portion of the second audio file based on the first start time to obtain a cut second audio file; and when the first end time is earlier than the second end time, cutting the end portion of the second audio file based on the first end time to obtain a cut second audio file.
[0119] In some embodiments, the audio file marking device 300 also includes: a second acquisition module, which obtains the target timeline file of the audio file to be marked; the target timeline file at least includes the start time and end time of the audio file to be marked; a second writing module, which is used to write the timeline label corresponding to the marking interval into the target timeline file to obtain the marked timeline file; the marked timeline file is used to locate the time corresponding to the abnormal audio segment.
[0120] In some embodiments, the sound recognition model includes a sound extraction network and a sound recognition network; the first recognition module 302 is further used to extract sound features of the first audio clip through the sound extraction network to obtain sound feature data of the first audio clip; the first audio clip is any audio clip of at least one audio clip of the audio file to be marked; the sound recognition network outputs a first probability parameter that the sound feature data belongs to human voice feature data; if the value of the first probability parameter is less than or equal to a first threshold, the recognition result is determined to be that the first audio clip is an abnormal audio clip.
[0121] In some embodiments, the abnormal audio segment includes a first abnormal segment and a second abnormal segment, the first abnormal segment is an audio segment with noise; the second abnormal segment is an audio segment without human voice; the first recognition module 302 is further used to determine that the recognition result is that the first audio segment is the first abnormal segment if the value of the first probability parameter is less than or equal to the first threshold and greater than the second threshold; wherein the first threshold is greater than or equal to the second threshold; if the value of the first probability parameter is less than or equal to the second threshold, the recognition result is determined that the first audio segment is the second abnormal segment.
[0122] To implement the method of the embodiment of the present application, Figure 4 A hardware structure diagram of an audio file marking device provided in an embodiment of the present application is shown as follows: Figure 4 As shown, an embodiment of the present application also provides an audio file marking device 40 that may include: a memory 401 for storing a computer program; and a processor 402 for implementing any of the above methods when executing the computer program. For example, the processor 402 may be used to implement: obtaining an audio file to be marked; the audio file to be marked includes an audio file of at least two channels with time codes aligned; identifying the audio file of any channel in the audio file to be marked according to a preset sound recognition model to obtain a recognition result; when the recognition result indicates that there is an abnormal audio segment in the audio file of any channel, determining the marking interval corresponding to the abnormal audio segment; writing the metadata corresponding to the marking interval to the audio file to be marked to obtain a target audio file, and the target audio file is used to locate the abnormal audio segment. The processor 402 may also implement the steps in any of the methods described above, which will not be repeated here.
[0123] It should be noted that the audio file marking device and the audio file marking method provided in the above embodiment belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0124] Of course, in actual application, Figure 4 As shown, the audio file marking device 40 may also include: at least one network interface 403. The various components in the audio file marking device are coupled together via a bus system 404. It is understood that the bus system 404 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 404 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, Figure 4Various buses are labeled as bus system 404. There may be at least one processor 402. Network interface 403 is used for wired or wireless communication between the audio file marking device and other devices. Memory 401 in the embodiments of this application is used to store various types of data to support the operation of the audio file marking device. The methods disclosed in the embodiments of this application can be applied to or implemented by processor 402. Processor 402 may be an integrated circuit chip with signal processing capabilities. During implementation, the steps of the methods described above can be performed by hardware integrated logic circuits or software instructions within processor 402. Processor 402 can be a general-purpose processor, a digital signal processor (DSP), or other programmable logic device, discrete gate or transistor logic device, discrete hardware components, etc. Processor 402 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly implemented and executed by a combination of hardware and software modules within a single-chip microcomputer. The software module may be located in a storage medium located in the memory 401. The processor 402 reads the information in the memory 401 and, in combination with its hardware, completes the steps of the aforementioned method. In an exemplary embodiment, the audio file marking device 40 may be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers (MCUs), microprocessors, or other electronic components to perform the aforementioned method.
[0125] Specifically, embodiments of the present application provide a computer-readable storage medium having a computer program stored thereon, such as a memory 401 storing the computer program. The computer program can be executed by a processor 402 to perform the steps of the aforementioned method. The computer-readable storage medium can be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface mount storage, optical disk, or CD-ROM.
[0126] In addition, all functional units in the embodiments of the present application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the above-mentioned integrated units can be implemented in the form of hardware or in the form of hardware plus software functional units.
[0127] Those skilled in the art will appreciate that all or part of the steps of the above-mentioned method embodiments may be implemented by hardware associated with program instructions, and the aforementioned program may be stored in a computer-readable storage medium. When the program is executed, the program executes the steps of the above-mentioned method embodiments. The aforementioned storage medium includes various media that can store program codes, such as mobile storage devices, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.
[0128] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.
[0129] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the detailed description of other embodiments above and will not be repeated here.
[0130] The above is a detailed introduction to the audio file marking method, device, electronic device and storage medium provided in the embodiments of the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. At the same time, for those skilled in the art, based on the ideas of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present application.
Claims
1. A method for marking an audio file, characterized in that: include: Get the audio file to be marked; The audio file to be marked includes an audio file of at least two channels with time codes aligned; The audio file to be marked is a single-audio multi-channel file; Identify the audio file of any channel in the audio file to be marked according to a preset sound recognition model to obtain a recognition result; If the recognition result indicates that an abnormal audio segment exists in the audio file of any channel of the audio file to be marked, determining a marking interval corresponding to the abnormal audio segment; Writing metadata corresponding to the marked interval into the audio file to be marked to obtain a target audio file, wherein the target audio file is used to locate the abnormal audio segment; The step of obtaining the audio file to be marked includes: Acquire a first audio file and at least one second audio file; the first audio file and the second audio file are from different memories; The first time code of the first audio file is aligned with the second time code of at least one of the second audio files to obtain the audio file to be marked.
2. The audio file marking method according to claim 1, characterized in that: The aligning the first time code of the first audio file with the second time code of at least one of the second audio files to obtain the audio file to be marked includes: Splitting at least one second audio file based on the first time code and at least one second time code to obtain a split second audio file; The audio file to be marked is generated according to the first audio file and at least one of the cut second audio files.
3. The audio file marking method according to claim 2, characterized in that: The step of segmenting at least one second audio file based on the first time code and at least one second time code to obtain a segmented second audio file includes: determining a first start time and a first end time of the first time code, and a second start time and a second end time of the second time code; When the first start time is later than the second start time, cutting the start portion of the second audio file based on the first start time to obtain the cut second audio file; In a case where the first end time is earlier than the second end time, the end portion of the second audio file is cut based on the first end time to obtain the cut second audio file.
4. The method for marking an audio file according to claim 1, wherein: The method further comprises: Obtain a target timeline file for the audio file to be marked; the target timeline file at least includes the start time and end time of the audio file to be marked; The timeline tag corresponding to the marked interval is written into the target timeline file to obtain a marked timeline file; the marked timeline file is used to locate the time corresponding to the abnormal audio segment.
5. The audio file marking method according to claim 1, characterized in that: The sound recognition model includes a sound extraction network and a sound recognition network; the audio file of any channel in the audio file to be marked is recognized according to the preset sound recognition model to obtain a recognition result, including: Performing sound feature extraction on a first audio segment through the sound extraction network to obtain sound feature data of the first audio segment; the first audio segment is any audio segment of the at least one audio segment of the audio file to be marked; Outputting, by the sound recognition network, a first probability parameter that the sound feature data belongs to human voice feature data; If the value of the first probability parameter is less than or equal to a first threshold, the recognition result is determined to be that the first audio segment is the abnormal audio segment.
6. The audio file marking method according to claim 5, characterized in that: The abnormal audio segment includes a first abnormal segment and a second abnormal segment, wherein the first abnormal segment is an audio segment containing noise; and the second abnormal segment is an audio segment without human voice. The method further includes: If the value of the first probability parameter is less than or equal to the first threshold and greater than a second threshold, determining that the recognition result is that the first audio segment is the first abnormal segment; wherein the first threshold is greater than or equal to the second threshold; If the value of the first probability parameter is less than or equal to the second threshold, the recognition result is determined to be that the first audio segment is the second abnormal segment.
7. A device for marking audio files, characterized in that: include: A first acquisition module is used to acquire the audio file to be marked; The audio file to be marked includes an audio file of at least two channels with time codes aligned; The audio file to be marked is a single-audio multi-channel file; A first recognition module is used to recognize an audio file of any channel in the audio file to be marked according to a preset sound recognition model to obtain a recognition result; a first determining module configured to, if the recognition result indicates that an abnormal audio segment exists in the audio file of any channel of the audio file to be marked, determine a marking interval corresponding to the abnormal audio segment; A first writing module is configured to write metadata corresponding to the marked interval into the audio file to be marked to obtain a target audio file, wherein the target audio file is used to locate the abnormal audio segment; Among them, the first acquisition module is also used to obtain a first audio file and at least one second audio file; the first audio file and the second audio file are from different memories; the first time code of the first audio file and the second time code of at least one second audio file are aligned to obtain the audio file to be marked.
8. An electronic device, characterized in that: The method comprises a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the method for marking an audio file according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium, characterized in that A computer program is stored thereon, and the computer program is loaded by a processor to execute the steps in the audio file marking method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Virtual Wireless Multitrack Recording System
US20100217414A1