An audio detection method, device, storage medium and electronic device
By identifying music events in audio segments and matching metadata information, this technology solves the problem of inaccurately collecting music-related data in audio, thus enabling accurate acquisition and statistics of music events.
Patent Information
- Application Number
- CN202210220184.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-08
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2042-03-08
AI Technical Summary
Current technology cannot accurately collect music-related statistics in the audio being tested, such as music duration and playback start and end times.
By acquiring audio segments, identifying music events, determining metadata information that matches the music events, and then determining statistical data based on the metadata information.
It achieves accurate identification and statistics of music events in audio, providing a statistical data reference for music events, which facilitates subsequent analysis.
Smart Images

Figure CN114596878B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure relate to the technical field of audio processing, and particularly relate to an audio detection method and device, a storage medium, and an electronic device. BACKGROUND
[0002] With the popularization of Internet technology and the rapid popularity of audio and video, users can play audio and video, such as live programs, songs, and audio novels, through electronic devices such as mobile phones and computers.
[0003] However, in the process of implementing the present disclosure, the inventors have found that at least the following technical problems exist in the prior art: The current audio detection method cannot count statistical data (such as music duration, music play start and end time, etc.) related to music in the audio to be detected. SUMMARY
[0004] Embodiments of the present disclosure provide an audio detection method, device, storage medium, and electronic device to accurately obtain statistical data in audio to be detected.
[0005] In a first aspect, embodiments of the present disclosure provide an audio detection method, comprising:
[0006] obtaining an audio segment in the audio to be detected, identifying a music event in the audio segment;
[0007] determining metadata information matched with the music event, and determining statistical data in the audio to be detected based on the metadata information.
[0008] In a second aspect, embodiments of the present disclosure further provide an audio detection device, comprising:
[0009] a music event identification module configured to obtain an audio segment in the audio to be detected, and identify a music event in the audio segment;
[0010] a statistical data determination module configured to determine metadata information matched with the music event, and determine statistical data in the audio to be detected based on the metadata information.
[0011] In a third aspect, embodiments of the present disclosure further provide an electronic device, comprising:
[0012] one or more processors;
[0013] a storage device configured to store one or more programs,
[0014] when the one or more programs are executed by the one or more processors, the one or more processors implement the audio detection method according to any of the embodiments of the present disclosure.
[0015] In a fourth aspect, the embodiments of the present disclosure further provide a storage medium containing computer executable instructions for executing the audio detection method according to any of the embodiments of the present disclosure when executed by a computer processor.
[0016] The technical solutions of the embodiments of the present disclosure achieve preliminary identification of music events in each audio segment by obtaining the audio segments in the detected audio and identifying the music events in the audio segments. Further, metadata information matched with the music events is determined, which achieves matching and obtaining of reference data and provides a reference basis for obtaining statistical data. The music events in each audio segment are counted according to the metadata information obtained by matching, and the statistical data in the detected audio is obtained, which achieves accurate obtaining of statistical data of music events. BRIEF DESCRIPTION OF DRAWINGS
[0017] The above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent upon understanding the following detailed description, taken in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals refer to the same or similar elements. It should be understood that the drawings are schematic, and the original and elements are not necessarily drawn according to the scale.
[0018] Figure 1 A flowchart of an audio detection method provided by the embodiments of the present disclosure is shown in FIG. 1.
[0019] Figure 2 A flowchart of an audio detection method provided by the embodiments of the present disclosure is shown in FIG. 1.
[0020] Figure 3 A flowchart of an audio detection method provided by the embodiments of the present disclosure is shown in FIG. 1.
[0021] Figure 4 A flowchart of an audio detection method provided by the embodiments of the present disclosure is shown in FIG. 1.
[0022] Figure 5 A flowchart of an audio detection method provided by the embodiments of the present disclosure is shown in FIG. 1.
[0023] Figure 6 A structural diagram of an audio detection device provided by the embodiments of the present disclosure is shown in FIG. 1.
[0024] Figure 7 A structural diagram of an electronic device provided by the embodiments of the present disclosure is shown in FIG. 1. DETAILED DESCRIPTION
[0025] Embodiments of the present disclosure will be described below in greater detail with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be construed as being limited to the embodiments set forth herein; rather, these embodiments are provided so as to more completely and thoroughly understand the present disclosure. It is understood that the drawings and embodiments of the present disclosure are for exemplary purposes only and are not intended to limit the scope of protection of the present disclosure.
[0026] It should be understood that each of the steps recited in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0027] As used herein, the term "comprising" and variations thereof, are intended to mean "including but not limited to". The term "based on" is intended to mean "based, at least in part, on". The term "one embodiment" means "at least one embodiment". The term "another embodiment" means "at least one additional embodiment". The term "some embodiments" means "at least some embodiments". Related definitions are given throughout the description.
[0028] It should be noted that the terms "first", "second", and the like in the present disclosure are merely used to distinguish different devices, modules or units, and do not imply the order or interdependence of the functions performed by these devices, modules or units.
[0029] It should be noted that the terms "one", "multiple" in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that "one" or "multiple" should be understood as "one or more" unless the context clearly indicates otherwise.
[0030] Figure 1 A flowchart of an audio detection method provided by an embodiment of the present disclosure is shown. The embodiment of the present disclosure is suitable for automatically obtaining statistical data of music events in audio. The method can be performed by an audio detection device provided by the embodiment of the present disclosure. The audio detection device can be implemented in the form of software and / or hardware, and can be implemented by an electronic device, which can be a mobile terminal or a PC, etc. Figure 1 The method of the present embodiment includes:
[0031] S110, obtaining an audio segment in the detected audio, and identifying a music event in the audio segment.
[0032] S120, determining metadata information matched with the music event, and determining statistical data in the detected audio based on the metadata information.
[0033] In the embodiments of the present disclosure, the electronic device can be any electronic device with audio and video playing function and / or audio and video processing function, which can include but is not limited to a smart phone, a wearable device, a computer, a server, and the like. The electronic device can obtain the detected audio in various ways. For example, the detected audio can be collected in real time by an audio collection device, or the detected audio can be called from a preset storage location or other devices, and the present disclosure does not limit the method of obtaining the detected audio.
[0034] The detected audio refers to audio that needs to be detected for statistical data, which can include but is not limited to audio in a live video, audio in a video, broadcast audio, and the like, and the present disclosure does not limit the audio. Accordingly, in some embodiments, obtaining the detected audio can be extracting audio data from a video (such as a live video or an offline video) as the detected audio.
[0035] In order to improve the recognition accuracy and efficiency of the audio, the detected audio is divided into multiple audio segments, and each audio segment is recognized and processed. Specifically, in the case of real-time data, the audio segments are sequentially divided from the real-time collected audio, and the obtained audio segments are recognized and processed in real time. In the case of offline data, each audio segment can be sequentially recognized and processed according to the time sequence of the divided audio segments. In some embodiments, for offline audio data, the obtained multiple audio segments can be parallel data to improve processing efficiency.
[0036] In the present embodiment, the audio segment can be audio data with a preset time length, which can contain one or more of music, environmental sound, speech, noise, and the like. The time length of the audio segment can be preset, for example, determined according to the recognition accuracy, and the present disclosure does not limit the time length of the audio segment. For example, the time length of the audio segment can be 20s. The music event can be a sound event represented by one or more of rhythm (such as beat, tempo, and pronunciation), pitch (such as melody and harmony), and dynamics (such as the volume of sound or notes), which can include but is not limited to background music, a cappella, and the like.
[0037] In order to identify the music event in the audio segment, feature analysis needs to be performed on the audio segment to determine whether the audio segment contains a music event. In some embodiments, at least one sound feature can be extracted by any feature extraction method (such as a mel-frequency cepstral coefficient extraction method, a linear prediction coefficient extraction method, etc.), the extracted sound feature is compared with a music feature in a music database, and it is determined whether the audio segment contains a music event according to the comparison result, wherein the music database refers to a database containing various music features. In some embodiments, the audio segment can be identified by a music recognition model, and it is determined whether the audio segment contains a music event according to the identification result, wherein the music recognition model can take music, chat sound, noise, etc. as training samples, wherein the audio data containing music is taken as a positive sample, and the samples not containing music, such as chat sound, noise audio data, are taken as negative samples. The music recognition model is trained based on the above sample data, and under the condition that the training end condition is met, a model with a music event recognition function is obtained. The method for identifying music events is not limited in this embodiment.
[0038] On the basis of the above-mentioned embodiments, for the audio segment containing a music event, metadata information matched with the music event is determined respectively to obtain the statistical data of the detected audio, and for the audio segment not containing a music event, the metadata information matching of the audio segment is not needed, thereby avoiding the waste of computing resources caused by invalid processing. The metadata information is the description information of the music metadata including the music feature in the music event, and in some embodiments, the metadata information can be a label formed by multiple description information of the music metadata, wherein the description information of the music metadata can include but is not limited to the spectral information of the music, the music name, the music type, the singer, the composer, etc. The information is not limited, and the metadata information can be in the form of music name-singer / performer, for example. It can be understood that the music metadata is represented by the metadata information, the metadata information has uniqueness, and the music metadata can be uniquely represented. Taking the metadata information as a statistical dimension can improve the reliability of the music event statistics in the audio, and then the music events in each audio segment are counted by the metadata information, which can improve the accuracy of the statistical data corresponding to the music events.
[0039] In some embodiments, the metadata information matched with the music event can be determined by extracting music features in the music event, matching the music features with music features corresponding to each music metadata, and determining metadata information of the music metadata matched successfully as the metadata information matched with the music event. In some embodiments, the music features include but are not limited to feature information of tone, rhythm, lyrics, etc. Accordingly, the above feature information extraction is performed on the music event, and the extracted feature information is matched in the preset metadata database to obtain the metadata information matched with the music event, wherein the preset metadata database can include multiple metadata information and feature information corresponding to the metadata information. In some embodiments, the music features can be audio fingerprint features. Accordingly, the audio fingerprint features of the audio segment are extracted, and the audio fingerprint features are matched in the fingerprint feature library based on the audio fingerprint features to determine the metadata information matched with the music event, wherein the fingerprint feature library includes music metadata and corresponding fingerprint features. It should be noted that the audio fingerprint features correspond one-to-one to the audio segments from which the audio fingerprint features are determined, and the music metadata in this embodiment can correspond to multiple fingerprint features. Specifically, the music metadata is divided into multiple music sub-data, and each music sub-data can overlap with part of the data. The fingerprint features corresponding to each music sub-data are determined respectively. Accordingly, matching the audio fingerprint features of the audio segment with the fingerprint features of the music metadata can be matching the audio fingerprint features of the audio segment with the fingerprint features of the multiple music sub-data in the music metadata respectively. If the audio fingerprint features of the audio segment match the fingerprint features of any music sub-data successfully, the metadata information of the music metadata to which the music sub-data belongs is determined as the metadata information corresponding to the music event in the audio segment.
[0040] The statistical data is the result of the statistics of the music events in each audio segment, i.e., the music statistical data. The statistical data can include but is not limited to the playing time of each music in the detected audio, the start time and stop time of the music playing, the number of users (e.g., audio listening users or video watching users of the audio) received during the playing of each music, etc. It should be noted that the type of statistical data can be determined according to business needs, and is not limited.
[0041] In some embodiments, the audio duration of all music events corresponding to each metadata information in the audio to be detected can be counted based on the metadata information corresponding to each music event. The number of continuous music events corresponding to the metadata information and the number of continuous music events corresponding to each metadata information in the audio to be detected can be determined according to the metadata information corresponding to the music event and the timestamp of the music event, so as to obtain the application situation of the music in the audio to be detected. The audio interval corresponding to each metadata information in the audio to be detected and the number of users received in each audio interval can be counted to evaluate the flow guiding ability of the music corresponding to each metadata information, etc.
[0042] On the basis of the above-mentioned embodiments, after obtaining the statistical data, the method further includes: obtaining music metadata corresponding to each music event according to the metadata information corresponding to the music event, repairing the music event in each audio segment according to the music metadata corresponding to the music event, and obtaining the repaired music event, so as to avoid the case that the noise interference included in the to-be-detected audio causes the music event to be unclear. Wherein, the repairing of the music event in each audio segment according to the music metadata corresponding to the music event can be to intercept music sub-data corresponding to the music event in the music metadata, and replace the audio data of the music event based on the music sub-data. Alternatively, after obtaining the statistical data, the method further includes: performing cutting, splicing or the like on the audio segment according to the metadata information corresponding to the music event in the audio segment, and obtaining one or more new audio segments, for example, the audio data corresponding to the same metadata information in the to-be-detected audio can be cut and spliced.
[0043] The audio detection method provided by the embodiments of the present disclosure realizes the preliminary identification of the music event in each audio segment by obtaining the audio segment in the detected audio and identifying the music event in the audio segment. Further, the metadata information matched with the music event is determined, the matching of the reference data is realized, and the reference basis for obtaining the statistical data is provided. The statistical data in the detected audio is obtained by counting the music event in each audio segment according to the matched metadata information, and the identification and statistics of the audio in the music dimension are realized, which facilitates the subsequent analysis of the to-be-detected audio based on the statistical data.
[0044] Reference Figure 2 , Figure 2 The audio detection method flowchart provided by the embodiments of the present disclosure can be combined with the various optional schemes of the audio detection method provided in the above-mentioned embodiments. The audio detection method provided by the embodiments of the present disclosure is further refined. Optionally, the identifying the music event in the audio segment includes: inputting the audio segment into a pre-trained music recognition model to obtain a music event identification result output by the music recognition model, wherein the music recognition model is trained based on an audio sample and an event label corresponding to the audio sample.
[0045] As Figure 2 , the method of the embodiments includes:
[0046] S210, obtaining an audio segment in a detected audio.
[0047] S220, inputting the audio segment into a pre-trained music recognition model to obtain a music event identification result output by the music recognition model, wherein the music recognition model is trained based on an audio sample and an event label corresponding to the audio sample.
[0048] S230, determine metadata information matched with the music event, and determine statistics in the detected audio based on the metadata information.
[0049] In this embodiment, the music recognition model has the ability to identify music events in audio data. For an input audio segment, the music recognition model can identify whether the audio segment includes a music event. Correspondingly, the training process of the music recognition model can include: obtaining an audio sample and an event label corresponding to the audio sample. The audio sample can include multiple different sound events, such as music, laughter, chat, noise, etc. Correspondingly, the event label corresponding to the audio sample can be an event identifier, such as a music identifier, a laughter identifier, a noise identifier, etc. Optionally, the audio sample including a music event is taken as a positive sample, and the audio sample including a laughter, chat, noise, etc. event is taken as a negative sample. Correspondingly, the event labels corresponding to the positive and negative samples can be positive and negative, respectively. The initial training model is trained based on the audio samples corresponding to the positive and negative samples and the event labels corresponding to the audio samples, and a music recognition model is obtained. The initial training model can include but is not limited to a long short-term memory network model, a support vector machine model, etc., which is not limited herein. Further, after the music recognition model is trained, an audio segment can be input into the pre-trained music recognition model to classify or recognize the sound event in the audio segment. The music recognition model can quickly output a music event recognition result. It should be noted that the pre-trained music recognition model does not need complex calculation in the online application of the audio detection device, and can quickly obtain a music event recognition result, thereby improving the speed of audio detection.
[0050] On the basis of the above embodiment, the music recognition model can also output the start and end timestamps of the recognized music event in the audio segment. Correspondingly, the training sample of the music recognition model also includes the start and end timestamps corresponding to the music event label in the audio sample. The music recognition model trained by the above training sample can identify whether the input audio segment includes a music event and the start and end timestamps of the music event.
[0051] Optionally, after identifying the music event in the audio segment, the method further includes: determining whether the duration of the music event in the audio segment is greater than a first preset duration, and if not, canceling the marking of the music event. The duration of the music event can be determined based on the start and end timestamps of the music event.
[0052] It should be noted that when the music event identification result includes a music event, it indicates that the audio segment in the detected audio can include an event of playing music or singing, etc., but it can also be caused by an interference sound, for example, the interference sound can be a short message prompt tone or a mobile phone ring, etc., and the interference sound can also contain music. In this case, it indicates that the music event in the audio segment is not a real music event, and the music event needs to be unmarked to avoid the case of misjudgment of the music event.
[0053] Specifically, the duration of the music event in the audio segment is judged. If the duration of the music event in the audio segment is greater than a first preset duration, it indicates that the music event meets the standard of music, and the music event marking of the audio segment remains unchanged. If the duration of the music event in the audio segment is less than the first preset duration, it indicates that the music event does not meet the standard of music, and the music event marking of the audio segment is cancelled. The first preset duration can be set according to historical experience. For example, the first preset duration can be 6s.
[0054] The audio detection method provided by the embodiments of the present disclosure classifies or identifies the sound event in the audio segment by inputting the audio segment into the pre-trained music recognition model to obtain a music event identification result. The music event less than the first preset duration, the interference sound misidentifying the music event, and the increase of the statistical workload caused by the short-time music event are removed.
[0055] Reference Figure 3 , Figure 3 The audio detection method flowchart provided by the embodiments of the present disclosure can be combined with the various optional schemes of the audio detection method provided in the above embodiments. The audio detection method provided by the embodiments of the present disclosure is further optimized. Optionally, the metadata information matched with the music event includes: extracting the audio fingerprint feature of the audio segment containing the music event; matching based on the audio fingerprint feature in the fingerprint feature library to determine the metadata information matched with the music event, wherein the fingerprint feature library includes music metadata and corresponding fingerprint features. Figure 3 The method of the present embodiment includes:
[0056] S310, acquiring an audio segment in the detected audio, and identifying a music event in the audio segment.
[0057] S320, for the audio segment containing the music event, extracting the audio fingerprint feature of the audio segment.
[0058] S330, based on the audio fingerprint feature, matching in the fingerprint feature library to determine the metadata information matched with the music event, wherein the fingerprint feature library includes music metadata and corresponding fingerprint features.
[0059] S340, determining statistical data in the detected audio based on the metadata information.
[0060] In the embodiment, the audio fingerprint feature refers to a digital feature of a music event, i.e., a music fingerprint feature, which is unique. Specifically, the audio fingerprint feature of the audio segment can be extracted by using an audio fingerprint technology, which includes but is not limited to Philips algorithm or Shazam algorithm.
[0061] The fingerprint feature library refers to a database containing music metadata and fingerprint features, and can pre-store a plurality of music metadata and fingerprint features corresponding to the music metadata. The fingerprint features corresponding to the music metadata can be used to match the audio fingerprint features. If the matching is successful, the metadata information matched by the music event is obtained. The fingerprint features can include but are not limited to frequency parameters and time parameters corresponding to the spectrum of the music metadata.
[0062] On the basis of the above embodiment, the extracting the audio fingerprint feature of the audio segment includes: according to the start and end time stamps of the music event in the audio segment, the audio segment is intercepted to obtain an intercepted audio segment, and the audio fingerprint feature of the intercepted audio segment is extracted. Specifically, in some embodiments, the start and end time stamps of the music event are included in the recognition result of the music event, the start time stamp and the end time stamp of the music event in the audio segment are obtained; according to the start and end time stamps of the music event, the audio segment is intercepted, the audio data corresponding to the music event in the audio segment is extracted, by eliminating the non-music event part of the audio data, only the audio data corresponding to the intercepted music event is determined to determine the audio fingerprint feature, avoiding the interference of the non-music event part of the audio data on the audio fingerprint feature, at the same time, reducing the amount of audio data for determining the audio fingerprint feature, which is conducive to the rapid extraction of the audio fingerprint feature.
[0063] On the basis of the above-mentioned embodiments, the extracting the audio fingerprint feature of the audio segment comprises: extracting audio data of the audio track where the music event is located in the audio segment, and extracting the audio fingerprint feature based on the audio data of the audio track where the music event is located. In some embodiments, the audio to be detected can include multiple audio tracks, that is, each audio segment includes multiple audio tracks. For example, the audio to be detected can include a background audio track and a speech audio track. In any audio segment, the audio data in the background audio track can be background music, and the audio data in the speech audio track can be the dialogue speech data of the host. For example, the audio data in the background audio track can be noise, and the audio data in the speech audio track can be the singing voice of the host. The different audio tracks can simultaneously include music events, or one or more of the audio tracks can independently include music events. By extracting the audio data of the audio track where the music event is located and excluding the music data of the audio track where the non-music event is located, the interference of the non-music event is reduced, and the accuracy of the subsequent extraction of the audio fingerprint feature is improved.
[0064] The audio detection method provided by the embodiments of the present disclosure extracts the audio fingerprint feature of the audio segment containing the music event, performs matching in the fingerprint feature library based on the extracted audio fingerprint feature, determines the metadata information matched with the music event, and obtains the metadata information corresponding to the music event through the fingerprint feature library matching. Therefore, the processing speed is fast, and the time for audio detection can be saved.
[0065] Reference Figure 4 , Figure 4 The audio detection method flowchart provided by the embodiments of the present disclosure can be combined with the various optional schemes of the audio detection method provided in the above-mentioned embodiments. The audio detection method provided by the embodiments of the present disclosure is further optimized. Optionally, the determining the statistical data in the detected audio based on the metadata information comprises: merging the music events corresponding to the same metadata information according to the start and end time stamps of the music events in each audio segment to obtain the statistical data in the detected audio.
[0066] As Figure 4 The method of the present embodiment comprises:
[0067] S410, acquiring an audio segment in the detected audio, and identifying a music event in the audio segment.
[0068] S420, determining metadata information matched with the music event.
[0069] S430, merging the music events corresponding to the same metadata information according to the start and end time stamps of the music events in each audio segment to obtain the statistical data in the detected audio.
[0070] The start and end time stamps of the music events refer to start time stamps and end time stamps of the music events. Specifically, if the metadata information of the music events is the same, it indicates that the music events are part of the same music or song, and the music events with the same metadata information can be merged. The statistical data in the detected audio is determined based on the merged music events, which avoids identification errors caused by division of the audio segments of the detected audio data, and improves the accuracy of the statistical data.
[0071] On the basis of the above embodiment, the statistical data in the detected audio is obtained by merging the music events corresponding to the same metadata information according to the start and end time stamps of the music events in each audio segment, including: for adjacent music events, if the metadata information corresponding to the adjacent music events is the same, and the interval duration of the adjacent music events is less than a second preset duration, the adjacent music events are merged; if the metadata information corresponding to the adjacent music events is different, or the metadata information corresponding to the adjacent music events is the same, and the interval duration of the adjacent music events is greater than or equal to the second preset duration, the adjacent music events are not merged.
[0072] The adjacent music events can be music events in adjacent audio segments, or adjacent music events in an audio segment, which is not limited herein.
[0073] For example, for adjacent music events, if the metadata information corresponding to the adjacent music events is the same, and the interval duration of the adjacent music events is less than a second preset duration, it indicates that the adjacent music events belong to the same song and the interval between the two music events is normal singing or playing pause, or identification error caused by division of the audio segment, and the adjacent music events can be merged to calibrate the identified music events; if the metadata information corresponding to the adjacent music events is different, it indicates that the adjacent music events do not belong to the same song, and the adjacent music events are not merged to distinguish and count different songs; if the metadata information corresponding to the adjacent music events is the same, and the interval duration of the adjacent music events is greater than or equal to the second preset duration, it indicates that the adjacent music events belong to the same song but the interval time is relatively long, for example, the same song is played twice, and the adjacent music events are not merged to avoid counting the same song with a long playing interval into the same statistical data.
[0074] The audio detection method provided by the embodiment of the present disclosure can obtain an audio segment containing merged music events by merging the music events corresponding to the same metadata information according to the start and end time stamps of the music events in each audio segment, so as to accurately obtain the music events in the detected audio and improve the accuracy of the statistical data.
[0075] On the basis of the above-mentioned embodiment, the audio to be detected is audio in a live video; the method further comprises: determining the viewing data of a live interval corresponding to each metadata information in the statistical data.
[0076] The live video can be a live video collected in real time, and can also be a historical live video. The audio to be detected is obtained by extracting audio from the live video. The use of music metadata in the live video is obtained by identifying and counting music events of the audio extracted from the live video.
[0077] The viewing data of the live interval refers to the viewing statistical data of the live room in a preset time period, which can include but is not limited to total viewing number, independent access number and average viewing time, etc. Specifically, the statistical data can be used as a viewing data matching condition, and the viewing data of the live interval is matched in the live database according to the matching condition, so as to accurately obtain the viewing data. The live database can include but is not limited to real-time statistical video viewing data. Optionally, the statistical data of the music metadata in the live video and the video viewing data corresponding to the statistical data are used to evaluate the flow diversion effect of the music metadata in the live video, or to predict the development trend of the music metadata.
[0078] Reference Figure 5 , Figure 5 The audio detection method flowchart provided by the embodiments of the present disclosure is provided, and a preferred example is provided on the basis of the above-mentioned embodiment to specifically describe the audio detection method of the above-mentioned embodiment.
[0079] As Figure 5 The method of the embodiment comprises:
[0080] Taking video live broadcast as an example, the audio in the live stream is divided to obtain a plurality of audio stream slices (i.e. the above-mentioned audio segments), and each audio stream slice can be processed in parallel.
[0081] The music event recognition on the audio stream slices specifically includes: extracting short-time features and long-time features from each audio stream slice, reducing the extracted short-time features and long-time features through a dimension reduction algorithm to remove redundant information of the short-time features and the long-time features, and obtaining main features. The dimension of the features after the dimension reduction is greatly reduced, and the performance is also improved to a certain extent. The main features are input into an SVM (Support Vector Machine) to obtain a recognition result. The short-time features at least include one of the following features: PLP (Perceptual Linear Predictive Coefficients), LPCC (Linear Predictive Cepstrum Coefficients), LFCC (Linear Frequency cepstral coefficients), Pitch, short-time energy (STE), sub-band energy distribution (SBED), brightness and bandwidth (BR and BW). The long-time features at least include one of the following features: spectrum flux (SF), long-term average spectrum (LTAS) and LPC entropy.
[0082] If the recognition result is a music event, it is further judged whether the duration of the current music event is greater than a first preset duration; if the recognition result is not a music event, the music event is unmarked. Further, if the duration of the current music event is greater than the first preset duration, the audio fingerprint feature of the music event is further extracted; if the duration of the current music event is not greater than the first preset duration, the music event is unmarked.
[0083] Further, the audio fingerprint feature of the music event is extracted through an audio fingerprint extraction algorithm, the audio fingerprint feature is matched in a fingerprint feature library, and metadata information is obtained. If the metadata information of adjacent music events is the same, i.e., the adjacent music events are the same song, and the interval duration of the adjacent music events is less than a second preset duration, it is indicated that the two belong to the same song and only normal singing or playing pause is in between, and the adjacent music events are merged; if the metadata information is the same, and the interval duration of the adjacent music events is not less than the second preset duration, it is indicated that the two belong to the same song but the pause time is long, and the adjacent music events are not merged. If the metadata information is not the same, i.e., the adjacent music events are not the same song, the adjacent music events are not merged.
[0084] After the adjacent music events are merged, the method further includes: obtaining statistical data of the merged music event, such as a play start time, a play end time, and the like of the music event. The statistical data can be used for music copyright billing.
[0085] Figure 6 is a structural schematic diagram of an audio detection apparatus provided by an embodiment of the present disclosure. As shown in Figure 6 The apparatus includes:
[0086] A music event identification module 610 is configured to obtain an audio segment in the detected audio, and identify a music event in the audio segment.
[0087] A statistical data determination module 620 is configured to determine metadata information matched with the music event, and determine statistical data in the detected audio based on the metadata information.
[0088] In some optional implementations of the embodiments of the present disclosure, the music event identification module 610 can be further configured to:
[0089] input the audio segment into a pre-trained music identification model to obtain a music event identification result output by the music identification model, wherein the music identification model is trained based on an audio sample and an event label corresponding to the audio sample.
[0090] In some optional implementations of the embodiments of the present disclosure, the apparatus can be further configured to:
[0091] determine whether a duration of the music event in the audio segment is greater than a first preset duration, and if not, cancel marking the music event.
[0092] In some optional implementations of the embodiments of the present disclosure, the statistical data determination module 620 can further include:
[0093] A fingerprint feature extraction unit is configured to extract an audio fingerprint feature of the audio segment containing the music event.
[0094] A metadata matching unit is configured to perform matching in a fingerprint feature library based on the audio fingerprint feature to determine metadata information matched with the music event, wherein the fingerprint feature library includes music metadata and corresponding fingerprint features.
[0095] In some optional implementations of the embodiments of the present disclosure, the fingerprint feature extraction unit can be further configured to:
[0096] According to the start and end time stamps of the music events in the audio segment, the audio segment is intercepted to obtain an intercepted audio segment, and an audio fingerprint feature of the intercepted audio segment is extracted; or
[0097] In the audio segment, audio data of the audio track in which the music event is located is extracted, and an audio fingerprint feature is extracted based on the audio data of the audio track in which the music event is located.
[0098] In some optional implementation of the embodiments of the present disclosure, the statistical data determination module 620 can further include:
[0099] A data merging unit is configured to merge music events corresponding to the same metadata information according to the start and end time stamps of the music events in each audio segment to obtain statistical data in the detected audio.
[0100] In some optional implementation of the embodiments of the present disclosure, the data merging unit can be further configured to:
[0101] For adjacent music events, if the metadata information corresponding to the adjacent music events is the same, and the interval duration of the adjacent music events is less than a second preset duration, the adjacent music events are merged.
[0102] If the metadata information corresponding to the adjacent music events is different, or the metadata information corresponding to the adjacent music events is the same, and the interval duration of the adjacent music events is greater than or equal to the second preset duration, the adjacent music events are not merged.
[0103] In some optional implementation of the embodiments of the present disclosure, the audio to be detected is audio in a live video; the apparatus can be further configured to determine viewing data of a live interval corresponding to each metadata information in the statistical data.
[0104] The audio detection apparatus provided in the embodiments of the present disclosure can execute the audio detection method provided in any of the embodiments of the present disclosure, and has the corresponding function modules and beneficial effects of executing the audio detection method.
[0105] It should be noted that each unit and module included in the apparatus is only divided according to the function logic, but is not limited to the above division, as long as the corresponding function can be implemented; in addition, the specific name of each functional unit is only for convenient distinction, and does not limit the protection scope of the embodiments of the present disclosure.
[0106] Reference will be made to the following description of the embodiments of the present disclosure Figure 7 which shows an electronic device (for example, a mobile phone) suitable for implementing the embodiments of the present disclosure. Figure 7The diagram below shows the structure of the terminal device (or server) 400. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. The electronic device shown in Figure 7 is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this disclosure.
[0107] like Figure 7 As shown, electronic device 400 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 402 or a program loaded from storage device 408 into random access memory (RAM) 403. RAM 403 also stores various programs and data required for the operation of electronic device 400. Processing device 401, ROM 402, and RAM 403 are interconnected via bus 404. Input / output (I / O) interface 405 is also connected to bus 404.
[0108] Typically, the following devices can be connected to I / O interface 405: input devices 406 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 407 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 408 including, for example, magnetic tapes, hard disks, etc.; and communication devices 409. Communication device 409 allows electronic device 400 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 An electronic device 400 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0109] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 409, or installed from storage device 408, or installed from ROM 402. When the computer program is executed by processing device 401, it performs the functions defined in the methods of embodiments of this disclosure.
[0110] The electronic device provided by the embodiments of the present disclosure and the audio detection method provided by the above embodiments belong to the same inventive concept, and the technical details not described in detail in the present embodiment can be referred to the above embodiments, and the present embodiment has the same beneficial effects as the above embodiments.
[0111] The present disclosure provides a computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the audio detection method provided by the above embodiments.
[0112] It should be noted that the computer readable medium of the present disclosure described above can be a computer readable signal medium or a computer readable storage medium or any combination of the two. The computer readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination thereof. More specific examples of computer readable storage media can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or component. In the present disclosure, the computer readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer readable program code. Such a propagated data signal can take various forms, including but not limited to an electromagnetic signal, an optical signal or any suitable combination thereof. The computer readable signal medium can also be any computer readable medium other than the computer readable storage medium, which can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or component. The program code contained in the computer readable medium can be transmitted by any suitable medium, including but not limited to a wire, a cable, an RF (radio frequency) or the like, or any suitable combination thereof.
[0113] In some embodiments, the client, server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future developed networks.
[0114] The computer readable medium can be included in the electronic device; or can exist separately from the electronic device.
[0115] The computer readable medium carries one or more programs which, when executed by the electronic device, cause the electronic device to:
[0116] obtain an audio segment in the detected audio, and identify a music event in the audio segment;
[0117] determine metadata information matched with the music event, and determine statistical data in the detected audio based on the metadata information.
[0118] Computer program code for carrying out operations of the present disclosure can be written in any of one or more programming languages or combinations of languages including object or visual programming languages specifically, assembly language, C, C++, Java, Visual Basic and / or JavaScript, etc. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0119] The flow diagrams and the block diagrams in the drawings are illustrations of possible architectures, functions, and operations for systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or operations, or combinations of special purpose hardware and computer instructions.
[0120] The units described in the embodiments of the present disclosure can be implemented in the form of software, or can be implemented in the form of hardware. In some cases, the name of the unit / module does not constitute a limitation on the unit itself.
[0121] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, example types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0122] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0123] According to one or more embodiments of the present disclosure, Example One provides an audio detection method, which comprises:
[0124] obtaining an audio segment in the detected audio, identifying a music event in the audio segment;
[0125] determining metadata information matched with the music event, and determining statistical data in the detected audio based on the metadata information.
[0126] According to one or more embodiments of the present disclosure, Example Two provides an audio detection method, which further comprises:
[0127] The identifying the music event in the audio segment comprises:
[0128] inputting the audio segment into a pre-trained music recognition model to obtain a music event recognition result output by the music recognition model, wherein the music recognition model is trained based on an audio sample and an event label corresponding to the audio sample.
[0129] According to one or more embodiments of the present disclosure, Example Three provides an audio detection method, further comprising:
[0130] After identifying the music event in the audio segment, the method further comprises:
[0131] Determining whether the duration of the music event in the audio segment is greater than a first preset duration, and if not, canceling the marking of the music event.
[0132] According to one or more embodiments of the present disclosure, Example Four provides an audio detection method, further comprising:
[0133] The determination of the metadata information matched with the music event comprises:
[0134] For an audio segment containing a music event, extracting an audio fingerprint feature of the audio segment;
[0135] Based on the matching of the audio fingerprint feature in the fingerprint feature library, determining the metadata information matched with the music event, wherein the fingerprint feature library includes music metadata and corresponding fingerprint features.
[0136] According to one or more embodiments of the present disclosure, Example Five provides an audio detection method, further comprising:
[0137] The extraction of the audio fingerprint feature of the audio segment comprises:
[0138] According to the start and end timestamps of the music event in the audio segment, the audio segment is intercepted to obtain an intercepted audio segment, and the audio fingerprint feature of the intercepted audio segment is extracted; or,
[0139] In the audio segment, the audio data of the audio track where the music event is located is extracted, and the audio fingerprint feature is extracted based on the audio data of the audio track where the music event is located.
[0140] According to one or more embodiments of the present disclosure, Example Six provides an audio detection method, further comprising:
[0141] The determination of the statistical data in the detected audio based on the metadata information comprises:
[0142] According to the start and end timestamps of the music event in each audio segment, the music events corresponding to the same metadata information are merged to obtain the statistical data in the detected audio.
[0143] According to one or more embodiments of the present disclosure, Example Seven provides an audio detection method, further comprising:
[0144] The statistical data in the detected audio is obtained by merging music events corresponding to the same metadata information according to the start and end time stamps of the music events in each audio segment, and includes:
[0145] For adjacent music events, if the metadata information corresponding to the adjacent music events is the same, and the interval duration of the adjacent music events is less than a second preset duration, the adjacent music events are merged.
[0146] If the metadata information corresponding to the adjacent music events is different, or the metadata information corresponding to the adjacent music events is the same, and the interval duration of the adjacent music events is greater than or equal to the second preset duration, the adjacent music events are not merged.
[0147] According to one or more embodiments of the present disclosure, Example Eight provides an audio detection method, further comprising:
[0148] The audio to be detected is audio in a live video.
[0149] The method further comprises:
[0150] Determining the viewing data of a live interval corresponding to each metadata information in the statistical data.
[0151] According to one or more embodiments of the present disclosure, Example Nine provides an audio detection device, which comprises:
[0152] A music event identification module, configured to acquire an audio segment in a detected audio, and identify a music event in the audio segment.
[0153] A statistical data determination module, configured to determine metadata information matched with the music event, and determine statistical data in the detected audio based on the metadata information.
[0154] The above description is only the preferred embodiments of the present disclosure and the explanation of the applied technical principles. Those skilled in the art should understand that the disclosure range involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features disclosed in the present disclosure (but not limited to) having similar functions to form technical solutions.
[0155] Moreover, while operations are depicted in a particular order, this should not be understood as requiring such an order nor infringing on the scope of the disclosure. Certain of the operations described in the discussion are combinable into a single operation, and certain operations can be separated into several operations. In some embodiments, the operations described in the discussion can be performed in an order different than presented in the discussion. In some embodiments, the operations described in the discussion can be performed concurrently. Also, while several specific implementation details are discussed in the discussion, these should not be interpreted as limiting the scope of the disclosure. Rather, certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination.
[0156] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. An audio detection method, characterized in that, include: Acquire audio segments from the detected audio and identify music events within those audio segments; Metadata information matching the music event is determined. Based on the start and end timestamps of the music events in each audio segment, music events corresponding to the same metadata information are merged to obtain statistical data in the detected audio. The music events to be merged are adjacent music events with the same metadata information and an interval of less than a second preset duration. The adjacent music events are music events located in adjacent audio segments. The statistical data includes the playback duration of each piece of music in the detected audio, the music playback start time, and the music playback stop time.
2. The method according to claim 1, characterized in that, The identification of music events in the audio segment includes: The audio segment is input into a pre-trained music recognition model to obtain the music event recognition result output by the music recognition model, wherein the music recognition model is trained based on the audio sample and the event label corresponding to the audio sample.
3. The method according to claim 1, characterized in that, After identifying musical events in the audio segment, the method further includes: Determine whether the duration of the music event in the audio segment is greater than a first preset duration; if not, unmark the music event.
4. The method according to claim 1, characterized in that, The metadata information that matches the music event includes: For audio segments containing music events, extract the audio fingerprint features of the audio segments; Based on the audio fingerprint features, a matching process is performed in the fingerprint feature database to determine the metadata information that matches the music event. The fingerprint feature database includes music metadata and corresponding fingerprint features.
5. The method according to claim 4, characterized in that, The extraction of audio fingerprint features from the audio segment includes: Based on the start and end timestamps of music events within the audio segment, the audio segment is truncated to obtain a truncated audio segment, and the audio fingerprint features of the truncated audio segment are extracted; or... The audio data of the track where the music event is located is extracted from the audio segment, and audio fingerprint features are extracted based on the audio data of the track where the music event is located.
6. The method according to claim 1, characterized in that, The method involves merging music events with the same metadata information based on the start and end timestamps of music events in each audio segment to obtain statistical data in the detected audio, including: For adjacent music events, if the metadata information corresponding to the adjacent music events is the same, and the interval between the adjacent music events is less than the second preset duration, then the adjacent music events will be merged. If the metadata information corresponding to the adjacent music events is different, or if the metadata information corresponding to the adjacent music events is the same, and the interval length of the adjacent music events is greater than or equal to the second preset length, then the adjacent music events will not be merged.
7. The method according to claim 1, characterized in that, The detected audio is the audio from the live video; The method further includes: Determine the viewing data for the live broadcast interval corresponding to each metadata information in the statistical data.
8. An audio detection device, characterized in that, include: The music event recognition module is used to acquire audio segments in the detected audio and recognize music events in the audio segments; The statistical data determination module is used to determine the metadata information that matches the music event. Based on the start and end timestamps of the music events in each audio segment, the music events corresponding to the same metadata information are merged to obtain the statistical data in the detected audio. The music events to be merged are adjacent music events with the same metadata information and an interval of less than a second preset duration. The adjacent music events are music events located in adjacent audio segments. The statistical data includes the playback duration of each piece of music in the detected audio, the music playback start time, and the music playback stop time.
9. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the audio detection method as described in any one of claims 1-7.
10. A storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to perform the audio detection method as described in any one of claims 1-7.
Citation Information
Patent Citations
Media identification system
GB202116661D0
Musical piece extraction device and musical piece recording device
JP2010078984A
Music analysis system and method for public spaces
WO2020176057A1