Data processing method, device, equipment, storage medium and computer program product
By performing segmented processing of audio and multi-threaded parallel processing of rhythm points, the problem of low efficiency of manually labeling rhythm points in the prior art is solved, and the automation and efficient calculation of music rhythm points are realized.
Patent Information
- Application Number
- CN202111022658.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-01
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-09-01
AI Technical Summary
The method of determining music rhythm points in the prior art relies on manual annotation, which has problems of poor subjectivity and low efficiency.
By performing segmented processing of the audio to be processed, the rhythm points of each audio segment are processed in parallel or concurrently using multiple data processing threads, and finally each rhythm point sequence is fused to obtain the target rhythm point sequence.
The automation and intelligence of music rhythm points are realized, which significantly improves the efficiency of determining rhythm points and reduces calculation time.
Smart Images

Figure CN114333899B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a data processing method, apparatus, device, storage medium and computer program product. Background Art
[0002] Musical beats are a regular combination of strong and weak beats. For example, each measure of 4 / 2 time has a strong beat and a weak beat that alternate regularly, and each measure of 4 / 4 time has a strong beat and three weak beats. Rhythm points refer to the time when music beats occur, and have a wide range of application scenarios, such as designing games based on music beats, making point-point videos in the direction of audio visualization, and switching videos according to the changes in music beats, and adding different sound effects to the original music in the direction of changing music styles to enhance the music atmosphere.
[0003] Manual labeling is a way to determine the rhythm points of music. However, due to subjective factors, different labeling standards will lead to inconsistent music rhythm points, and the required human resources are high and inefficient. Therefore, how to efficiently determine the rhythm points of music is of great research significance. Summary of the invention
[0004] The embodiments of the present application provide a data processing method, apparatus, device, storage medium and computer program product, which can realize the automation and intelligence of determining the rhythm points in the audio, thereby effectively improving the efficiency of determining the rhythm points in the audio.
[0005] On the one hand, an embodiment of the present application provides a data processing method, including:
[0006] Segment the audio to be processed to obtain at least two audio segments;
[0007] Using at least two data processing threads to process at least two audio clips to obtain a rhythm point sequence of each audio clip;
[0008] The rhythm point sequences of each audio clip are fused to obtain a target rhythm point sequence of the audio to be processed.
[0009] An embodiment of the present application provides a data processing device, including:
[0010] A segmentation module, used for segmenting the audio to be processed to obtain at least two audio segments;
[0011] A processing module, used to process at least two audio clips using at least two data processing threads to obtain a rhythm point sequence of each audio clip;
[0012] The fusion module is used to fuse the rhythm point sequences of each audio clip to obtain a target rhythm point sequence of the audio to be processed.
[0013] On the one hand, an embodiment of the present application provides a computer device, including: a processor, a memory, and a network interface; the processor is connected to the memory and the network interface, wherein the network interface is used to provide network communication functions, the memory is used to store program codes, and the processor is used to call program codes to execute the data processing method in the embodiment of the present application.
[0014] On the one hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, the data processing method in the embodiment of the present application is executed.
[0015] Accordingly, an embodiment of the present application provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the data processing method provided in one aspect of an embodiment of the present application.
[0016] In an embodiment of the present application, the audio to be processed (or the original audio) is segmented to obtain audio segments, and then multiple data processing threads can be used to simultaneously process the multiple audio segments to obtain the rhythm point sequence of each audio segment. In this way, the rhythm point sequence of the original audio is decomposed into multiple rhythm point sequences that are calculated simultaneously. By fusing these rhythm point sequences, the rhythm point sequence of the original audio is quickly obtained, thereby realizing the automation and intelligence of determining the rhythm points in the audio, and effectively improving the efficiency of determining the rhythm points. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 is an architecture diagram of a data processing system provided in an embodiment of the present application;
[0019] Figure 2 It is a flowchart of a data processing method provided in an embodiment of the present application;
[0020] Figure 3This is a schematic diagram of a process for determining a music rhythm point provided by an embodiment of the present application;
[0021] Figure 4 It is a flowchart of another data processing method provided in an embodiment of the present application;
[0022] Figure 5 It is a flowchart of another data processing method provided in an embodiment of the present application;
[0023] Figure 6 It is a schematic diagram of the result of a spectrogram processing process provided in an embodiment of the present application;
[0024] Figure 7 is a structural schematic diagram of a data processing device provided in an embodiment of the present application;
[0025] Figure 8 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0026] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0027] Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool and be used on demand, which is flexible and convenient. With the rapid development and application of the Internet industry, each item may have its own identification mark in the future, and all need to be transmitted to the background system for logical processing. Data of different levels will be processed separately. All kinds of industry data need strong system backing support, which can only be achieved through cloud computing.
[0028] Cloud computing is a computing model that distributes computing tasks on a resource pool composed of a large number of computers, so that various application systems can obtain computing power, storage space and information services as needed. According to the logical function division, the PaaS (Platform as a Service) layer can be deployed on the IaaS (Infrastructure as a Service) layer, and the SaaS (Software as a Service) layer can be deployed on the PaaS layer. SaaS can also be directly deployed on IaaS. PaaS is a platform for software operation, such as databases, web containers, etc. SaaS is a variety of business software, such as web portals, SMS mass senders, etc. Generally speaking, SaaS and PaaS are upper layers relative to IaaS. The data processing solution provided in this application can be a function provided by the PaaS service, which can support third-party applications to call relevant interfaces to determine music rhythm points by executing the data processing solution, and third-party applications can use the music rhythm points to make card point music, card point videos, etc.
[0029] See also Figure 1 , Figure 1 is an architecture diagram of a data processing system provided in an embodiment of the present application, such as Figure 1 As shown, it includes a terminal device 101 and a server 100. The terminal device 101 and the server 100 can be connected to each other by wire or wirelessly.
[0030] The terminal device 101 can collect audio data through a sound pickup device or collect video data in combination with a shooting device to generate an audio or video file. The audio included in the audio or video may be music. The terminal device 101 uploads the audio or video file through a running application client (such as an online web application or a third-party APP), or inputs a local storage path indicating the file. The server 100 can directly obtain the audio and video data or obtain the audio and video data according to the local storage path.
[0031] The server 100 can obtain audio data from the terminal device 101 or other databases (for videos, the audio data included therein is extracted) and segment the audio data to obtain multiple audio segments, and then enable multiple data processing threads in the server 100 to process the audio segments in parallel or concurrently to obtain a rhythm point sequence, and finally merge the rhythm point sequences obtained from the audio segments according to corresponding rules to obtain the final music rhythm point. The processing result is then returned to the terminal device 101.
[0032] In this way, it is not necessary to perform serial calculations on the entire piece of music to locate the rhythm points. Instead, the rhythm points of multiple audio clips can be calculated simultaneously, which speeds up the positioning of the rhythm points and greatly reduces the calculation time.
[0033] It is understandable that the above-mentioned terminal device 101 can be a smart phone, a tablet computer, a car terminal, an intelligent voice interaction device, a smart home appliance, a smart wearable device, a personal computer and other devices, and the server 100 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms.
[0034] Further, for ease of understanding, the embodiments mentioned below in this application are all based on the server (such as the above Figure 1 The server 100 in the corresponding embodiment is used as an example for explanation. Figure 2 , Figure 2 1 is a flow chart of a data processing method provided in an embodiment of the present application. The data processing method may at least include the following steps S101-S103, wherein:
[0035] S101, segment the audio to be processed to obtain at least two audio segments.
[0036] In one embodiment, the audio to be processed may be audio obtained by the server from a database, or audio uploaded to the server by a terminal device, or audio extracted from a video. The audio is digitized voice data and may be any of music, reading audio, noise audio, and silent audio. The content of the audio to be processed and the method of obtaining the audio are not limited here. Segmenting the audio to be processed as a pre-processing step of data processing can obtain two or more audio segments, each of which may be uniform or non-uniform. Correspondingly, the segmentation method may be equal or unequal cutting according to any time length, which is not limited here.
[0037] Optionally, the specific implementation steps of the segmentation processing may include: obtaining the audio to be processed, performing average segmentation processing on the audio to be processed according to the set number of segments and the time length of the audio to be processed, and obtaining at least two audio segments. The set number of segments is recorded as N, and the value range can be 3 to 5. The time length of the audio to be processed is recorded as T (abbreviated as duration). The average segmentation processing can obtain an audio segment with a duration of T / N. It should be noted that if the duration of the audio to be processed cannot be divided by N, T / N can be rounded up, and the audio to be processed is segmented according to the duration after rounding up. The audio segment at the end of the time sorting after segmentation is usually shorter than the duration. For example, a 20-second piece of music (time range is 0s-20s) needs to be divided into 3 equal parts (that is, the set number of segments is 3). Then, the audio segments can be segmented according to the duration of each audio segment of 7 seconds (20 / 3 rounded up), and the obtained audio segments are recorded as V1, V2, and V3 respectively, among which the time range of V1 is 0s-7s, the time range of V2 is 8s-15s, and the time range of V3 is 16s-20s. The duration of the last audio segment V3 is 4s, which is less than the evenly divided duration of 7s.
[0038] As an optional implementation method, the audio to be processed can also be evenly divided according to the specified audio segment length. The specified audio segment length is recorded as t, and the number of audio segments obtained is T / t. If the result of T / t is not an integer, T / t can be rounded up or down under certain conditions to determine whether the audio segment arranged last in time is retained. For example, if the remainder obtained by T / t is less than the specified audio segment length t, the last audio segment is discarded, and the duration of each audio segment retained is the same. For example, the specified audio segment length t is 6s, and the duration of the audio to be processed is 20s. According to T / t, four audio segments can be obtained. The duration of the last audio segment is 2s, which is less than the specified audio segment length 6s, so the audio segment is deleted.
[0039] When the audio to be processed is music, the audio fragments obtained after segmentation processing are called sliced music (or slices). Since the specific position of each music beat is related to a segment of audio around the beat, and has nothing to do with the long-distance global music signal, subsequent processing can still accurately locate the rhythm points (or beats) of the segmented sliced music.
[0040] S102: Process at least two audio clips using at least two data processing threads to obtain a rhythm point sequence of each audio clip.
[0041] In one embodiment, the server may start at least two data threads (referred to as threads) to process the audio segments obtained by the above segmentation processing. Optionally, one data processing thread may process one audio segment (one-to-one), and multiple data processing threads may process one audio segment (many-to-one). In the case of many-to-one, multiple threads may process sub-segments of the audio segment one-to-one, or multiple data processing threads may process the same audio segment at the same time and then select the best processing result. Regardless of which method is adopted, such a processing method can improve the speed of calculating the rhythm point sequence.
[0042] The processing of multiple audio clips by each data processing thread can be parallel processing or concurrent processing, wherein parallel processing refers to two or more data processing threads executing simultaneously on different CPU (Central Processing Unit) resources at the same time, that is, the threads in the parallel state are distributed on different CPUs (or multiple processes are assigned to different CPU resources), so that there is no time difference in the execution of the data processing threads, and there is no competition for CPU resources. Concurrent processing refers to multiple threads being between started and completed in the same period of time. At the micro level, the concurrent processing of multiple threads can be treated as a sequence, including queuing, waking up, and executing. At the macro level, multiple threads that arrive almost at the same time look like they are being processed at the same time, so multiple processes are not carried out at the same time, but they are carried out at the same time, and concurrently use a CPU resource, and each thread needs to compete for CPU resources. Whether it is concurrent or parallel processing, compared with the serial processing of audio in time order, the use of multi-threaded processing of audio clips can maximize the use of CPU resources, improve the speed of data processing, and thus efficiently determine the rhythm point sequence.
[0043] It should be noted that the specific processing algorithm executed by the data processing thread is not limited in the embodiment of the present application, that is, any method that can accurately obtain the rhythm point sequence can be applied to this step. The rhythm point sequence of the audio clip can refer to a sequence in which the beat time in the audio clip is a strong beat or a weak beat, that is, a beat sequence in which the rhythm points are arranged in chronological order, and the time unit is a frame. For example, the first frame can be a strong beat, and the second frame, the third frame, and the fourth frame can be a sequence of weak beats. Here, the time corresponding to one frame is the time of one beat.
[0044] Normally, the time to calculate the rhythm points of music is proportional to the duration of the music, which is about 1 / 4 of the duration of the music. For example, when a 4-minute music is processed by a serial algorithm without segmentation, the user needs to wait for about 1 minute to get the rhythm points. Such a wait is very long for the user. Therefore, the original audio can be divided into equal parts in the pre-processing stage, and then the audio segments can be processed in parallel or concurrently using multi-threading, and the calculation can be accelerated to determine the rhythm points at the same time. Experiments have shown that this can achieve a 2 to 4 times acceleration effect on the speed of rhythm point determination while ensuring the accuracy of the rhythm points, that is, compressing the time to half or a quarter of the original.
[0045] S103, fusing the rhythm point sequences of the audio clips to obtain a target rhythm point sequence of the audio to be processed.
[0046] In one embodiment, the target rhythm point sequence is the rhythm point sequence of the audio to be processed restored by using the rhythm point sequences of each audio segment. The rhythm point sequence can be used to represent the beat of the entire audio to be processed. The final rhythm point sequence can be obtained by fusing the rhythm point sequences of each slice. For the specific fusion method, in addition to directly combining the rhythm points at each moment, the embodiment of the present application can also design a specific fusion rule to ensure the rationality and accuracy of the final rhythm point sequence. For the fusion rule, please refer to the following Figure 4 The content of the corresponding embodiment may also be implemented in other ways, which are not limited here.
[0047] Based on the above data processing scheme, in most cases the processing object (i.e. the audio to be processed) is music, and its processing flow can be summarized as follows: Figure 3 The content shown includes four steps: music input, pre-processing, rhythm point calculation and post-processing: a) music input: input video or audio files, extract audio tracks as music input for the algorithm; b) pre-processing: divide the original music into equal parts to obtain M evenly sliced music, where M can be the same as the set number of segments N, and send them to the rhythm point calculation process at the same time; c) rhythm point calculation: parallel processing of the M sliced music input in step b) and the position of the rhythm points of each slice; d) post-processing: merge the rhythm points of each slice obtained in step c) to obtain the final music rhythm point. This solution can be used to quickly calculate music rhythm points. For example, the calculation process that originally takes 1 minute can be compressed to less than 15 seconds under the premise of ensuring the effect, greatly reducing user waiting time and improving user experience.
[0048] The data processing scheme provided by the present application is applied to the automatic and rapid calculation of music rhythm points in various forms. Taking the web interface as an example, the specific operation steps and product manifestation can be as follows: first, the user uploads a video or audio URL (uniform resource locator), and the relevant algorithm in the background server calculates the rhythm points of the music, and then returns the music rhythm point information in the form of json through the web interface, such as encapsulating the rhythm point information in a json file and returning it. The function of determining the music rhythm point can be deployed on the PAAS (Platform as a Service) service platform, and a calling interface can be provided for third-party applications. The interface is used to realize the transmission between data, so that the third-party application can obtain and parse the music rhythm point information in the form of json, obtain the location of the beat, and apply it to the developed function. When the user uses the function developed based on the music rhythm point in the third-party application (such as the card point video production), the video or audio file can be directly uploaded in the online application or application of the web end, and the background automatically extracts the indication address of the video or audio file (such as the above URL), so that the PAAS service obtains the video or audio data according to the indication address and processes it, and returns the processing result to the background server of the third-party application to realize the corresponding function.
[0049] In summary, the embodiments of the present application have at least the following advantages:
[0050] By processing the audio to be processed in segments to obtain audio segments, multiple data processing threads are used to process the audio segments in parallel or concurrently, the speed of calculating the rhythm point sequence of each audio segment is improved, and then the rhythm point sequence is quickly obtained by fusion, so as to realize the automation and intelligence of determining the rhythm points in the audio. Due to the regularity of the position of the beat (i.e., related to the close-range audio signal), the accuracy of the rhythm point sequence can be guaranteed. Parallel or concurrent calculation not only greatly reduces the processing time, effectively improves the efficiency of determining the rhythm points, but also improves the utilization of resources.
[0051] See also Figure 4 , Figure 4 2 is a flow chart of a data processing method provided in an embodiment of the present application. The data processing method may at least include the following steps S201-S205, wherein:
[0052] S201, segment the audio to be processed to obtain at least two audio segments.
[0053] S202: Process at least two audio clips using at least two data processing threads to obtain a rhythm point sequence of each audio clip.
[0054] The optional implementation of steps S201-S202 can be found in the above Figure 2 S101-S102 in the corresponding embodiment will not be described in detail here.
[0055] S203, obtaining a first rhythm point in the first rhythm point sequence, and obtaining a second rhythm point in the second rhythm point sequence.
[0056] In one embodiment, the first rhythm point sequence and the second rhythm point sequence are rhythm point sequences in any combination of adjacent audio segments in at least two audio segments, and the first rhythm point sequence is arranged before the second rhythm point sequence in time order; the first rhythm point is the last rhythm point in the first rhythm point sequence, and the second rhythm point is the first rhythm point in the second rhythm point sequence. The combination of adjacent audio segments can be any two audio segments that are adjacent in time among the multiple audio segments obtained by the audio segmentation to be processed, such as audio segment V 1 (Time range 11s-30s) and audio clip V 2 (Time range is 31s-40s) can be used as a combination of adjacent audio clips. The rhythm point sequence of each audio clip is sorted according to time, which is the same as the sorting of each audio clip. For example, the above audio clip V 1 The rhythm point sequence is in the audio clip V 2 According to the definition of the first rhythm point sequence and the second rhythm point sequence, the audio segment V in the above example is 1 The rhythm point sequence of can be used as the first rhythm point sequence, the audio segment V 2 The rhythm point sequence of can be used as the second rhythm point sequence. The first rhythm point and the second rhythm point are the positions where two rhythm point sequences intersect, and the first rhythm point and the second rhythm point are the rhythm points at the end of the previous section and the beginning of the next section respectively.
[0057] S204: According to the first rhythm point and the second rhythm point, a fusion process is performed on the first rhythm point sequence and the second rhythm point sequence to obtain a fused rhythm point sequence.
[0058] In one embodiment, considering that the rhythm points at non-boundary positions are far apart and have little influence, the solution can fuse the results of two adjacent rhythm point sequences at the boundary to obtain a fused rhythm point sequence. The result of two adjacent rhythm point sequences at the boundary can be to select rhythm points according to the first rhythm point and the second rhythm point, and then fuse the first rhythm point sequence and the second rhythm point sequence to obtain a fused rhythm point sequence. It should be noted that the number of rhythm points included in the fused rhythm point sequence can be equal to or less than the sum of the number of rhythm points included in the first rhythm point sequence and the second rhythm point sequence. For the specific rules for fusing two adjacent rhythm point sequences, please refer to the following content.
[0059] Optionally, since adjacent rhythm point sequences may locate the same rhythm point, and the beat of the rhythm point may be misjudged, the rhythm points may be screened and combined according to the following rules to ensure the rationality and accuracy of the fusion of the rhythm point sequence. The specific steps may include: obtaining the time interval between the first rhythm point and the second rhythm point, and obtaining the beat type of the first rhythm point and the second rhythm point; if the time interval is greater than or equal to the interval threshold, and the beat type of the first rhythm point and the second rhythm point meets the beat setting rule, the second rhythm point sequence is fused with the first rhythm point sequence to obtain a fused rhythm point sequence.
[0060] The interval threshold here can be taken as 1 / 2 of the shortest slice rhythm point interval, and the shortest slice rhythm point interval refers to the minimum time interval of adjacent rhythm points selected in the rhythm point sequence of each audio clip. The specific implementation method can be that for the rhythm point sequence of any audio clip, the time interval between any two adjacent rhythm points is determined, and then the minimum time interval is selected, so that the rhythm point sequence of each audio clip corresponds to a minimum value of a rhythm point (time) interval. For example, there are 3 rhythm point sequences, corresponding to 3 minimum time intervals, and then the minimum time intervals of the rhythm point sequences of all audio clips are compared, and one is selected from the three minimum time intervals as mentioned above, and the final minimum time interval is used as the shortest slice rhythm point interval. Beat types include strong beats and weak beats. The beat setting rule can mean that a strong beat will be followed by a weak beat, and a strong beat and a strong beat will not appear adjacent to each other, that is, the beat setting rule does not include a regular combination of "strong strong weak".
[0061] The time position of the rhythm point and the beat type it belongs to are the key to forming the rhythm point sequence. The time position of the rhythm point can determine the time interval between adjacent rhythm points, and the beat type of the rhythm point can determine the beat rule of adjacent rhythm points. The rhythm points are selected by judging whether the time interval between adjacent rhythm points is greater than or equal to the time threshold and whether the beat rule satisfies the beat setting rule (i.e., whether the beat type satisfies the beat setting rule). Under the condition that the time interval is greater than or equal to the time threshold and the beat type satisfies the beat setting rule, the two adjacent rhythm point sequences can be directly spliced to obtain a fused rhythm point sequence, the number of which is the sum of the number of rhythm points included in the two adjacent rhythm point sequences.
[0062] If any of the above conditions cannot be met, further processing is required, namely: if the time interval is less than the interval threshold, or the beat types of the first rhythm point and the second rhythm point do not meet the beat setting rule, the second rhythm point is deleted from the second rhythm point sequence; a new first rhythm point is determined from the deleted second rhythm point sequence; based on the first rhythm point and the new first rhythm point, the first rhythm point sequence and the deleted second rhythm point sequence are fused to obtain a fused rhythm point sequence.
[0063] That is to say, the interval between the rhythm point at the end of the previous segment (i.e., the first rhythm point of the first rhythm point sequence) and the rhythm point at the beginning of the next segment (i.e., the second rhythm point of the second rhythm point sequence) is particularly small, less than the interval threshold (e.g., half of the interval of the shortest slice rhythm point), or the beat types of the two rhythm points do not meet the beat setting rule. The specific screening process can be to delete the second rhythm point in the second rhythm point sequence, and then determine the new first rhythm point from the rhythm points included in the remaining rhythm point sequence, and use the same method to obtain the beat type and time interval of the new first rhythm point and the first rhythm point, compare the time interval and the interval threshold, and whether the beat type meets the beat setting rule, and then decide whether to delete the second rhythm point. If any of the above conditions cannot be met, continue to delete the new first rhythm point, and then determine the new first rhythm point from the remaining rhythm points, and repeat in sequence until both of the above conditions are met, and then perform fusion processing to obtain a fused rhythm point sequence. In short, the rhythm points at the beginning of the next segment are deleted one by one until the interval between the first rhythm point in the remaining rhythm point sequence and the rhythm point at the end of the previous segment is greater than the interval threshold and the first rhythm point in the remaining rhythm point sequence is a weak beat (when the first rhythm point is a strong beat), then the deletion can be stopped.
[0064] Optionally, if the time interval or beat type does not meet the conditions, the first rhythm point in the first rhythm point sequence can also be deleted, and then a new last rhythm point is determined from the remaining rhythm points in the first rhythm point sequence, and the fused rhythm point sequence is determined based on the new last rhythm point and the second rhythm point. That is, the rhythm points at the end of the previous segment are deleted in sequence until the time interval between the last rhythm point in the remaining rhythm point sequence and the rhythm point at the beginning of the next segment is greater than the interval threshold and the first rhythm point in the remaining rhythm point sequence belongs to a weak beat, then the deletion can be stopped. The specific content is similar to the above and will not be repeated here.
[0065] It should be noted that deleting the rhythm points that do not meet the conditions in the rhythm point sequence does not affect the time position of other rhythm points. In addition, the reason for deleting a rhythm point that is smaller than the interval threshold is that two adjacent slice rhythm points with too small an interval are actually located at the same rhythm point, and deletion can remove duplicates. The reason for not meeting the beat setting rule is to use the a priori fact of music beats, that is, strong beats and strong beats will not appear adjacent to each other.
[0066] S205: Determine a target rhythm point sequence of the audio to be processed according to the fused rhythm point sequence.
[0067] In one embodiment, according to the implementation method of step S204, the rhythm point sequences of any two adjacent audio clips can be fused to obtain a fused rhythm point sequence. For multiple audio clips, there can be one or more fused rhythm point sequences. For each fused rhythm point sequence, the target rhythm point sequence can be determined by splicing each fused rhythm point sequence in the same way as the fused rhythm point sequence, that is, the time interval and beat type of the rhythm points at the junction of adjacent rhythm point sequences are detected to see if they meet the conditions. If they meet the conditions, they can be directly spliced. If not, they need to be processed before splicing. No further details are given here.
[0068] All audio segments obtained by segmenting the audio to be processed are recorded as an audio segment set S = {s 1 ,s 2 ,…,s L}, including L audio segments (where L can be equal to the number of segments N set in the above embodiment). Optionally, for all audio segments of the audio to be processed, each two adjacent audio segments can be combined to obtain an adjacent audio segment combination {s 1 ,s 2}、{s 3 ,s 4}…{s L-1 ,s L}, according to the above rules, the rhythm point sequence of each audio clip combination can be fused to obtain the corresponding fused rhythm point sequence. In the corresponding ideal case, L / 2 rhythm point sequences can be obtained by non-repeating combinations of two audio clips without any remaining audio clips. For example, L=4, the corresponding fused rhythm point sequence is 2, and then these two fused rhythm point sequences are regarded as the first rhythm point sequence and the second rhythm point sequence, and are fused according to the same rules. The final rhythm point sequence is recorded as the target rhythm point sequence of the audio to be processed. If the number of audio clips is an odd number, for example, L=3, for the combination of adjacent audio clips {s 1 ,s 2}、{s 3}, get and the remaining single audio segment {s 3}, at this time, the fused rhythm point sequence can be regarded as the first rhythm point sequence, and the rhythm point sequence of the separate audio clip can be regarded as the second rhythm point sequence, and they are fused according to the same rules.
[0069] Optionally, the method for determining the target fused rhythm point sequence may also be to determine the valid rhythm points in the rhythm point sequences of each audio clip, and then combine the valid rhythm points of all rhythm point sequences into one rhythm point sequence, wherein the method for determining the valid rhythm points may be to adopt the aforementioned rule for fusing two rhythm point sequences. Exemplarily, for the rhythm point sequences of three audio clips, all the rhythm points included in the rhythm point sequence of the first audio clip may be regarded as valid rhythm points, and the rhythm points in the rhythm point sequence of the second audio clip are screened based on the last rhythm point of the rhythm point sequence, specifically, the first rhythm point of the second audio clip is selected, that is, whether the time interval and beat type between the rhythm point and the last rhythm point in the previous rhythm point sequence meet the conditions. If not, the rhythm point needs to be deleted again until it meets the conditions, so that the valid rhythm points in the rhythm point sequence of the second audio clip may be all or part of the rhythm points, and for the rhythm point sequence of the third audio clip, the valid rhythm points are screened based on the last rhythm point of the rhythm point sequence of the second audio clip according to the same rule. Finally, these screened rhythm point sequences are combined to obtain the target fused rhythm point sequence. It can be found that this method omits the step of obtaining the fusion rhythm point sequence in the middle, and can directly and quickly obtain the target result. It should be noted that the above method of determining the effective rhythm points can be processed in parallel to quickly determine the rhythm point sequence of the entire audio segment.
[0070] In summary, the embodiments of the present application have at least the following advantages:
[0071] The rhythm point sequences of each audio clip are fused by judging the rhythm point information (including time interval and beat type) of adjacent rhythm point sequences at the boundary position, wherein the same rhythm points at the boundary of the two previous and next rhythm point sequences can be deduplicated according to corresponding rules, and the rationality of the beat combination rules can be verified by using the prior facts of the beats. Such a fusion method can automatically determine the rhythm points in the audio and ensure the accuracy and rationality of the final rhythm point sequence.
[0072] See also Figure 5 , Figure 5 301 is a flow chart of a data processing method provided in an embodiment of the present application. The data processing method may at least include the following steps S301-S304, wherein:
[0073] S301, segment the audio to be processed to obtain at least two audio segments.
[0074] For details on this step, see Figure 2 The content of step S101 of the corresponding embodiment is not described in detail here.
[0075] S302: Use at least two data processing threads to perform spectrum conversion processing on at least two audio segments respectively to obtain a spectrogram of each audio segment.
[0076] In one embodiment, the data processing thread may process the audio clip in a one-to-one manner or in a many-to-one manner, and the details may refer to the contents of the aforementioned embodiments. In the process of calculating the rhythm points, the data processing thread may process the audio data or intermediate data (such as a spectrogram) in parallel or concurrently, which is not limited here. The audio clip processed by the data processing thread may be the audio PCM (Pulse Code Modulation) data extracted after the input audio to be processed is segmented, that is, the standard digital audio data converted from the analog signal after sampling, quantization, and encoding. The same spectrum transformation process is applied to each audio clip to obtain the spectrogram of each audio clip, and the spectrum change process here may be a fast Fourier transform or a short-time Fourier transform (STFT), which is not limited here.
[0077] S303: Use at least two data processing threads to respectively call the beat detection model to perform rhythm point detection processing on the spectrograms of at least two audio clips to obtain a rhythm point sequence of each audio clip.
[0078] In one embodiment, the data processing thread, as a context execution instruction, can call the beat detection model to perform rhythm point detection processing on the spectrogram, and the specific processing flow for the spectrogram of each audio segment can be the same. Therefore, taking the spectrogram of any audio segment to perform rhythm point detection processing to obtain a rhythm point sequence as an example, the specific steps may include: for the spectrogram of any audio segment in at least two audio segments, using the target data processing thread in at least two data processing threads to call the beat detection model to process the spectrogram of any audio segment, determine the feature information of any audio segment, and determine the beat probability value of the audio corresponding to each beat unit time in any audio segment according to the feature information, and determine the beat type of the audio corresponding to each beat unit time according to the probability threshold and the beat probability value, and determine the rhythm point sequence of any audio segment according to the beat type of the audio corresponding to each beat unit time; wherein the beat probability value includes a strong beat probability value and a weak beat probability value, and the beat type includes a strong beat or a weak beat.
[0079] Among them, the target data processing thread can be one or more of at least two data processing threads, which is not limited here. When the target data processing thread is multiple data processing threads, the processing object can be the spectrogram of the same audio clip. The beat detection model called by the data processing thread is a trained beat detection model. The training process of the beat detection model belongs to supervised learning, that is, the model is trained using audio data with labels, including audio data with strong beats (downbeat), weak beats (beat) or no beats (non-beat). The beat detection model here can be a RNN (Recurrent Neural Network) time series depth model, or other models for processing time series data such as long short-term memory networks (Long short-term memory, LSTM), or models for processing image data, such as residual graph convolutional neural networks, which are not limited here. The spectrogram can obtain a rhythm point sequence through the model, and the intermediate processing process of the model includes determining feature information, determining beat probability values, and determining beat types in sequence.
[0080] In this embodiment, the spectrogram can be regarded as image data, and the processing of the spectrogram by calling the beat detection model (such as the above-mentioned RNN time series depth model) using the target data processing thread can be essentially understood as extracting features from the image to obtain feature information, and then obtaining the beat probability value of each beat unit time according to the feature information, including the strong beat probability value and the weak beat probability value. The beat unit time here refers to the time required for a beat, which can be obtained by detecting the BPM (Beat Per Minute) of the audio. For example, a BPM of 60 means that there are 60 beats per minute, and the time of one beat is 1 second. A BPM of 120 means that there are 120 beats per minute, and each beat is 0.5 seconds. To determine whether the beat of the audio corresponding to the beat unit time is a strong beat or a weak beat, a probability threshold can be set to select one of the two. For example, if the strong beat probability value in the beat probability value is greater than the probability threshold, the beat of the audio corresponding to the beat unit time is determined to be a strong beat, otherwise, it is a weak beat. In this way, the beat type of the audio of each beat unit time can be determined and then a rhythm point sequence can be obtained and output. It should be noted that the processing of an audio clip by the beat detection model can be time-sequential, that is, each time step (i.e., beat unit time) corresponds to the output of a beat probability value, and the beats included in the rhythm point sequence are arranged in chronological order.
[0081] See also Figure 6 , Figure 6It is a schematic diagram of the results of a spectrogram processing process provided in an embodiment of the present application. In order of the arrows, there are the spectrogram of the audio clip, the shallow features obtained by processing the spectrogram through the shallow network of the beat detection model, the relative deep features obtained by further processing the shallow features, and the final output rhythm point sequence. In the output rhythm point sequence, the peak value at the bottom corresponds to the strong beat probability value, and the peak value at the top corresponds to the weak beat probability value.
[0082] Furthermore, the above-mentioned beat detection model may include a first timing network and a second timing network, the spectrogram of any audio segment may include a first spectrogram and a second spectrogram, the beat probability value includes a first beat probability value and a second beat probability value, and the feature information of any audio segment is determined, and the beat probability value is determined based on the feature information. The implementation method may be: inputting the first spectrogram into the first timing network for processing to obtain the beat feature of any audio segment, and inputting the second spectrogram into the second timing network for processing to obtain the harmonic feature of any audio segment; determining the first beat probability value of the audio corresponding to each beat unit time in any audio segment based on the beat feature, and determining the second beat probability value of the audio corresponding to each beat unit time in any audio segment based on the harmonic feature.
[0083] The first spectrogram and the second spectrogram can be obtained by applying the same spectrum transformation method to the audio data, which can be a short-time Fourier transform. That is to say, the first spectrogram and the second spectrogram are two completely identical spectrum graphs. The first timing network and the second timing network can be two parallel recurrent neural networks, and the first spectrogram and the second spectrogram can be input into the first timing network and the second timing network respectively and simultaneously. The beat feature can be represented by calculating the multi-band spectrum flux, which expresses the rhythm content. The spectrum flux refers to the degree of change between adjacent frames of the signal, which can be used to calculate the characteristics of the starting point of the note. The specific processing of the first spectrogram in the first time series network can be to apply a logarithmic filter group to process the amplitude spectrogram (i.e., spectrogram) obtained by short-time Fourier transform to achieve amplitude compression and reduce computational complexity, and then for each frame, calculate the amplitude difference (i.e., spectral flux, used to represent the beat feature) between the sampling points of the current frame and the previous frame (i.e., between two adjacent frames). However, in order to further reduce the amount of data processed by the network, the time granularity can be increased to ensure that the subsequent network processing is faster. Specifically, the window average can be performed according to the position of the beat, that is, the average value of the frequency amplitude is calculated for the window of length Δb / np to synchronize the beat of the feature sequence, where Δb is the beat period (i.e., the time of one beat), and np is the number of beat divisions. Harmonic features can also be called harmony features, and the harmonic content of the entire slice can be represented by chromatic features. The chromatic features can be obtained based on the second spectrogram. Harmony here can be understood as adding chords to the position of the rhythm point.
[0084] In the specific training process, based on the Western music data set Ballroom, the beat features and harmonic features can be modeled using the first timing network and the second timing network, respectively, to obtain the corresponding activation function values (i.e., probability values), and the activation function values of the two are subsequently input into the decoding network (such as the dynamic Bayesian network) to decode the probability values into the time series of rhythm points. Therefore, the beat detection model includes not only the first timing network and the second timing network, but also the decoding network. The trained model is applied to the processing of the audio clip, that is, the first timing network and the second timing network can be used in parallel to process the spectrogram, respectively extract the beat features and harmonic features, and then determine the first beat probability value of the audio corresponding to each beat unit time according to the beat features and the second beat probability value of the audio corresponding to each beat unit time according to the harmonic features, where the first beat probability value and the second beat probability value both include the probability value indicating that the beat of the audio corresponding to the beat unit time is a strong beat or a weak beat. The subsequent first beat probability value and the second beat probability value are averaged and then input into the decoding network for processing to obtain the final rhythm point sequence. The principle of the decoding network for the averaged beat probability value can be the same as the above method, that is, using the probability threshold to determine whether the rhythm point is a strong beat or a weak beat. A more complex method can also be to combine the probability threshold with the beat law to determine the final rhythm point sequence. The beat law here means that strong beats and strong beats will not appear adjacent to each other.
[0085] Optionally, the spectrum transformation methods used by the first spectrogram and the second spectrogram may be different, that is, different spectrum transformation methods are used for the same audio clip, so that the first spectrogram and the second spectrogram obtained are also different. For example, the first spectrogram is obtained by fast Fourier transform or short-time Fourier transform, and the second spectrogram is obtained by constant-Q transform. Compared with the linear distribution spectrum diagram obtained by fast Fourier transform, constant-Q transform is a time-frequency transform with the same exponential distribution law. By calculating the second spectrogram, the amplitude value of the music signal at each note frequency can be directly obtained. The spectrum diagrams obtained by different spectrum transformations are respectively input into the first timing network and the second timing network, and the beat characteristics and harmonic characteristics can also be determined, thereby obtaining the first beat probability and the second beat probability.
[0086] S304: Fusing the rhythm point sequences of the audio clips to obtain a target rhythm point sequence of the audio to be processed.
[0087] For details on this step, see Figure 2 Corresponding to step S103 of the embodiment or Figure 4 The content of step S205 of the corresponding embodiment is not described in detail here.
[0088] In summary, the embodiments of the present application have at least the following advantages:
[0089] Using multiple data processing threads to call the beat detection model to perform parallel processing on the spectrogram of each audio clip can maximize the use of resources, effectively improve the efficiency of calculation, and greatly shorten the time required for data processing; by extracting the feature information of the spectrogram, including beat features and harmonic features, since the beat features can describe the starting point of the beat and the harmonic features can also describe the beat position, these feature information can be used to enrich the rhythm point information, ensure the accuracy of the beat of the audio corresponding to the beat unit time, and then ensure the accuracy of the final rhythm point sequence.
[0090] See also Figure 7 , Figure 7 : is a schematic diagram of the structure of a data processing device provided in an embodiment of the present application. The above-mentioned data processing device may be a computer program (including program code) running in a computer device, for example, the data processing device is an application software; the device may be used to execute the corresponding steps in the method provided in an embodiment of the present application. Figure 7 As shown, the data processing device 70 may include: a segmentation module 701 , a processing module 702 , and a fusion module 703 .
[0091] A segmentation module 701, used for segmenting the audio to be processed to obtain at least two audio segments;
[0092] A processing module 702 is used to process at least two audio clips using at least two data processing threads to obtain a rhythm point sequence of each audio clip;
[0093] The fusion module 703 is used to fuse the rhythm point sequences of the various audio clips to obtain a target rhythm point sequence of the audio to be processed.
[0094] In one embodiment, the fusion module 703 is specifically used to: obtain a first rhythm point in a first rhythm point sequence, and obtain a second rhythm point in a second rhythm point sequence; fuse the first rhythm point sequence and the second rhythm point sequence according to the first rhythm point and the second rhythm point to obtain a fused rhythm point sequence; determine a target rhythm point sequence of the audio to be processed according to the fused rhythm point sequence; wherein the first rhythm point sequence and the second rhythm point sequence are rhythm point sequences in any combination of adjacent audio clips in at least two audio clips, and the first rhythm point sequence is arranged before the second rhythm point sequence in chronological order; the first rhythm point is the last rhythm point in the first rhythm point sequence, and the second rhythm point is the first rhythm point in the second rhythm point sequence.
[0095] In one embodiment, the fusion module 703 is further specifically used to: obtain the time interval between the first rhythm point and the second rhythm point, and obtain the beat type of the first rhythm point and the second rhythm point, the beat type including a strong beat or a weak beat; if the time interval is greater than or equal to the interval threshold, and the beat type of the first rhythm point and the second rhythm point meets the beat setting rule, then the second rhythm point sequence is fused with the first rhythm point sequence to obtain a fused rhythm point sequence.
[0096] In one embodiment, the data processing device 70 further includes a deletion module 704 and a determination module 705, wherein:
[0097] A deleting module 704 is configured to delete the second rhythm point from the second rhythm point sequence if the time interval is less than the interval threshold, or the beat types of the first rhythm point and the second rhythm point do not satisfy the beat setting rule;
[0098] A determination module 705 is used to determine a new first rhythm point from the deleted second rhythm point sequence;
[0099] The fusion module 703 is used to fuse the first rhythm point sequence and the deleted second rhythm point sequence according to the first rhythm point and the new first rhythm point to obtain a fused rhythm point sequence.
[0100] In one embodiment, the segmentation module 701 is specifically used to: obtain the audio to be processed; and perform average segmentation processing on the audio to be processed according to the set segment number and the time length of the audio to be processed to obtain at least two audio segments.
[0101] In one embodiment, the processing module 702 is specifically used to: use at least two data processing threads to perform spectrum transformation processing on at least two audio clips respectively to obtain spectrograms of each audio clip; use at least two data processing threads to call the beat detection model to perform rhythm point detection processing on the spectrograms of at least two audio clips respectively to obtain rhythm point sequences of each audio clip.
[0102] In one embodiment, the processing module 702 is further specifically used to: for the spectrogram of any audio segment in at least two audio segments, use the target data processing thread in at least two data processing threads to call the beat detection model to process the spectrogram of any audio segment, determine the feature information of any audio segment, and determine the beat probability value of the audio corresponding to each beat unit time in any audio segment based on the feature information, and determine the beat type of the audio corresponding to each beat unit time based on the probability threshold and the beat probability value, and determine the rhythm point sequence of any audio segment based on the beat type of the audio corresponding to each beat unit time; wherein the beat probability value includes a strong beat probability value and a weak beat probability value, and the beat type includes a strong beat or a weak beat.
[0103] In one embodiment, the beat detection model includes a first timing network and a second timing network, the spectrogram of any audio segment includes a first spectrogram and a second spectrogram, and the beat probability value includes a first beat probability value and a second beat probability value; the processing module 702 is also specifically used to: input the first spectrogram into the first timing network for processing to obtain the beat feature of any audio segment; input the second spectrogram into the second timing network for processing to obtain the harmonic feature of any audio segment; determine the first beat probability value of the audio corresponding to each beat unit time in any audio segment according to the beat feature, and determine the second beat probability value of the audio corresponding to each beat unit time in any audio segment according to the harmonic feature.
[0104] It is understandable that the functions of the functional modules of the data processing device described in the embodiment of the present application can be specifically implemented according to the method in the above method embodiment, and the specific implementation process can refer to the relevant description of the above method embodiment, which will not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated.
[0105] See also Figure 8 , Figure 8 80 is a schematic diagram of the structure of a computer device 80 provided in an embodiment of the present application. The computer device 80 may include an independent device (e.g., one or more of a server, a node, a terminal, etc.), or may include components inside an independent device (e.g., a chip, a software module, or a hardware module, etc.). The computer device 80 may include at least one processor 801 and a communication interface 802. Further, optionally, the computer device 80 may also include at least one memory 803 and a bus 804. The processor 801, the communication interface 802, and the memory 803 are connected via a bus 804.
[0106] Among them, the processor 801 is a module that performs arithmetic operations and / or logical operations, and can specifically be a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a complex programmable logic device (CPLD), a coprocessor (assisting the central processing unit to complete corresponding processing and applications), a microcontroller unit (MCU) and other processing modules, or a combination of multiple of them.
[0107] The communication interface 802 may be used to provide information input or output for the at least one processor. And / or, the communication interface 802 may be used to receive data sent externally and / or send data externally, and may be a wired link interface including an Ethernet cable, etc., or a wireless link (Wi-Fi, Bluetooth, general wireless transmission, and other short-range wireless communication technologies, etc.) interface.
[0108] The memory 803 is used to provide a storage space in which data such as an operating system and a computer program can be stored. The memory 803 can be a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), or a portable read-only memory (CD-ROM), etc., or a combination of multiple thereof.
[0109] At least one processor 801 in the computer device 80 is used to call a computer program stored in at least one memory 803 to execute the aforementioned data processing method, such as the aforementioned Figure 2 , Figure 4 , Figure 5 The data processing method described in the illustrated embodiment.
[0110] In a possible implementation, the processor 801 in the computer device 80 is used to call a computer program stored in at least one memory 803 to perform the following operations:
[0111] Segment the audio to be processed to obtain at least two audio segments;
[0112] Using at least two data processing threads to process at least two audio clips to obtain a rhythm point sequence of each audio clip;
[0113] The rhythm point sequences of each audio clip are fused to obtain a target rhythm point sequence of the audio to be processed.
[0114] In one embodiment, when the processor 801 fuses the rhythm point sequences of the audio clips to obtain the target rhythm point sequence of the audio to be processed, it is specifically used to: obtain the first rhythm point in the first rhythm point sequence, and obtain the second rhythm point in the second rhythm point sequence; fuse the first rhythm point sequence and the second rhythm point sequence according to the first rhythm point and the second rhythm point to obtain a fused rhythm point sequence; determine the target rhythm point sequence of the audio to be processed according to the fused rhythm point sequence; wherein the first rhythm point sequence and the second rhythm point sequence are rhythm point sequences in any combination of adjacent audio clips in at least two audio clips, and the first rhythm point sequence is arranged before the second rhythm point sequence in chronological order; the first rhythm point is the last rhythm point in the first rhythm point sequence, and the second rhythm point is the first rhythm point in the second rhythm point sequence.
[0115] In one embodiment, the processor 801 fuses the first rhythm point sequence and the second rhythm point sequence according to the first rhythm point and the second rhythm point to obtain the fused rhythm point sequence, which is specifically used to: obtain the time interval between the first rhythm point and the second rhythm point, and obtain the beat type of the first rhythm point and the second rhythm point, where the beat type includes a strong beat or a weak beat; if the time interval is greater than or equal to the interval threshold, and the beat types of the first rhythm point and the second rhythm point meet the beat setting rule, then fuse the second rhythm point sequence with the first rhythm point sequence to obtain the fused rhythm point sequence.
[0116] In one embodiment, the processor 801 is further used to: if the time interval is less than the interval threshold, or the beat types of the first rhythm point and the second rhythm point do not satisfy the beat setting rule, then delete the second rhythm point from the second rhythm point sequence; determine a new first rhythm point from the deleted second rhythm point sequence; and, based on the first rhythm point and the new first rhythm point, fuse the first rhythm point sequence and the deleted second rhythm point sequence to obtain a fused rhythm point sequence.
[0117] In one embodiment, when the processor 801 performs segment processing on the audio to be processed and obtains at least two audio segments, it is specifically used to: obtain the audio to be processed; and perform average segment processing on the audio to be processed according to the set number of segments and the time length of the audio to be processed to obtain at least two audio segments.
[0118] In one embodiment, when the processor 801 uses at least two data processing threads to process at least two audio clips to obtain a rhythm point sequence of each audio clip, it is specifically used to: use at least two data processing threads to perform spectrum transformation processing on the at least two audio clips respectively to obtain a spectrogram of each audio clip; use at least two data processing threads to call the beat detection model to perform rhythm point detection processing on the spectrograms of the at least two audio clips respectively to obtain a rhythm point sequence of each audio clip.
[0119] In one embodiment, the processor 801 uses at least two data processing threads to respectively call the beat detection model to perform rhythm point detection processing on the spectrograms of at least two audio clips to obtain the rhythm point sequence of each audio clip. Specifically, it is used to: for the spectrogram of any audio clip in the at least two audio clips, use the target data processing thread in the at least two data processing threads to call the beat detection model to process the spectrogram of any audio clip, determine the feature information of any audio clip, and determine the beat probability value of the audio corresponding to each beat unit time in any audio clip according to the feature information, and determine the beat type of the audio corresponding to each beat unit time according to the probability threshold and the beat probability value, and determine the rhythm point sequence of any audio clip according to the beat type of the audio corresponding to each beat unit time; wherein the beat probability value includes a strong beat probability value and a weak beat probability value, and the beat type includes a strong beat or a weak beat.
[0120] In one embodiment, the beat detection model includes a first timing network and a second timing network, the spectrogram of any audio segment includes a first spectrogram and a second spectrogram, and the beat probability value includes a first beat probability value and a second beat probability value; when the processor 801 determines the feature information of any audio segment, and determines the beat probability value of the audio corresponding to each beat unit time in any audio segment based on the feature information, it is specifically used to: input the first spectrogram into the first timing network for processing to obtain the beat feature of any audio segment; input the second spectrogram into the second timing network for processing to obtain the harmonic feature of any audio segment; determine the first beat probability value of the audio corresponding to each beat unit time in any audio segment based on the beat feature, and determine the second beat probability value of the audio corresponding to each beat unit time in any audio segment based on the harmonic feature.
[0121] It should be understood that the computer device 80 described in the embodiment of the present application can execute the above Figure 2 , Figure 4 as well as Figure 5 The description of the data processing method in the corresponding embodiment can also be performed as described above. Figure 7 The description of the data processing device 70 in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of the same method will not be repeated here either.
[0122] In addition, it should be pointed out here that: the embodiment of the present application also provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program executed by the computer device 80 of the data processing method mentioned above, and the computer program includes program instructions. When the processor executes the program instructions, the above-mentioned data processing method can be executed. Figure 2 , Figure 4 as well as Figure 5 The description of the above data processing method in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of the same method will not be repeated. For technical details not disclosed in the computer-readable storage medium embodiment involved in this application, please refer to the description of the method embodiment of this application.
[0123] The computer-readable storage medium may be the data processing device provided in any of the aforementioned embodiments or the internal storage unit of the computer device, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart memory card (smart media card, SMC), a secure digital (secure digital, SD) card, a flash card (flash card), etc. equipped on the computer device. Further, the computer-readable storage medium may also include both the internal storage unit of the computer device and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store data that has been output or is to be output.
[0124] In one aspect of the present application, a computer program product or a computer program is provided, the computer program product or the computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method provided in one aspect of the embodiments of the present application.
[0125] In one aspect of the present application, another computer program product is provided, which includes a computer program or computer instructions, and when the computer program or computer instructions are executed by a processor, the steps of the data processing method provided in the embodiment of the present application are implemented.
[0126] Finally, it should be noted that the terms in the specification and claims of the present application and the above-mentioned drawings, such as first and second, etc., are used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "including one..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.
[0127] The above disclosure is only the preferred embodiment of the present application, which certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.
Claims
1. A data processing method, It is characterized in that The method comprises: Segment the audio to be processed to obtain at least two audio segments; Using at least two data processing threads to process the at least two audio clips to obtain a rhythm point sequence of each audio clip; Acquire a first rhythm point in a first rhythm point sequence, and acquire a second rhythm point in a second rhythm point sequence; wherein the first rhythm point sequence and the second rhythm point sequence are rhythm point sequences in any combination of adjacent audio segments of the at least two audio segments, and the first rhythm point sequence is arranged before the second rhythm point sequence in time sequence; the first rhythm point is the last rhythm point in the first rhythm point sequence, and the second rhythm point is the first rhythm point in the second rhythm point sequence; Acquire the time interval between the first rhythm point and the second rhythm point, and acquire the beat type of the first rhythm point and the second rhythm point, wherein the beat type includes a strong beat or a weak beat; If the time interval is greater than or equal to the interval threshold, and the beat types of the first rhythm point and the second rhythm point satisfy the beat setting rule, the second rhythm point sequence is fused with the first rhythm point sequence to obtain a fused rhythm point sequence; A target rhythm point sequence of the audio to be processed is determined according to the fused rhythm point sequence.
2. The method according to claim 1, It is characterized in that The method further comprises: If the time interval is less than the interval threshold, or the beat types of the first rhythm point and the second rhythm point do not satisfy the beat setting rule, deleting the second rhythm point from the second rhythm point sequence; Determine a new first rhythm point from the deleted second rhythm point sequence; According to the first rhythm point and the new first rhythm point, the first rhythm point sequence and the deleted second rhythm point sequence are fused to obtain a fused rhythm point sequence.
3. The method according to claim 1, It is characterized in that The step of performing segment processing on the audio to be processed to obtain at least two audio segments includes: Get the audio to be processed; According to the set number of segments and the time length of the audio to be processed, the audio to be processed is segmented evenly to obtain at least two audio segments.
4. The method according to claim 1, It is characterized in that The step of processing the at least two audio clips by using at least two data processing threads to obtain a rhythm point sequence of each audio clip includes: Using at least two data processing threads to perform spectrum conversion processing on the at least two audio segments respectively to obtain a spectrogram of each audio segment; The at least two data processing threads are used to respectively call the beat detection model to perform rhythm point detection processing on the spectrograms of the at least two audio segments to obtain the rhythm point sequences of the respective audio segments.
5. The method according to claim 4, It is characterized in that The step of using the at least two data processing threads to respectively call the beat detection model to perform rhythm point detection processing on the spectrograms of the at least two audio clips to obtain the rhythm point sequences of the respective audio clips includes: For a spectrogram of any audio segment of the at least two audio segments, a target data processing thread of the at least two data processing threads is used to call a beat detection model to process the spectrogram of the any audio segment, determine feature information of the any audio segment, and determine a beat probability value of the audio corresponding to each beat unit time in the any audio segment according to the feature information, and, Determine the beat type of the audio corresponding to each beat unit time according to the probability threshold and the beat probability value, and determine the rhythm point sequence of any audio segment according to the beat type of the audio corresponding to each beat unit time; The beat probability value includes a strong beat probability value and a weak beat probability value, and the beat type includes a strong beat or a weak beat.
6. The method according to claim 5, It is characterized in that The beat detection model includes a first timing network and a second timing network; the spectrogram of any audio segment includes a first spectrogram and a second spectrogram; the beat probability value includes a first beat probability value and a second beat probability value; determining feature information of any audio segment, and determining the beat probability value of the audio corresponding to each beat unit time in any audio segment according to the feature information, includes: Inputting the first spectrogram into the first temporal network for processing to obtain the beat feature of any audio segment; Inputting the second spectrogram into the second time series network for processing to obtain the harmonic features of any audio segment; The first beat probability value of the audio corresponding to each beat unit time in any audio segment is determined according to the beat feature, and the second beat probability value of the audio corresponding to each beat unit time in any audio segment is determined according to the harmonic feature.
7. A data processing device, It is characterized in that include: A segmentation module, used for segmenting the audio to be processed to obtain at least two audio segments; A processing module, configured to process the at least two audio clips using at least two data processing threads to obtain a rhythm point sequence of each audio clip; A fusion module is used to: obtain a first rhythm point in a first rhythm point sequence, and obtain a second rhythm point in a second rhythm point sequence; wherein the first rhythm point sequence and the second rhythm point sequence are rhythm point sequences in any combination of adjacent audio segments in the at least two audio segments, and the first rhythm point sequence is arranged before the second rhythm point sequence in chronological order; the first rhythm point is the last rhythm point in the first rhythm point sequence, and the second rhythm point is the first rhythm point in the second rhythm point sequence; obtain a time interval between the first rhythm point and the second rhythm point, and obtain a beat type of the first rhythm point and the second rhythm point, wherein the beat type includes a strong beat or a weak beat; if the time interval is greater than or equal to an interval threshold, and the beat types of the first rhythm point and the second rhythm point meet a beat setting rule, the second rhythm point sequence is fused with the first rhythm point sequence to obtain a fused rhythm point sequence; and determine a target rhythm point sequence of the audio to be processed according to the fused rhythm point sequence.
8. A computer device, It is characterized in that include: Processor, memory, and network interface; The processor is connected to the memory and the network interface, wherein the network interface is used to provide a network communication function, the memory is used to store program code, and the processor is used to call the program code to execute the data processing method described in any one of claims 1-6.
9. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the data processing method according to any one of claims 1 to 6 is executed.
10. A computer program product, It is characterized in that The computer program product comprises a computer program or computer instructions, and when the computer program or computer instructions are executed by a processor, the steps of the data processing method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Audio abnormality monitoring method, audio abnormality monitoring device, equipment and storage medium
CN110085213A
Method and device for triggering display through audio, computer equipment and storage medium
CN110377212A
Beat detection model training method, beat detection method and beat detection device
CN113223485A