Audio recognition method and apparatus, device, and medium
By segmenting the target audio based on historical audio segment cache information, identifying identical segments and obtaining their content recognition results, and only recognizing changing segments, the high cost and low efficiency of audio recognition in existing technologies are solved, achieving efficient audio recognition and subtitle acquisition.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-11
- Publication Date
- 2026-04-02
AI Technical Summary
Existing audio recognition methods are costly and inefficient, especially in multimedia editing scenarios where changes to the audio require re-identification of all audio content, resulting in high costs and low efficiency.
By segmenting the target audio based on historical audio segment cache information, identical segments are identified, and the content recognition results of these segments are directly obtained using historical audio segment cache information. Only the segments that have changed are further identified.
It effectively reduces audio recognition costs and improves audio recognition efficiency, especially in subtitle addition scenarios, enhancing subtitle acquisition efficiency.
Smart Images

Figure CN2025120701_02042026_PF_FP_ABST
Abstract
Description
Audio recognition method, device, equipment and medium
[0001] Cross-reference to Related Applications
[0002] This application claims priority to Chinese Patent Application No. 202411356731.X, filed on September 26, 2024, the disclosure of which is incorporated herein in its entirety as part of the present application. TECHNICAL FIELD
[0003] The present disclosure relates to an audio recognition method, device, equipment and medium. BACKGROUND
[0004] In some scenarios, it is necessary to identify the audio content and apply the content recognition result of the audio. For example, in order to facilitate the user to clearly know the content of the audio-video, it is generally necessary to add subtitles to the audio-video in the current multimedia editing scenario, that is, to display the content recognition result of the currently played audio in the form of text on the audio-video playing interface. However, the inventors have found through research that the audio recognition method has problems such as high cost and low efficiency, and therefore still needs to be improved. SUMMARY
[0005] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides an audio recognition method, device, equipment and medium.
[0006] The present disclosure provides an audio recognition method, which comprises: performing slicing processing on a target audio based on slicing cache information of a historical audio to obtain a plurality of target audio segments; wherein the target audio is an audio obtained by editing the historical audio; the slicing cache information of the historical audio is information of a plurality of historical audio segments obtained by performing slicing processing on the historical audio based on a content recognition result of the historical audio; comparing the plurality of target audio segments and the plurality of historical audio segments, and determining a first audio segment based on a comparison result; wherein the first audio segment is a segment that is the same in the plurality of target audio segments and the plurality of historical audio segments; obtaining a content recognition result of the first audio segment based on the slicing cache information of the historical audio, and obtaining a content recognition result of the target audio based on the content recognition result of the first audio segment.
[0007] Optionally, the slicing cache information of the historical audio comprises content identification information of the historical audio segments; and the comparing the plurality of target audio segments and the plurality of historical audio segments comprises: obtaining content identification information corresponding to each of the plurality of target audio segments; and comparing the content identification information corresponding to the plurality of target audio segments and the content identification information corresponding to the plurality of historical audio segments.
[0008] Optionally, the obtaining of the content identification information corresponding to each of the plurality of target audio segments comprises: obtaining feature information of the content corresponding to each of the plurality of target audio segments by using a preset feature extraction algorithm, and obtaining the content identification information corresponding to each of the plurality of target audio segments based on the feature information of the content corresponding to each of the plurality of target audio segments.
[0009] Optionally, the obtaining of the content identification result of the target audio based on the content identification result of the first audio segment comprises: determining a candidate segment to be identified based on the segments other than the first audio segment in the plurality of target audio segments; determining a segment to be identified from the candidate segment; obtaining a content identification result corresponding to the segment to be identified, and obtaining the content identification result of the target audio based on the content identification result corresponding to the segment to be identified and the content identification result of the first audio segment.
[0010] Optionally, the determining of the segment to be identified from the candidate segment comprises: taking each of the candidate segments as the segment to be identified; or searching whether the content identification information of the candidate segment is recorded in a preset cache table, and taking the candidate segment whose content identification information is not recorded in the cache table as the segment to be identified; wherein the cache table records the content identification information of the audio segment whose content has been identified in advance.
[0011] Optionally, the obtaining of the content identification result corresponding to the segment to be identified comprises: in a case where the segment to be identified satisfies a preset segment merging condition, determining an associated segment of the segment to be identified from the plurality of target audio segments; merging the segment to be identified and the associated segment to obtain a merged segment; and obtaining a content identification result of the merged segment to obtain the content identification result corresponding to the segment to be identified based on the content identification result of the merged segment.
[0012] Optionally, the segment merging condition comprises: a segment duration being lower than a preset duration threshold, and / or adjacent segments being still the segment to be identified; and the associated segment of the segment to be identified comprises adjacent segments of the segment to be identified.
[0013] Optionally, a duration of the merged segment is greater than a preset duration threshold.
[0014] Optionally, the obtaining of the content identification result of the merged segment comprises: sending the merged segment to a server, and receiving a content identification result returned by the server for the merged segment.
[0015] Optionally, the obtaining the content recognition result of the target audio based on the content recognition result corresponding to the to-be-identified segment and the content recognition result of the first audio segment comprises: obtaining the content recognition result of the target audio based on the content recognition result corresponding to the to-be-identified segment, the content recognition result of the first audio segment, and the starting time and the ending time corresponding to each of the plurality of target audio segments.
[0016] Optionally, the slice cache information of the historical audio comprises the starting time and the ending time of the historical audio slice; and the slicing the target audio based on the slice cache information of the historical audio to obtain the plurality of target audio segments comprises: obtaining audio editing information, wherein the audio editing information is information of editing the historical audio to obtain the target audio; obtaining segmentation information corresponding to the target audio based on the audio editing information and the starting time and the ending time of the historical audio segment; and slicing the target audio based on the segmentation information to obtain the plurality of target audio segments.
[0017] Optionally, the method further comprises: obtaining and storing slice cache information of the target audio based on the content recognition result of the target audio; and the slice cache information of the target audio comprises the starting time, the ending time, the content recognition result, and the content identification information corresponding to each of the plurality of target audio segments.
[0018] The embodiments of the present disclosure further provide an audio recognition device, comprising: an audio slicing module configured to slice a target audio based on slice cache information of a historical audio to obtain a plurality of target audio segments; wherein the target audio is an audio obtained by editing the historical audio; and the slice cache information of the historical audio is information of a plurality of historical audio segments obtained by slicing the historical audio based on a content recognition result of the historical audio; a segment comparison module configured to compare the plurality of target audio segments and the plurality of historical audio segments, and determine a first audio segment based on a comparison result; wherein the first audio segment is a same segment in the plurality of target audio segments and the plurality of historical audio segments; and an audio recognition module configured to obtain a content recognition result of the first audio segment based on the slice cache information of the historical audio, and obtain a content recognition result of the target audio based on the content recognition result of the first audio segment.
[0019] The embodiments of the present disclosure further provide an electronic device, comprising: a processor; a memory configured to store executable instructions of the processor; and the processor configured to read the executable instructions from the memory and execute the instructions to implement the audio recognition method provided by the embodiments of the present disclosure.
[0020] The embodiment of the present disclosure further provides a computer readable storage medium, the storage medium stores a computer program, and the computer program is used for executing the audio recognition method provided by the embodiment of the present disclosure.
[0021] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0022] The accompanying drawings incorporated in and forming a part of the specification illustrate embodiments consistent with the present disclosure and serve to explain the principles of the present disclosure together with the description.
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the drawings needed to be used in the embodiments will be briefly introduced as follows. Obviously, for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.
[0024] Fig. 1 is a flowchart of an audio recognition method provided by an embodiment of the present disclosure;
[0025] Fig. 2 is a slice initialization diagram provided by an embodiment of the present disclosure;
[0026] Fig. 3 is a slice start and end time processing diagram provided by an embodiment of the present disclosure;
[0027] Fig. 4 is a slice start and end time processing diagram provided by an embodiment of the present disclosure;
[0028] Fig. 5 is a slice start and end time processing diagram provided by an embodiment of the present disclosure;
[0029] Fig. 6 is a segment merging processing diagram provided by an embodiment of the present disclosure;
[0030] Fig. 7 is a segment merging processing diagram provided by an embodiment of the present disclosure;
[0031] Fig. 8 is a segment merging processing diagram provided by an embodiment of the present disclosure;
[0032] Fig. 9 is a slice cache information updating diagram provided by an embodiment of the present disclosure;
[0033] Fig. 10 is a subtitle processing flow diagram provided by an embodiment of the present disclosure;
[0034] Fig. 11 is a subtitle processing interaction diagram provided by an embodiment of the present disclosure;
[0035] Fig. 12 is a subtitle processing flow diagram provided by an embodiment of the present disclosure;
[0036] FIG. 13 is a structural schematic diagram of an audio recognition device according to an embodiment of the present disclosure; and
[0037] FIG. 14 is a structural schematic diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0038] In order to more clearly understand the above-mentioned purposes, features and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.
[0039] In the following description, a large number of specific details are set forth in order to facilitate a thorough understanding of the present disclosure, but the present disclosure can also be implemented in other different manners from those described herein; obviously, the embodiments described in the specification are only a part of the embodiments of the present disclosure, rather than all the embodiments.
[0040] The inventors have found that in a multimedia editing scenario, once the audio changes, it is necessary to re-identify all the audio content, which has a high recognition cost and low efficiency, and the user waiting time is also relatively long. In order to improve this problem, the present disclosure provides an audio recognition method, device, equipment and medium, which will be described in detail below.
[0041] FIG. 1 is a flowchart of an audio recognition method according to an embodiment of the present disclosure, which can be executed by a subtitle processing device. The device can be implemented by software and / or hardware, and can be integrated in an electronic device. As shown in FIG. 1, the method mainly includes the following steps S102-S108:
[0042] In step S102, the target audio is sliced based on the slice cache information of the historical audio to obtain a plurality of target audio segments; wherein the target audio is an audio obtained by editing the historical audio; and the slice cache information of the historical audio is information of a plurality of historical audio segments obtained by slicing the historical audio based on a content recognition result of the historical audio.
[0043] In some embodiments, the target audio is an audio obtained based on a current target draft, and the historical audio is an audio obtained by editing the target draft once before. The target draft contains multimedia materials and editing information of the multimedia materials. The target draft can be regarded as an engineering file of multimedia editing. In actual application, multimedia content such as audio and video can be exported based on the target draft, and the draft can be saved so as to further edit the corresponding multimedia content based on the draft when needed. In actual application, a subtitle adding operation can be further performed based on the content recognition result of the audio, so as to directly export a video configured with subtitles based on the target draft.
[0044] Exemplarily, the content recognition result of the historical audio can be used for sentence segmentation processing, so as to obtain slice cache information of the historical audio. The slice cache information of the historical audio can include the start time, the end time, the content recognition result and the content identification information of each historical audio segment. The content recognition result, i.e., the speech recognition result of the audio segment, can be represented in the form of text; the content identification information can be expressed by the content features of the audio segment, so as to uniquely identify the content of the audio segment. For example, the content identification information can be an MD5 value calculated based on the content of the audio segment. The MD5 value is a 128-bit (16-byte) hash value calculated by the MD5 (Message-Digest Algorithm 5) information digest algorithm, which can uniquely identify a file or a piece of data. Even if the original data has a slight change, the generated MD5 value will be completely different. In actual application, the slice cache information can be stored in the draft or uploaded to the cloud for multi-terminal synchronous caching. The draft protocol can be added in the draft to indicate the parsing and processing mode of the slice cache information. The slice cache information in the draft can be updated as the draft changes.
[0045] It is worth noting that, instead of directly dividing the target audio according to the preset time length, the content recognition result of the historical audio is used to determine the division mode of the target audio. Specifically, the content recognition result of the historical audio can be segmented, and the division mode of the target audio can be determined based on the segmentation result. The segmentation result can be obtained by segmenting based on punctuation, pause and the like. The division information obtained by the above method can include the start time and the end time of each to-be-divided segment in the target audio. Different segments are divided in units of sentences, and the above division mode is more convenient for parsing and processing.
[0046] In step S104, the plurality of target audio segments and the plurality of historical audio segments are compared, and a first audio segment is determined based on the comparison result. The first audio segment is the same segment in the plurality of target audio segments and the plurality of historical audio segments. That is, the first audio segment appears in both the plurality of target audio segments and the plurality of historical audio segments.
[0047] By comparing the plurality of target audio segments and the plurality of historical audio segments, the same segment in the plurality of target audio segments and the plurality of historical audio segments can be quickly found, i.e., the first audio segment that does not need to be recognized again can be found from the plurality of target audio segments, so as to effectively save the recognition cost of the first audio segment in the target audio.
[0048] In step S106, a content recognition result of the first audio segment is obtained based on the slice cache information of the historical audio, and a content recognition result of the target audio is obtained based on the content recognition result of the first audio segment.
[0049] The embodiments of the present disclosure can determine the to-be-recognized segment based on the segments other than the first audio segment in the plurality of target audio segments. In the related art, even if the audio is sliced, content recognition is performed on all audio segments obtained by slicing in parallel. In the embodiments of the present disclosure, the slice cache information of the historical audio can be referred to, so that it is not necessary to recognize all target audio segments, but the to-be-recognized segment is determined therefrom, so that subsequent content recognition can be performed only on the to-be-recognized segment. Then, the content recognition result of the target audio can be directly obtained based on the content recognition result of the first audio segment and the content recognition result of the to-be-recognized segment obtained.
[0050] In actual application, the content recognition result corresponding to the to-be-recognized segment can be obtained by using the ASR (Automatic Speech Recognition) technology. Specifically, the content recognition result can be obtained by local recognition or uploading to a server for recognition, so that the powerful processing capability of the server is used to quickly obtain the content recognition result. The specific manner of obtaining the content recognition result of the to-be-recognized segment is not limited herein. Based on the content recognition result of the historical audio, the content recognition result of the segment other than the to-be-recognized segment in the plurality of target audio segments can be obtained, and then the content recognition result corresponding to the to-be-recognized segment is combined, so that the content recognition result corresponding to the target audio can be conveniently and quickly obtained.
[0051] The above technical solution provided by the embodiments of the present disclosure can slice the target audio based on the slice cache information of the historical audio, and find the same segment (i.e., the first audio segment) in the plurality of target audio segments and the plurality of historical audio segments by comparison, so that it is not necessary to recognize the first audio segment again, but the content recognition result of the first audio segment can be directly obtained based on the slice cache information of the historical audio, which is helpful to efficiently obtain the content recognition result of the target audio. The above manner does not need to perform content recognition on the entire target audio, so that the audio recognition cost can be effectively reduced and the audio recognition efficiency can be improved. The above manner can be effectively applied to the application scenario of the audio recognition result, such as the scenario of adding subtitles, so that the subtitle obtaining efficiency is improved and the subtitle obtaining cost is reduced.
[0052] In some embodiments, the slice cache information of the historical audio includes the start time and the end time of the historical audio slice; and the step S102, i.e., the step of slicing the target audio based on the slice cache information of the historical audio to obtain the plurality of target audio segments, can be performed with reference to the following steps 1 to 3.
[0053] Step 1, obtaining audio editing information; wherein the audio editing information is information of editing a historical audio to obtain a target audio. In actual application, if the historical audio and the target audio are obtained based on a target draft, the audio editing information can be obtained based on change information of the target draft. The change information of the target draft is the difference between the target draft before and after editing, for example, assuming that the target audio is obtained based on the current target draft, and the historical audio is obtained based on the previous editing of the target draft, the change information of the target draft represents the difference between the content of the current target draft and the content of the previous editing of the target draft. In actual application, the target draft can record the editing information of the user each time, so the change information can be obtained directly based on the editing information, and the audio editing information between the historical audio and the target audio can be obtained accurately and reliably.
[0054] Step 2, obtaining segmentation information corresponding to the target audio based on the audio editing information and the start time and end time of the historical audio slice. In actual application, the audio editing information can be used to know what editing operation is performed by the user based on the historical audio, such as adding audio content, modifying audio content, deleting audio content, etc. Then, based on the editing operation, the start time and end time of the historical audio slice can be adjusted, so as to obtain the start time and end time of the target audio slice corresponding to the target audio, that is, to obtain the segmentation information corresponding to the target audio. The above method can make the target audio slice and the historical audio slice have an accurate corresponding relationship. It can be understood that in the scenario of adding / deleting audio content, the content of the target audio will be dislocated relative to the content of the historical audio. If the target audio is directly segmented according to the start time and end time of the historical audio slice, the part of the target audio after adding / deleting the audio content cannot be accurately corresponding to the historical audio segment. Therefore, the embodiment of the present disclosure does not directly segment the target audio slice according to the start time and end time of the historical audio slice, but combines the audio editing information to accurately determine the segmentation method of the target audio, so that the target audio segment obtained by segmentation can have an accurate corresponding relationship with the historical audio segment.
[0055] Step 3, performing slice processing on the target audio based on the segmentation information to obtain a plurality of target audio segments. The target audio can be more objectively and reliably sliced based on the above segmentation information. Compared with the way of directly segmenting with a preset time length in the related art, the above segmentation method is more convenient for subsequent identification and processing.
[0056] To facilitate the understanding of the above, the following first introduces the related content of slice initialization. Specifically, when the audio content corresponding to the first editing of the target draft is recognized, that is, after obtaining the content recognition result for the first time, the slice granularity can be determined based on the punctuation granularity of the content recognition result, so as to perform slice processing on the audio corresponding to the first editing of the target draft. Specifically, refer to a slice initialization diagram shown in FIG. 2, which illustrates a plurality of slices (i.e., audio segments) and the key value corresponding to each slice. The key value corresponding to each slice is the content identification information corresponding to each audio segment as described above. The MD5 value of each slice can be used as the key value of the slice to uniquely identify the content of the slice. In addition, FIG. 2 also identifies the recognition text result (i.e., the content recognition result in the form of text) corresponding to the slice. It should be noted that each slice corresponds to a recognition text result. FIG. 2 only symbolically illustrates the recognition text result of part of the slice. In actual application, the key value and the recognition text result corresponding to each slice can be stored in the cache table in the form of key / value, that is, the content identification information and the content recognition result corresponding to the audio segment can be stored in a specified location in a one-to-one association.
[0057] On the basis of slice initialization, the following introduces the way of slice time maintenance. Specifically, when the draft changes, the start and end times of each slice need to be updated, mainly in three scenarios: adding audio content, modifying audio content, and deleting audio content. The scenario of adding audio content can be seen in a slice start and end time processing diagram shown in FIG. 3. The draft indicates that the user inserts content in slice 5, slice 5 becomes longer, and therefore the start and end times of slice 5 and the start and end times of the subsequent slices need to be offset backward. The scenario of modifying audio content can be seen in a slice start and end time processing diagram shown in FIG. 4. The draft indicates that the user only modifies the audio content in a certain time range, and the start and end times of the slices do not change, so the start and end times of the slices are not processed, that is, the start and end times of the slices do not change. The scenario of deleting audio content can be seen in a slice start and end time processing diagram shown in FIG. 5. Part of the content of slice 4, slice 5, and part of the content of slice 6 are deleted, and therefore the start and end times of slice 4 and the subsequent slices need to be updated.
[0058] Through the above-mentioned manner, the audio editing information can be obtained based on the draft change, and the start and end times of the target audio segment can be obtained in combination with the start and end times of the historical audio segment, that is, the accurate segmentation information can be obtained.
[0059] In some embodiments, the slice cache information of the historical audio includes content identification information of the historical audio segment; the step of comparing the plurality of target audio segments and the plurality of historical audio segments in the foregoing step S104 can be performed with reference to the following steps A to B:
[0060] Step A, obtaining the content identification information corresponding to each of the plurality of target audio segments. In some specific embodiments, a preset feature extraction algorithm can be used to obtain the feature information of the content corresponding to each of the plurality of target audio segments, and the content identification information corresponding to each of the plurality of target audio segments is obtained based on the feature information of the content corresponding to each of the plurality of target audio segments. The embodiments of the present disclosure do not limit the feature extraction algorithm. For example, the feature extraction algorithm includes MD5 algorithm. The feature information (such as MD5) obtained by the above-mentioned method can uniquely identify the segment content. The feature information of the target audio segment can be directly used as the content identification information of the target audio segment, so as to accurately and conveniently determine whether the content of the audio segment changes, and the maintenance cost is low. The final generated audio content can be mainly concerned, and the draft can be decoupled, and there is no need to concern about the business change.
[0061] Step B, comparing the content identification information corresponding to the plurality of target audio segments and the content identification information corresponding to the plurality of historical audio segments. Specifically, for each target audio segment, the content identification information of the target audio segment and the historical audio segment corresponding to the target audio segment is inconsistent, and then the first audio segment is screened out according to the comparison result, that is, the audio segment with consistent content identification information in the content identification information corresponding to the plurality of target audio segments and the content identification information corresponding to the plurality of historical audio segments is taken as the first audio segment. In this way, the same segment in the plurality of target audio segments and the plurality of historical audio segments can be efficiently and accurately found.
[0062] In some examples, the step of obtaining the content identification result of the target audio based on the content identification result of the first audio segment in step S106 can be performed by referring to the following steps (1) to (3):
[0063] Step (1), determining the candidate segment to be identified based on the segments other than the first audio segment in the plurality of target audio segments; such as, taking all the segments other than the first audio segment in the plurality of target audio segments as candidate segments, wherein the content identification information of the candidate segment is inconsistent with the content identification information of the historical audio segment corresponding to the candidate segment, such as, taking the target audio segment with MD5 value different from the MD5 value of the corresponding historical audio segment as the candidate segment. It can be understood that if the content identification information changes, it means that the content of the target audio segment has changed compared with the historical audio segment, and then it can be taken as the candidate segment to be identified.
[0064] Step (2), determining the to-be-identified segment from the candidate segments. In some specific embodiments, all the candidate segments can be taken as the to-be-identified segments; in some other specific embodiments, it can be checked from a preset cache table whether the content identification information of the candidate segments is recorded, and the candidate segments whose content identification information is not recorded in the cache table are taken as the to-be-identified segments; wherein, the cache table records the content identification information of the audio segments whose contents have been identified in advance; further, the content identification information in the cache is stored in association with the content identification result. As for the candidate segments whose content identification information is recorded in the cache table, the content identification result corresponding to the content identification information recorded in the cache table can be directly taken as the content identification result of the segment, and the segment does not need to be re-identified.
[0065] In actual application, the cache table can be a local cache or a cloud cache. For example, for the cache table of the local cache, the content identification information of the segments corresponding to the audio edited by all the users can be recorded, although the content of the target audio segment is changed compared with the historical audio segment edited in the previous draft, the user can have used the content of the target audio segment in the previous draft editing and recorded in the cache, and thus the content identification result of the target audio segment can still be recorded. For example, the audio edited by the user for the first time contains segment A, the audio edited by the user for the second time contains segment B which is modified from segment A, and the audio edited by the user for the third time contains segment A which is modified from segment B, although the content of segment A in the audio edited by the user for the third time is inconsistent with the content of segment B in the audio edited by the user for the second time, since segment A appears in the audio edited by the user for the first time, the content identification information of segment A is still recorded in the cache table, and thus the content identification result stored in association with the content identification information of segment A can be directly found, and segment A does not need to be identified again. In addition, in the embodiments of the present disclosure, the cache table of the cloud can also be used, the cache table of the cloud records the content identification information and the content identification result of all the identified audio segments, for example, the cache table of the cloud records the content identification information and the content identification result of all the identified public materials, and if the current to-be-identified audio contains an audio segment with the same content identification information as the identified public material, for example, the same segment is contained in the respective audios, the audio segment can be directly obtained from the cache table of the cloud, which is very convenient and efficient.
[0066] Step (3), obtaining the content identification result corresponding to the to-be-identified segment, and obtaining the content identification result of the target audio based on the content identification result corresponding to the to-be-identified segment and the content identification result of the first audio segment.
[0067] For example, the content recognition result corresponding to the to-be-recognized segment can be obtained by using the foregoing ASR technology, can be recognized locally, or can be uploaded to a server for recognition, and the recognition result is not limited herein. Then, based on the content recognition result corresponding to the to-be-recognized segment and the content recognition result of the first audio segment, the content recognition result of the target audio can be obtained by combination. In this way, it is not necessary to recognize all the target audio segments, but to-be-recognized segments that have not been recognized before are searched from the target audio segments for recognition, so that the recognition cost is effectively saved, the recognition efficiency is improved, and the user's waiting time for recognition can be reduced.
[0068] In some embodiments, the step of obtaining the content recognition result corresponding to the to-be-recognized segment in step (3) can be performed according to the following steps a to c.
[0069] In step a, when the to-be-recognized segment meets a preset segment merging condition, the associated segment of the to-be-recognized segment is determined from the plurality of target audio segments.
[0070] In some embodiments, the segment merging condition includes that the segment duration is less than a preset duration threshold and / or the adjacent segment is also a to-be-recognized segment. The adjacent segment of the target audio segment can be a preceding adjacent segment and / or a following adjacent segment. The preset duration threshold can be set according to requirements, such as 15 s, and the threshold is not limited herein. For example, as long as the preceding adjacent segment or the following adjacent segment of the target audio segment is also a to-be-recognized segment, the target audio segment meets the segment merging condition. In addition, as long as the duration of the target audio segment is less than the preset duration threshold, the target audio segment meets the segment merging condition.
[0071] The associated segment of the to-be-recognized segment includes the adjacent segment of the to-be-recognized segment, and the adjacent segment includes the preceding adjacent segment and / or the following adjacent segment. For each to-be-recognized segment, if the duration thereof is short, the cutting cost is high and the recognition accuracy is low. By merging the to-be-recognized segment with the adjacent segment, the recognition accuracy of the merged segment with a long duration can be improved, and the cutting cost can be effectively reduced. In addition, at least two continuous to-be-recognized segments can also be merged, so that the audio cutting operation is reduced and the efficiency is improved.
[0072] In step b, the to-be-recognized segment and the associated segment are merged to obtain a merged segment. For example, the duration of the merged segment is greater than the preset duration threshold, so as to ensure the recognition accuracy of the merged segment.
[0073] It should be noted that the associated segment of the to-be-identified segment can not be limited to one segment, and the associated segment can be one or more segments adjacent to the to-be-identified segment in front or behind, such as merging processing of continuous N to-be-identified segments, or merging processing of continuous M adjacent segments in front and / or behind the to-be-identified segment, until the time length of the merged segment is greater than the preset time length threshold, and the values of N and M can be dynamically adjusted according to actual conditions.
[0074] In step c, the content recognition result of the merged segment is obtained, so as to obtain the content recognition result corresponding to the to-be-identified segment based on the content recognition result of the merged segment. Since the merged segment contains the to-be-identified segment, the content recognition result corresponding to the to-be-identified segment can be directly extracted based on the content recognition result of the merged segment.
[0075] In actual application, in order to improve the identification processing efficiency, the server with strong data processing capability can be used for identification processing, so that when the content recognition result of the merged segment is obtained, the merged segment can be sent to the server, and the content recognition result returned by the server for the merged segment can be received.
[0076] In order to facilitate understanding of the merging processing mode, the following will be described in three scenes of adding audio content, modifying audio content, and deleting audio content. Specifically, based on FIGS. 3-5, the segment merging processing diagrams shown in FIGS. 6-8 will be described. FIGS. 3 and 6 correspond to the scene of adding audio content. As shown in FIG. 3, the inserted content of slice 5 is shown in FIG. 6 as key5', and slice 5 needs to be added to the to-be-identified audio segment queue as a to-be-identified segment. FIGS. 4 and 7 correspond to the scene of modifying audio content. As shown in FIG. 4, the audio content of slice 5 and slice 6 is modified in a certain time range. In FIG. 7, the characteristic information of slice 5 after modification is key5', and the characteristic information of slice 6 after modification is key6', both of which are to-be-identified segments. Therefore, slice 5 and slice 6 can be merged to obtain a merged segment (i.e., slice 5'), and the newly obtained slice 5' is added to the to-be-identified audio segment queue. FIGS. 5 and 8 correspond to the scene of deleting audio content. As shown in FIG. 5, part of the content of slice 4, all the content of slice 5, and part of the content of slice 6 are deleted. The content identification information of the remaining slice 4 becomes key4', and the content identification information of the remaining slice 6 becomes key6'. Since the remaining slice 4 is less than the preset time length threshold (assuming 15 seconds), the remaining slice 4 and the remaining slice 6 can be merged to obtain slice 4', and then the new slice 4' obtained by merging is added to the to-be-identified audio segment queue. Through the above method, the identification efficiency and accuracy can be improved, and the cutting cost can be reduced.
[0077] After obtaining the content recognition result corresponding to the to-be-identified segment, the content recognition result of the target audio can be obtained based on the content recognition result corresponding to the to-be-identified segment and the content recognition result of the first audio segment. Specifically, the content recognition result of the target audio can be obtained based on the content recognition result corresponding to the to-be-identified segment, the content recognition result of the first audio segment, and the starting time and the ending time corresponding to each of the plurality of target audio segments. In actual application, the content recognition result of the first audio segment can also be directly obtained based on the content identification information of the first audio segment, by searching the cache table for the content recognition result associated with the content identification information, so as to quickly obtain the content recognition result of the first audio segment. The content recognition result corresponding to the target audio can be obtained by sequentially splicing the starting time and the ending time corresponding to each of the target audio segments.
[0078] Further, the method provided by the embodiments of the present disclosure further includes obtaining and storing the slice cache information of the target audio based on the content recognition result of the target audio. Specifically, the content recognition result of the target audio can be processed by sentence segmentation, and the slice cache information of the target audio can be obtained and stored based on the sentence segmentation result. The slice cache information of the target audio includes the starting time, the ending time, the content recognition result and the content identification information corresponding to each of the plurality of target audio segments. The slice cache information of the target audio can be used to update the slice cache information of the historical audio recorded locally or on the server. That is, when the user edits the audio next time, the slice cache information of the target audio obtained and stored by the foregoing method is used as the slice cache information of the historical audio to be referred to for the audio obtained by the next editing, and the audio obtained by the next editing is regarded as a new target audio to be processed.
[0079] To facilitate understanding of the updating method of the slice cache information, taking the content recognition result of the audio as a subtitle as an example, the slice cache information updating diagram shown in FIG. 9 is referred to, in which it is shown that the slice 3 and the slice 5 are to-be-identified slices and are re-identified to obtain the corresponding recognition result. The slice 5 is a merged segment, and the recognition result thereof indicates that the slice 5 substantially corresponds to three sentences. The slice subtitle information (i.e., the content recognition result of the slice) can be updated based on the new recognition result of the slice 3 and the new recognition result of the slice 5. Specifically, the slice 5 can be updated by being divided into three slices according to the sentence segmentation result, i.e., the slice 5 to the slice 7, and each slice independently corresponds to the corresponding starting time, the ending time, the content recognition result and the content identification information. That is, based on the subtitle sentence segmentation result corresponding to the target audio, the most accurate slice cache information can be obtained and the slice cache list (i.e., the foregoing cache table) can be updated.
[0080] To facilitate understanding of the audio recognition method provided by the embodiments of the present disclosure, an application scenario example of taking the content recognition result of audio as subtitles is provided. Referring to a subtitle processing flow diagram shown in FIG. 10, multimedia information such as a video, recording 1, recording 2, subtitles, and the like is recorded in a draft. A complete audio can be obtained by synthesizing audio. Subtitle sentence information can be obtained by the subtitle sentences of the historical audio cached in the draft. Then, the audio to be recognized can be determined based on the audio modified by the user. In FIG. 10, slices 1, 2, and 4 are taken as examples of the to-be-recognized segments. The to-be-recognized segments are extracted and uploaded to an ASR recognition server for speech recognition. The slice 1 and the slice 2 can be merged and cropped. The content recognition result of the slice 1, the slice 2, and the slice 4 can be obtained by the server, and uploaded to the draft recognition cache data table to update the cache. The cache can be read and the subtitles in the draft can be updated subsequently. In the above manner, the content recognition of the entire target audio is not required. Instead, the content recognition of the to-be-recognized segments is sufficient. The audio recognition cost, that is, the subtitle acquisition cost, can be effectively reduced, and the subtitle acquisition efficiency can be improved.
[0081] Based on FIG. 10, referring to a subtitle processing interaction diagram shown in FIG. 11, a subtitle processing flow is described by taking the interaction between the user end and the server as an example. The subtitle processing flow is also an audio recognition flow. However, a step of taking the content recognition result of audio as subtitles is added. In FIG. 11, three processing modes of first recognition, draft change, and no draft change are shown. The specific steps S1-S14 are described as follows.
[0082] S1, the user end synthesizes an audio file based on the draft.
[0083] S2, the user end uploads the audio file to the server.
[0084] S3, the server obtains the content recognition result of the audio file by using an ASR technology.
[0085] S4, the server returns the content recognition result of the audio file in response to the query request of the user terminal.
[0086] S5, the user end creates slice cache information according to the text sentence timestamp. The slice cache information includes slice start and end time, slice key value, and slice recognition result.
[0087] S6, in the case where the draft is changed, the user end adjusts the slice start and end time and recalculates the slice key.
[0088] S7, the user end adds the slice with the changed key to the to-be-recognized slice queue.
[0089] S8, the user terminal extracts the to-be-identified slice and performs a slice key repetition detection operation. If there is no cache repetition, step S9 is performed, and if there is cache repetition, step S12 is performed. The slice key repetition detection operation corresponds to the related content of step (2) described above, that is, a secondary determination is performed using the cache table, and the slice with a key value not recorded in the cache table is regarded as a slice that needs to be uploaded to the server.
[0090] S9, the user terminal uploads the current to-be-identified slice.
[0091] S10, the server obtains the content recognition result of the to-be-identified slice using an ASR technology.
[0092] S11, the server returns the content recognition result of the to-be-identified slice in response to the query request of the user terminal.
[0093] S12, the user terminal updates the slice cache information. Then, the next to-be-identified slice can be processed by returning to step S8. It should be noted that FIG. 11 is only used to illustrate the processing of to-be-identified slices in series, and in actual applications, multiple to-be-identified slices can be processed in parallel, that is, multiple slices can be polled and recognized at the same time, so as to update the subtitle cache data.
[0094] S13, all slice recognition is completed, and the cache data is read to update the subtitle.
[0095] S14, in the case where the draft has not been changed, the user terminal uses the local cache information.
[0096] It should be noted that FIG. 11 is only used to illustrate the interaction mode of the user terminal and the server, and some steps are omitted, such as the query request initiated by the user terminal to the server and the association step of the slice and the identification task identifier. The specific implementation mode of the above steps can be referred to the related content described above, and will not be described here.
[0097] For ease of understanding, the content of FIG. 11 can be further summarized, and the subtitle processing flow diagram shown in FIG. 12 can be referred to. In the case where the draft initiates identification, the subtitle sentence slice can be read, and the slice list for re-identification can be exported. Then, the to-be-identified slice is uploaded, the identification task is submitted, the identification result is queried, and the slice cache is updated. Then, the to-be-identified slice uploading operation and the like can be performed, and finally, the sentence slice can be updated based on the identification result, so as to read the subtitle sentence slice again. Similarly, FIG. 12 is only used to illustrate the processing of to-be-identified slices in series, and in actual applications, multiple to-be-identified slices can be processed in parallel.
[0098] In summary, in the process of obtaining the audio subtitle corresponding to the draft, the content recognition of the entire audio is not required, the audio recognition cost can be effectively reduced, that is, the subtitle obtaining cost is reduced, and the subtitle obtaining efficiency is improved.
[0099] Corresponding to the foregoing audio recognition method, the embodiment of the disclosure further provides an audio recognition device. FIG. 13 is a structural schematic diagram of an audio recognition device provided by an embodiment of the disclosure. The device can be implemented by software and / or hardware, and can be integrated in an electronic device. As shown in FIG. 13, the audio recognition device comprises:
[0100] The audio slicing module 1302 is configured to perform slicing processing on the target audio based on the slicing cache information of the historical audio to obtain a plurality of target audio segments; wherein the target audio is an audio obtained by editing the historical audio; and the slicing cache information of the historical audio is information of a plurality of historical audio segments obtained by performing slicing processing on the historical audio based on a content recognition result of the historical audio.
[0101] The segment comparison module 1304 is configured to compare the plurality of target audio segments and the plurality of historical audio segments, and determine a first audio segment based on a comparison result; wherein the first audio segment is a segment that is the same in the plurality of target audio segments and the plurality of historical audio segments.
[0102] The audio recognition module 1306 is configured to obtain a content recognition result of the first audio segment based on the slicing cache information of the historical audio, and obtain a content recognition result of the target audio based on the content recognition result of the first audio segment.
[0103] The above device does not require content recognition of the entire target audio, and can effectively reduce the audio recognition cost and improve the audio recognition efficiency.
[0104] In some embodiments, the slicing cache information of the historical audio comprises content identification information of the historical audio segments; and the segment comparison module 1304 is specifically configured to: obtain content identification information corresponding to each of the plurality of target audio segments; and compare the content identification information corresponding to the plurality of target audio segments and the content identification information corresponding to the plurality of historical audio segments.
[0105] In some embodiments, the segment comparison module 1304 is specifically configured to: obtain feature information of the content corresponding to each of the plurality of target audio segments by using a preset feature extraction algorithm, and obtain the content identification information corresponding to each of the plurality of target audio segments based on the feature information of the content corresponding to each of the plurality of target audio segments.
[0106] In some embodiments, the audio recognition module 1306 is specifically configured to: determine candidate segments to be recognized based on the segments other than the first audio segment in the plurality of target audio segments; determine a segment to be recognized from the candidate segments; obtain a content recognition result corresponding to the segment to be recognized; and obtain the content recognition result of the target audio based on the content recognition result corresponding to the segment to be recognized and the content recognition result of the first audio segment.
[0107] In some embodiments, the audio recognition module 1306 is specifically configured to: take each of the candidate segments as the segment to be recognized; or, search a preset cache table for whether content identification information of the candidate segments is recorded, and take the candidate segment whose content identification information is not recorded in the cache table as the segment to be recognized; and wherein the cache table records content identification information of audio segments that have been pre-recognized.
[0108] In some embodiments, the audio recognition module 1306 is specifically configured to: in a case where the segment to be recognized satisfies a preset segment merging condition, determine an associated segment of the segment to be recognized from the plurality of target audio segments; perform merging processing on the segment to be recognized and the associated segment to obtain a merged segment; obtain a content recognition result of the merged segment; and obtain the content recognition result corresponding to the segment to be recognized based on the content recognition result of the merged segment.
[0109] In some embodiments, the segment merging condition includes: a segment duration being lower than a preset duration threshold, and / or, adjacent segments being segments to be recognized; and the associated segment of the segment to be recognized includes an adjacent segment of the segment to be recognized.
[0110] In some embodiments, a duration of the merged segment is greater than a preset duration threshold.
[0111] In some embodiments, the audio recognition module 1306 is specifically configured to: send the merged segment to a server, and receive a content recognition result returned by the server for the merged segment.
[0112] In some embodiments, the audio recognition module 1306 is specifically configured to: obtain the content recognition result of the target audio based on the content recognition result corresponding to the segment to be recognized, the content recognition result of the first audio segment, and respective start times and end times of the plurality of target audio segments.
[0113] In some embodiments, the slice cache information of the historical audio includes a start time and an end time of the historical audio slice; the audio slicing module 1302 is specifically configured to: obtain audio editing information; wherein the audio editing information is information of editing the historical audio to obtain the target audio; obtain the segmentation information corresponding to the target audio based on the audio editing information and the start time and the end time of the historical audio segment; and perform slice processing on the target audio based on the segmentation information to obtain the plurality of target audio segments.
[0114] In some embodiments, the apparatus further includes a cache storage module configured to obtain and store slice cache information of the target audio based on the content recognition result of the target audio; the slice cache information of the target audio includes a start time, an end time, a content recognition result, and content identification information corresponding to each of the plurality of target audio segments.
[0115] The audio recognition apparatus provided by the embodiments of the present disclosure can perform the audio recognition method provided by any of the embodiments of the present disclosure, and has the corresponding function modules and beneficial effects of performing the method.
[0116] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the apparatus embodiments described above can refer to the corresponding process in the method embodiments, which will not be described here.
[0117] The embodiments of the present disclosure provide an electronic device, which includes a storage device having a computer program stored thereon, and a processing device configured to execute the computer program in the storage device to implement the steps of any method of the present disclosure.
[0118] Reference is made to FIG. 14, which shows a structural schematic diagram of an electronic device 1400 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure can include, but is not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablets), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and the like, and fixed terminals such as digital TVs, desktop computers, and the like. The electronic device shown in FIG. 14 is only an example, and should not bring any limitation to the functions and use range of the embodiments of the present disclosure.
[0119] As shown in FIG. 14, the electronic device 1400 can include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 1401 that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 1402 or loaded into a random access memory (RAM) 1403 from a storage device 1408. Various programs and data required for the operation of the electronic device 1400 are also stored in the RAM 1403. The processing device 1401, the ROM 1402, and the RAM 1403 are connected to each other through a bus 1404. An input / output (I / O) interface 1405 is also connected to the bus 1404.
[0120] In general, the following devices can be connected to the I / O interface 1405: input devices 1406 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 1407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 1408 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 1409. The communication devices 1409 can allow the electronic device 1400 to communicate wirelessly or wired with other devices to exchange data. While FIG. 14 shows the electronic device 1400 with various devices, it should be understood that all of the shown devices are not required to be implemented or possessed. More or fewer devices can be alternatively implemented or possessed.
[0121] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 1409, or installed from the storage devices 1408, or installed from the ROM 1402. When the computer program is executed by the processing device 1401, the above-described functions defined in the methods of embodiments of the present disclosure are performed.
[0122] In addition to the method and device described above, the embodiments of the present disclosure can also be a computer program product, which includes computer program instructions that make the processor execute the image processing method provided by the embodiments of the present disclosure when the processor runs. The computer program product can be written in any combination of one or more programming languages to execute the program code of the embodiments of the present disclosure, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as "C" language or similar programming languages. The program code can be executed entirely on a user computing device, partially on a user device, as an independent software package, partially on a user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0123] In addition, the embodiments of the present disclosure can also be a computer readable storage medium, which stores computer program instructions, and the computer program instructions make the processor execute the audio recognition method provided by the embodiments of the present disclosure when the processor runs.
[0124] The computer readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples (non-exhaustive list) of readable storage medium include: electrical connection with one or more conductive wires, portable disk, hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination of the above.
[0125] The embodiments of the present disclosure also provide a computer program product, which includes computer programs / instructions that are executed by a processor to implement the audio recognition method in the embodiments of the present disclosure.
[0126] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type, use range, use scenario, etc. of the personal information involved in the present disclosure should be informed to the user and the authorization of the user should be obtained according to relevant laws and regulations.
[0127] For example, in response to receiving an active request of a user, a prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed by the user will need to acquire and use personal information of the user. Thus, the user can autonomously select whether to provide the personal information to the software or hardware such as an electronic device, an application program, a server or a storage medium performing the operation of the technical solution of the present disclosure according to the prompt information.
[0128] As an optional but non-limiting implementation, in response to receiving an active request of a user, the manner of sending a prompt information to the user may, for example, be a pop-up window manner, and the prompt information may be presented in the pop-up window in the form of text. In addition, the pop-up window may also carry a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0129] It can be understood that the above notification and acquisition of user authorization process is only illustrative and does not limit the implementation of the present disclosure, and other manners meeting the relevant laws and regulations can also be applied to the implementation of the present disclosure.
[0130] It should be noted that in this document, relational terms such as "first" and "second", and the like can be used solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any actual such relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without more limitations, an element preceded by "comprises a" does not, without more constraints, foreclose the existence of additional identical elements in the process, method, article, or apparatus that comprises the recited element.
[0131] The above description is merely one specific implementation of the present disclosure, which enables those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to these embodiments described herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An audio recognition method, comprising: slicing a target audio based on slice cache information of a historical audio to obtain a plurality of target audio segments; wherein the target audio is an audio obtained by editing the historical audio; the slice cache information of the historical audio is information of a plurality of historical audio segments obtained by slicing the historical audio based on a content recognition result of the historical audio; comparing the plurality of target audio segments and the plurality of historical audio segments, and determining a first audio segment based on a comparison result; wherein the first audio segment is a same segment in the plurality of target audio segments and the plurality of historical audio segments; obtaining a content recognition result of the first audio segment based on the slice cache information of the historical audio, and obtaining a content recognition result of the target audio based on the content recognition result of the first audio segment.
2. The method of claim 1, wherein, The slice cache information of the historical audio comprises content identification information of historical audio segments; the comparing the plurality of target audio segments and the plurality of historical audio segments comprises: obtaining content identification information corresponding to each of the plurality of target audio segments; comparing the content identification information corresponding to the plurality of target audio segments and the content identification information corresponding to the plurality of historical audio segments.
3. The method of claim 2, wherein, The obtaining the content identification information corresponding to each of the plurality of target audio segments comprises: obtaining feature information of content corresponding to each of the plurality of target audio segments by using a preset feature extraction algorithm, and obtaining the content identification information corresponding to each of the plurality of target audio segments based on the feature information of the content corresponding to each of the plurality of target audio segments.
4. The method according to any one of claims 1 to 3, wherein, The obtaining the content recognition result of the target audio based on the content recognition result of the first audio segment comprises: determining a candidate segment to be recognized based on segments other than the first audio segment in the plurality of target audio segments; determining a segment to be recognized from the candidate segment; obtaining a content recognition result corresponding to the segment to be recognized, and obtaining the content recognition result of the target audio based on the content recognition result corresponding to the segment to be recognized and the content recognition result of the first audio segment.
5. The method of claim 4, wherein, The determining the segment to be recognized from the candidate segment comprises: regarding each of the candidate segments as the segment to be recognized; or searching whether the content identification information of the candidate segment is recorded in a preset cache table, and regarding the candidate segment whose content identification information is not recorded in the cache table as the segment to be recognized; wherein the cache table records content identification information of audio segments whose contents have been recognized in advance. The obtaining the content recognition result corresponding to the segment to be recognized comprises:
6. The method of claim 4 or 5, wherein, determining an associated segment of the segment to be recognized from the plurality of target audio segments in a case where the segment to be recognized satisfies a preset segment merging condition; merging the segment to be recognized and the associated segment to obtain a merged segment; obtaining a content recognition result of the merged segment, and obtaining the content recognition result corresponding to the segment to be recognized based on the content recognition result of the merged segment. 7. The method of claim 6, wherein, The segment merging condition comprises: a segment duration being lower than a preset duration threshold, and / or adjacent segments being still to-be-identified segments. The associated segment of the to-be-identified segment comprises an adjacent segment of the to-be-identified segment.
8. The method of claim 6, wherein, The duration of the merged segment is greater than a preset duration threshold.
9. The method according to any one of claims 6-8, wherein, The obtaining of the content recognition result of the merged segment comprises: The merged segment is sent to a server, and a content recognition result returned by the server for the merged segment is received.
10. The method according to any one of claims 4-9, wherein, The obtaining of the content recognition result of the target audio based on the content recognition result corresponding to the to-be-identified segment and the content recognition result of the first audio segment comprises: The content recognition result of the target audio is obtained based on the content recognition result corresponding to the to-be-identified segment, the content recognition result of the first audio segment, and the starting time and the ending time corresponding to each of the plurality of target audio segments.
11. The method of any one of claims 1-10, wherein, The slice cache information of the historical audio comprises the starting time and the ending time of the historical audio segment; and the slice processing of the target audio based on the slice cache information of the historical audio to obtain the plurality of target audio segments comprises: The audio editing information is obtained; wherein the audio editing information is information of editing the historical audio to obtain the target audio; The slice information corresponding to the target audio is obtained based on the audio editing information and the starting time and the ending time of the historical audio segment; The target audio is slice-processed based on the slice information to obtain the plurality of target audio segments.
12. The method according to any one of claims 1-11, further comprising: The slice cache information of the target audio is obtained and stored based on the content recognition result of the target audio; The slice cache information of the target audio comprises the starting time, the ending time, the content recognition result, and the content identification information corresponding to each of the plurality of target audio segments.
13. An audio recognition apparatus, comprising: An audio slicing module configured to slice-process a target audio based on slice cache information of a historical audio to obtain a plurality of target audio segments; wherein the target audio is an audio obtained by editing the historical audio; and the slice cache information of the historical audio is information of a plurality of historical audio segments obtained by slice-processing the historical audio based on a content recognition result of the historical audio; A segment comparison module configured to compare the plurality of target audio segments and the plurality of historical audio segments, and determine a first audio segment based on a comparison result; wherein the first audio segment is a same segment in the plurality of target audio segments and the plurality of historical audio segments; An audio recognition module configured to obtain a content recognition result of the first audio segment based on the slice cache information of the historical audio, and obtain a content recognition result of the target audio based on the content recognition result of the first audio segment.
14. An electronic device, comprising: A storage device having a computer program stored thereon; A processing device configured to execute the computer program in the storage device to implement the audio recognition method according to any one of claims 1-12.
15. A computer readable storage medium storing a computer program, wherein, The computer program is for performing the audio recognition method of any of claims 1-12.
16. A computer program product comprising a computer program, wherein, The computer program, when executed by a processor, implements the audio recognition method of any of claims 1-12.
Citation Information
Patent Citations
Audio and video editing method and terminal
CN108337558A
Audio editing method, server and storage medium
CN111554329A
Sound and text realignment and information presentation method and device, electronic equipment and storage medium
CN113761865A
Audio data processing method and device, computer equipment and storage medium
CN116978381A
Intelligent video editing method and system based on large model
CN117812386A