A method, apparatus, electronic device, and readable storage medium for audio conversion

By obtaining the audio set duration and conversion time, determining the margin duration and sorting the audio filtering, the problem of poor offline audio conversion time is solved, more efficient audio conversion is achieved, and hardware resource investment is reduced.

CN114093362BActive Publication Date: 2025-07-04阳光保险集团股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111448918.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-01
Publication Date
2025-07-04
Estimated Expiration
2041-12-01

AI Technical Summary

Technical Problem

In the prior art, the high delay of offline audio conversion leads to poor conversion time. How to improve the timeline of offline audio conversion is an urgent problem.

Method used

By obtaining the set completion time and audio conversion time of each audio in the target audio set, the margin time is determined, and the audio is sorted and filtered according to the margin time, and the audio is converted separately according to the audio conversion sort, including priority filtering and parallel audio slicing and merging processing.

Benefits of technology

It improves the timeliness of audio conversion, is close to the efficiency of online audio conversion, makes full use of hardware resources, and reduces hardware resource investment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114093362B_ABST
    Figure CN114093362B_ABST
Patent Text Reader

Abstract

This application belongs to the field of audio technology and discloses a method, device, electronic device and readable storage medium for audio conversion. The method includes respectively obtaining the set completion duration and audio conversion duration corresponding to each audio in the target audio set; determining the margin duration of each audio respectively according to the set completion duration and audio conversion duration corresponding to each audio in the target audio set; sorting the audios according to the margin duration of each audio in the target audio set to obtain an audio conversion sorting; and performing audio conversion on each audio respectively according to the audio conversion sorting of each audio in the target set. In this way, the efficiency of audio conversion can be improved when performing offline audio conversion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio technology, and more particularly, to a method, apparatus, electronic device, and readable storage medium for audio conversion. Background Art

[0002] With the rapid development of the Internet, there are more and more scenarios for audio conversion applications. For example, in the sales industry or the customer service industry, it is necessary to contact customers by phone or voice. During the process, a large amount of audio is generated. Usually, the audio needs to be converted into text for content analysis, and the analysis results are used to better serve customers.

[0003] In the prior art, online speech recognition or offline speech recognition is usually adopted for audio conversion.

[0004] However, online audio conversion has low latency but also low efficiency, resulting in high usage costs; offline audio conversion has high latency, resulting in poor conversion timeliness.

[0005] Therefore, when performing offline audio conversion, how to improve the timeliness of audio conversion is a technical problem that needs to be solved. Summary of the Invention

[0006] The purpose of the embodiments of this application is to provide a method, apparatus, electronic device, and readable storage medium for audio conversion, so as to improve the timeliness of audio conversion when performing offline audio conversion.

[0007] On the one hand, a method for audio conversion is provided, including:

[0008] Obtain the set completion duration and audio conversion duration corresponding to each audio in the target audio set respectively;

[0009] Determine the margin duration of each audio respectively according to the set completion duration and audio conversion duration corresponding to each audio in the target audio set;

[0010] Sort the audios according to the margin durations of the audios in the target audio set to obtain an audio conversion sorting;

[0011] Perform audio conversion on each audio respectively according to the audio conversion sorting of each audio in the target set.

[0012] In the above implementation process, when performing audio conversion, the audios are sorted according to the margin durations of the audios, and each audio is converted respectively according to the audio conversion sorting, which improves the timeliness of audio conversion.

[0013] In one implementation, it further includes:

[0014] When it is determined that a preset set update condition is reached, execute the following steps:

[0015] Obtain the priority of each audio in the audio set to be processed respectively;

[0016] According to the priorities of the audios in the audio set to be processed, screen the audios in the audio set to be processed to obtain a target audio set, where the target audio set is a set of the screened audios.

[0017] In the above implementation process, screen the audios according to the priorities of the audios.

[0018] In one implementation manner, determining that a preset set update condition is met includes:

[0019] When it is determined that a set update instruction issued by the user is received, it is determined that the preset set update condition is met;

[0020] When it is determined that all the audios in the target audio set have completed audio conversion, it is determined that the preset set update condition is met;

[0021] When it is determined that a preset set update period is reached, it is determined that the preset set update condition is met.

[0022] In the above implementation process, when it is determined that the preset set update condition is met, update the set audios.

[0023] In one implementation manner, according to the priorities of the audios in the audio set to be processed, screen the audios in the audio set to be processed to obtain the target audio set, including:

[0024] Sort the audios in the audio set to be processed according to the priorities of the audios in the audio set to be processed to obtain an audio priority sorting;

[0025] According to the audio priority sorting, for each audio in the audio set to be processed in turn, perform the following steps:

[0026] Determine the sum of the audio durations of the audios in the target audio set to obtain a first total audio duration;

[0027] Determine the sum of the audio duration of an audio in the audio set to be processed and the first total audio duration to obtain a second total audio duration;

[0028] If the second total audio duration is lower than a set audio duration threshold, add an audio to the target audio set.

[0029] In the above implementation process, when adding an audio to the target audio set, ensure that the total audio duration of the target audio set does not exceed the set audio duration threshold.

[0030] In one implementation, before sorting the audio based on the margin duration of each audio in the target audio set to obtain the audio conversion sorting, the following steps are also included:

[0031] Filter out the parallel conversion audio with a margin duration less than the set duration threshold from the target audio set;

[0032] Split the parallel conversion audio to obtain multiple audio segments for each parallel conversion audio respectively;

[0033] Add the multiple audio segments of each parallel conversion audio to the target audio set;

[0034] Obtain the audio conversion duration of the multiple audio segments of each parallel conversion audio;

[0035] Determine the margin duration of the multiple audio segments of each parallel conversion audio respectively according to the audio conversion duration of the multiple audio segments of each parallel conversion audio.

[0036] In the above implementation process, if there is parallel conversion audio with a margin duration less than the set duration threshold, then split the parallel audio and determine the margin duration of each audio segment respectively.

[0037] In one implementation, after performing audio conversion on each audio respectively according to the audio conversion sorting of each audio, the following steps are also included:

[0038] If there are audio segments of at least one parallel conversion audio in the target audio set, then merge the audio conversion results of the audio segments after splitting each parallel conversion audio respectively to obtain the merged conversion result of each parallel conversion audio.

[0039] In the above implementation process, merge the audio conversion results of each audio segment to obtain the merged conversion result of the parallel conversion audio.

[0040] On the one hand, a device for audio conversion is provided, including:

[0041] An acquisition unit for respectively acquiring the set completion duration and the audio conversion duration corresponding to each audio in the target audio set;

[0042] A setting unit for respectively determining the margin duration of each audio according to the set completion duration and the audio conversion duration corresponding to each audio in the target audio set;

[0043] A sorting unit for sorting the audio according to the margin duration of each audio in the target audio set to obtain the audio conversion sorting;

[0044] A conversion unit for performing audio conversion on each audio respectively according to the audio conversion sorting of each audio in the target set.

[0045] In one implementation, the obtaining unit is further configured to:

[0046] When it is determined that a preset set update condition is met, perform the following steps:

[0047] Obtain the priority of each audio in the audio set to be processed respectively;

[0048] According to the priorities of the audios in the audio set to be processed, screen the audios in the audio set to be processed to obtain a target audio set, where the target audio set is a set of the screened audios.

[0049] In one implementation, the obtaining unit is specifically configured to:

[0050] When it is determined that a set update instruction issued by the user is received, determine that the preset set update condition is met;

[0051] When it is determined that audio conversion of each audio in the target audio set is completed, determine that the preset set update condition is met;

[0052] When it is determined that a preset set update period is reached, determine that the preset set update condition is met.

[0053] In one implementation, the obtaining unit is specifically configured to:

[0054] According to the priorities of the audios in the set to be processed, sort the audios in the set to be processed to obtain an audio priority sorting;

[0055] According to the audio priority sorting, for each audio in the audio set to be processed in sequence, perform the following steps:

[0056] Determine the sum of the audio durations of the audios in the target audio set to obtain a first total audio duration;

[0057] Determine the sum of the audio duration of an audio in the audio set to be processed and the first total audio duration to obtain a second total audio duration;

[0058] If the second total audio duration is lower than a set audio duration threshold, add an audio to the target audio set.

[0059] In one implementation, the sorting unit is further configured to:

[0060] Screen out parallel conversion audios with a margin duration less than a set duration threshold from the target audio set;

[0061] Split the parallel conversion audios to obtain multiple audio segments of each parallel conversion audio respectively;

[0062] Add multiple audio segments of each parallel conversion audio to the target audio set;

[0063] Obtain the audio conversion duration of multiple audio segments of each parallel conversion audio;

[0064] Determine the margin duration of multiple audio segments of each parallel conversion audio respectively according to the audio conversion duration of multiple audio segments of each parallel conversion audio.

[0065] In one implementation, the conversion unit is further configured to:

[0066] If there are audio segments of at least one parallel conversion audio in the target audio set, then merge the audio conversion results of the audio segments after splitting each parallel conversion audio respectively, and obtain the merged conversion results of each parallel conversion audio respectively.

[0067] On the one hand, an electronic device is provided, including a processor and a memory. The memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the steps of the method provided in any of the above optional implementation manners of audio conversion are run.

[0068] On the one hand, a readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps of the method provided in any of the above optional implementation manners of audio conversion are run.

[0069] On the one hand, a computer program product is provided. When the computer program product runs on a computer, the computer is enabled to execute the steps of the method provided in any of the above optional implementation manners of audio conversion.

[0070] In the method, device, electronic device and readable storage medium for audio conversion provided by the embodiments of the present application, the set completion duration and the audio conversion duration corresponding to each audio in the target audio set are obtained respectively; the margin duration of each audio is determined respectively according to the set completion duration and the audio conversion duration corresponding to each audio in the target audio set; the audios are sorted according to the margin durations of the audios in the target audio set to obtain an audio conversion sorting; and each audio is subjected to audio conversion respectively according to the audio conversion sorting of the audios in the target set. In this way, when performing audio conversion, the audios can be sorted according to the margin durations of the audios, and each audio can be converted respectively according to the audio conversion sorting, thereby improving the efficiency of audio conversion.

[0071] Other features and advantages of the present application will be set forth in the following description, and in part will be obvious from the description, or may be learned by practice of the present application. The objectives and other advantages of the present application may be realized and attained by the structure particularly pointed out in the written description, claims, as well as the drawings. Description of the Drawings

[0072] In order to illustrate the technical solutions of the embodiments of the present application more clearly, the following briefly introduces the drawings required to be used in the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0073] Figure 1 A schematic diagram of an audio conversion application scenario provided for an embodiment of the present application;

[0074] Figure 2 A flowchart of an implementation of a method for audio conversion provided for an embodiment of the present application;

[0075] Figure 3 An example diagram of a method for obtaining an audio margin duration provided for an embodiment of the present application;

[0076] Figure 4 A detailed flowchart of an implementation of a method for audio conversion provided for an embodiment of the present application;

[0077] Figure 5 A block diagram of a structure of a device for audio conversion provided for an embodiment of the present application;

[0078] Figure 6 A schematic diagram of a structure of an electronic device in an embodiment of the present application. Detailed Embodiments

[0079] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some, but not all, of the embodiments of the present application. The components of the embodiments of the present application described and illustrated in the drawings herein can be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the present application claimed, but merely represents selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.

[0080] It should be noted that like reference numerals and letters refer to like items in the following figures, and thus, once an item is defined in one figure, further definition and explanation thereof is not required in subsequent figures. Also, in the description of the present application, terms such as "first", "second", etc. are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance.

[0081] First, some terms involved in the embodiments of the present application are described to facilitate understanding by those skilled in the art.

[0082] Terminal device: It can be a mobile terminal, a fixed terminal or a portable terminal, such as a mobile phone, a site, a unit, a device, a multimedia computer, a multimedia tablet, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system device, a personal navigation device, a personal digital assistant, an audio / video player, a digital camera / video camera, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a game device or any combination thereof, including accessories and peripherals of these devices or any combination thereof. It is also foreseeable that the terminal device can support any type of interface for users (such as wearable devices), etc.

[0083] Server: It can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, and big data and artificial intelligence platforms.

[0084] Automatic Speech Recognition (ASR): It is a technology that converts human speech into text.

[0085] Natural Language Processing (NLP): A technology for human-machine interaction communication using the natural language used by humans for communication.

[0086] In order to improve the timeliness of audio conversion when performing audio conversion, the embodiments of the present application provide a method, an apparatus, an electronic device, and a readable storage medium for audio conversion.

[0087] Refer to Figure 1 As shown, it is a schematic diagram of an audio conversion application scenario provided by the embodiments of the present application. Figure 1 In this, the server generates a target audio set based on the set of audio to be processed, and performs audio conversion on the audio in the target audio set.

[0088] In one implementation, the server respectively obtains the set completion duration and the audio conversion duration corresponding to each audio in the target audio set, determines the margin duration of each audio respectively according to the set completion duration and the audio conversion duration corresponding to each audio in the target audio set, sorts the audios according to the margin durations of the audios in the target audio set to obtain an audio conversion sorting, and performs audio conversion on each audio respectively according to the audio conversion sorting of each audio in the target set.

[0089] For example, the server respectively obtains the priorities of four audios in the collected audio set to be processed, sorts the four audios in the audio set to be processed according to the priorities of the four audios in the audio set to be processed to obtain an audio priority sorting, and sequentially performs the following steps for the four audios in the audio set to be processed according to the audio priority sorting: determines the sum of the audio durations of three audios in the target audio set to obtain a first total audio duration, and determines the sum of the audio duration of Audio 1 in the audio set to be processed and the first total audio duration to obtain a second total audio duration. If the second total audio duration is lower than the set audio duration threshold, then Audio 1 is added to the target audio set. The server respectively obtains the set completion duration and the audio conversion duration corresponding to three audios in the target audio set, and determines the margin durations of the four audios respectively according to the set completion duration and the audio conversion duration corresponding to the three audios in the target audio set. The server filters out the parallel conversion audios with margin durations less than the set duration threshold from the target audio set, splits the parallel conversion audios to respectively obtain multiple audio segments of each parallel conversion audio, adds the multiple audio segments of each parallel conversion audio to the target audio set, and obtains the audio conversion durations of the multiple audio segments of each parallel conversion audio, and determines the margin durations of the multiple audio segments of each parallel conversion audio respectively according to the audio conversion durations of the multiple audio segments of each parallel conversion audio. The server sorts the audios according to the margin durations of the four audios in the target audio set to obtain an audio conversion sorting, and performs audio conversion on each audio respectively according to the audio conversion sorting of the four audios in the target set to obtain an audio conversion result. If there are audio segments of at least one parallel conversion audio in the target audio set, the server merges the audio conversion results of the audio segments after splitting each parallel conversion audio respectively to obtain the merged conversion result of each parallel conversion audio.

[0090] In this way, when performing audio conversion, the audios can be sequentially added to the target audio set according to the priorities of the audios, the audios can be sorted according to the margin durations of the audios in the target audio, and each audio can be converted respectively according to the audio conversion sorting, thereby improving the timeliness of audio conversion.

[0091] In the embodiments of the present application, only the server is taken as an example of the execution entity for illustration. In practical applications, the execution entity can also be other electronic devices such as terminal devices, which is not limited herein.

[0092] Refer to Figure 2 As shown, it is the implementation flowchart of a method for audio conversion provided by the embodiments of the present application. The specific implementation process of this method is as follows:

[0093] Step 200: Obtain the set completion duration and audio conversion duration corresponding to each audio in the target audio set respectively.

[0094] Specifically, the server obtains the set completion duration and audio conversion duration corresponding to each audio in the target audio set respectively.

[0095] Among them, the set completion duration is pre-set and is the duration that a user expects for an audio to complete audio conversion starting from the system successfully confirming receipt.

[0096] Optionally, the set completion duration can be set according to the actual application situation, which is not limited herein.

[0097] Among them, the audio conversion duration is the duration of an audio for audio conversion operation and is determined based on the audio conversion efficiency.

[0098] Among them, the audio conversion efficiency can be the duration of the audio converted by the audio conversion service working for one hour, which can be represented by R.

[0099] Optionally, the set completion duration can be represented by Tf, and the audio conversion duration can be represented by Tl.

[0100] Furthermore, before performing step 200, it is also possible to update the target audio set when a preset set update condition is reached.

[0101] Among them, when updating the target audio set, the following steps can be executed:

[0102] Step one: Obtain the priority of each audio in the to-be-processed audio set respectively.

[0103] Specifically, the server obtains the priority of each audio in the to-be-processed audio set respectively.

[0104] Among them, the priority of the audio can be the seat, team, regional department, and business scenario, etc., and the priority of the audio is set in advance.

[0105] Step two: Screen the audios in the to-be-processed audio set according to the priorities of the audios in the to-be-processed audio set to obtain the target audio set.

[0106] Specifically, when performing step 2, the following steps can be adopted:

[0107] Step 1: Sort each audio according to the priority of each audio to obtain the audio priority sorting.

[0108] Specifically, the server sorts each audio in the audio set to be processed according to the priority of each audio in the audio set to be processed, and obtains the audio priority sorting of each audio in the audio set to be processed.

[0109] Among them, the audio priority sorting of each audio in the audio set to be processed can be obtained by comparing the priorities of each audio in the audio set to be processed pairwise.

[0110] In practical applications, other sorting methods can be adopted according to the actual application situation, which is not limited here.

[0111] Step 2: According to the audio priority sorting, for each audio in the audio set to be processed in turn, perform the following steps:

[0112] Step A: Determine the sum of the audio durations of each audio in the target audio set to obtain the first total audio duration.

[0113] Specifically, the server adds up the audio durations of each audio in the target audio set to obtain the total audio duration of each audio in the target audio set as the first total audio duration.

[0114] In one implementation, there are five audios in the target audio set, and the corresponding audio durations of the five audios are 1 hour, 1 hour, 2 hours, 3 hours, and 2 hours respectively. Then, add up the audio durations of each audio in the target audio set, and the total audio duration of each audio in the target audio set is obtained as 9 hours, so the first total audio duration is 9 hours.

[0115] In practical applications, other methods can be adopted according to the actual application situation to determine the sum of the audio durations of each audio in the target audio set, which is not limited here.

[0116] Step B: Determine the sum of the audio duration of an audio in the audio set to be processed and the first total audio duration to obtain the second total audio duration.

[0117] Specifically, the server adds the audio duration of an audio in the audio set to be processed and the obtained first total audio duration to obtain the second total audio duration.

[0118] Step C: If the second total audio duration is lower than the set audio duration threshold, add the above-mentioned one audio to the target audio set.

[0119] Specifically, the server compares the total duration of the second audio with a set audio duration threshold. When it determines that the total duration of the second audio is not higher than the set audio duration threshold, it adds an audio to the target audio set.

[0120] Among them, the target audio set is a set of filtered audios.

[0121] In practical applications, the set audio duration can be set according to the actual application situation. For example, the set audio duration is 1 hour, and there is no limit here.

[0122] Furthermore, when the total duration of the second audio is higher than the set audio duration threshold, it refuses to add an audio to the target audio set and issues an exception warning.

[0123] Specifically, the server compares the total duration of the second audio with a set audio duration threshold. When it determines that the total duration of the second audio is higher than the set audio duration threshold, it refuses to add an audio to the target audio set and issues an exception warning.

[0124] Optionally, the exception warning can be an error message, which prompts the user in the form of a pop-up window, etc. In practical applications, other methods can be used for exception warning according to the actual application situation, and there is no limit here.

[0125] In one implementation, the audio duration of an audio in the set of audios to be processed is added to the total audio duration of the target audio set to obtain the total duration of the first audio, and the total duration of the first audio is compared with the total duration of the second audio. If the total duration of the first audio is less than the total duration of the second audio, an audio in the set of audios to be processed is added to the target audio set; otherwise, an audio in the set of audios to be processed is refused to be added to the target audio set. Among them, the total duration of the second audio is the maximum audio duration capacity value of the target audio sum.

[0126] Among them, the maximum audio duration capacity value can be 500 hours.

[0127] Among them, determining that the preset set update condition is met can be any one of the following conditions:

[0128] Condition 1: When it is determined that a set update instruction issued by the user is received, it is determined that the preset set update condition is met.

[0129] Specifically, the user sets a set update instruction. When the server receives the set update instruction issued by the user, it can be determined that the preset set update condition is met.

[0130] Among them, the preset set update condition is the condition for updating the target audio set, and the preset set update condition is set according to the actual application situation, and there is no limit here.

[0131] For example, the preset set update condition can be priority.

[0132] Condition 2: When it is determined that the audio conversion of each audio in the target audio set is completed, it is determined that the preset set update condition is met.

[0133] Specifically, when it is determined that the audio conversion of each audio in the target audio set is completed, the server can determine that the preset set update condition is met.

[0134] Condition 3: When it is determined that the preset set update period is reached, it is determined that the preset set update condition is met.

[0135] Specifically, when it is determined that the preset set update period is reached, the server can determine that the preset set update condition is met, and then update the conditions for screening the target audio.

[0136] Among them, the preset set update period can be a time period or a period based on the number of audios for which audio recognition is performed.

[0137] In one implementation, when the audio conversion of all the target audios in the target audio set is completed, the server determines that the preset set update condition is met, then updates the conditions for screening the audios in the target audio set, and according to the updated conditions, screens audios and adds them to the target audios.

[0138] In one implementation, according to the priority of the audios, each audio in the audio set to be processed is screened, and the audios with higher priority are screened out to obtain the target audio set. When it is determined that the audio conversion of each audio in the target audio set is completed, the server determines that the preset set update condition is met, then selects the remaining audios with lower priority from the audio set to be processed, and adds the above-mentioned audios to the target audio set in sequence for audio conversion.

[0139] In practical applications, the preset set update period can be set according to the actual application situation, and no limitation is made here.

[0140] In practical applications, meeting the preset set update condition can also be other conditions, which can be set according to the actual application situation, and no limitation is made here.

[0141] Step 201: According to the set completion duration and audio conversion duration corresponding to each audio in the target audio set, determine the margin duration of each audio respectively.

[0142] Specifically, the server subtracts the audio conversion duration from the set completion duration corresponding to each audio in the obtained target audio set, and the result is the margin duration of each audio.

[0143] Among them, the margin duration is the extra duration after the audio conversion is completed within the set completion duration.

[0144] Optionally, the margin duration can be represented by To.

[0145] Refer to Figure 3 As shown, it is an example diagram of a method for obtaining the audio margin duration provided by an embodiment of the present application. The example diagram includes two modules, a traffic module and an audio recognition module. After the voice call is hung up, the traffic module stores the audio. After the storage is completed, the set completion duration of the audio is obtained. The audio recognition module downloads the audio. Among them, the set completion duration includes the conversion waiting duration, the audio conversion duration, and the margin duration. The server obtains the margin duration of the audio according to the set completion duration, the conversion waiting duration, and the audio conversion duration of the audio.

[0146] Among them, after the traffic module stores the audio, it records the audio information.

[0147] Optionally, the traffic module uses data structures such as circular queues or red-black trees to store the audio.

[0148] Among them, the audio information can include the priority of the audio, the audio duration, the recognition duration, the time when the recording is received, and the set completion duration, etc.

[0149] Among them, the conversion waiting duration can be the duration from when the audio is ready to when the audio starts to be converted.

[0150] Optionally, it can be represented by Tw.

[0151] In this way, the margin duration of the audio can be obtained through the set completion duration, the conversion waiting duration, and the audio conversion duration of the audio.

[0152] In one implementation, the set completion duration of an audio is 3 hours. Among them, the audio conversion duration of this audio is 2 hours, then the margin duration of this audio is 1 hour.

[0153] Step 202: Sort the audio according to the margin duration of each audio in the target audio set to obtain the audio conversion sorting.

[0154] Specifically, the server sorts each audio in the obtained target audio set according to the margin duration of each audio in the target audio set to obtain the audio conversion sorting of each audio in the target audio set.

[0155] Among them, the audio conversion sorting of each audio in the target audio set can be obtained by comparing each audio in the target audio set pairwise.

[0156] In practical applications, other sorting methods can also be adopted according to the actual application situation, which is not limited herein.

[0157] Furthermore, before performing step 202, parallel conversion audios with a margin duration less than a set duration threshold can also be segmented.

[0158] Among them, when segmenting parallel conversion audios with a margin duration less than a set duration threshold, the following steps can be executed:

[0159] Step 1: Screen out parallel conversion audios with a margin duration less than a set duration threshold from the target audio set.

[0160] Specifically, the server compares the margin duration of each audio in the target audio set with the set duration threshold, and screens out parallel conversion audios with a margin duration less than the set duration threshold.

[0161] In practical applications, the set duration threshold can be set according to the actual application situation, which is not limited herein.

[0162] Step 2: Segment the parallel conversion audios to obtain multiple audio segments of each parallel conversion audio respectively.

[0163] Specifically, the server segments the screened parallel conversion audios to obtain multiple audio segments of each parallel conversion audio respectively.

[0164] Furthermore, special marks can also be made on the multiple audio segments after segmenting the parallel conversion audios. For example, number the multiple sub-steps. In practical applications, other special marks can also be adopted, which is not limited herein.

[0165] Step 3: Add the multiple audio segments of each parallel conversion audio to the target audio set.

[0166] Specifically, the server adds the multiple audio segments obtained after segmenting each parallel conversion audio to the target audio set.

[0167] Step 4: Obtain the audio conversion duration of the multiple audio segments of each parallel conversion audio.

[0168] Specifically, the server obtains the audio conversion duration of the multiple audio segments of each parallel conversion audio.

[0169] Among them, the audio conversion duration is related to the audio duration.

[0170] Step 5: Determine the margin duration of the multiple audio segments of each parallel conversion audio respectively according to the audio conversion duration of the multiple audio segments of each parallel conversion audio.

[0171] Specifically, the server determines the margin duration of multiple audio segments of each parallel conversion audio based on the obtained audio conversion duration of the multiple audio segments of each parallel conversion audio and the previously obtained set completion duration.

[0172] Among them, the difference between the set completion duration and the audio conversion duration is determined to obtain the margin duration.

[0173] In one implementation, the server separately obtains the margin duration of each audio in the target audio set. If the margin durations of all audios are greater than 0, then according to the margin durations of the audios in the target audio set, the audios are sorted to obtain an audio conversion sorting, and audio conversion operations are performed according to the audio conversion sorting. If there is an audio with a margin duration less than 0 among the audios, then the audio with a margin duration less than 0 is screened out, and the audio is segmented, and the margin duration after segmentation is recalculated, and then the audio with the smallest margin duration is selected for audio conversion operations.

[0174] Among them, the method of segmenting the audio can be to evenly segment it into at least two segments, or the product of the margin duration and the set coefficient can be determined as the new margin duration, and the audio duration is calculated according to the new margin duration, so that the audio can be segmented.

[0175] Among them, the set coefficient can be a coefficient greater than 0 and less than 1, or can be set according to the actual application situation. For example, the set coefficient can be 0.5, and there is no limitation here.

[0176] Step 203: Perform audio conversion on each audio according to the audio conversion sorting of the audios in the target set.

[0177] Specifically, the server performs audio conversion on each audio according to the audio conversion sorting of the audios in the target set to obtain an audio conversion result.

[0178] Among them, after performing step 203, if there are at least one audio segment of a parallel conversion audio in the target audio set, then the audio conversion results of the audio segments after splitting each parallel conversion audio are merged respectively to obtain the merged conversion result of each parallel conversion audio.

[0179] Specifically, if there are at least one audio segment of a parallel conversion audio in the target audio set, then the server merges the audio conversion results of the audio segments after splitting each parallel conversion audio respectively to obtain the merged conversion result of each parallel conversion audio.

[0180] In one implementation, the server can perform special marking on multiple audio segments after each parallel conversion audio is segmented, and after performing audio conversion on the audio segments after each parallel conversion audio is segmented, the server can merge the audio conversion results of the audio segments after each parallel conversion audio is segmented according to the special marking of the multiple audio segments after each parallel conversion audio is segmented, and respectively obtain the merged conversion results of each parallel conversion audio.

[0181] For example, number the multiple sub-steps. In practical applications, other special markings can also be used, which are not limited here.

[0182] Refer to Figure 4 As shown, it is a detailed implementation flowchart of a method for audio conversion provided by an embodiment of the present application. The specific implementation process of this method is as follows:

[0183] Step 400: The server respectively obtains the priority of each audio in the audio set to be processed.

[0184] Step 401: The server sorts the audios in the audio set to be processed according to the priorities of the audios in the audio set to be processed, and obtains an audio priority sorting.

[0185] Step 402: The server obtains the total duration of the current target audio set.

[0186] Step 403: The server determines whether the sum of the audio duration of an audio in the audio priority sorting and the total duration of the current target audio set is not less than a set audio duration threshold. If so, execute Step 405; otherwise, execute Step 404.

[0187] Step 404: The server adds the above audio to the target audio set according to the audio priority sorting.

[0188] It should be noted that after executing Step 404, execute Step 402.

[0189] Step 405: The server refuses to add the above audio to the target audio set.

[0190] Step 406: The server respectively obtains the set completion duration and the audio conversion duration corresponding to each audio in the target audio set.

[0191] Step 407: The server respectively determines the margin duration of each audio according to the set completion duration and the audio conversion duration corresponding to each audio in the target audio set.

[0192] Step 408: The server filters out the parallel conversion audios with a margin duration less than the set duration threshold from the target audio set.

[0193] Step 409: The server segments each parallel converted audio to obtain multiple audio segments of each parallel converted audio respectively.

[0194] Step 410: The server adds the multiple audio segments of each parallel converted audio to the target audio set.

[0195] Step 411: The server obtains the audio conversion duration of the multiple audio segments of each parallel converted audio, and determines the margin duration of the multiple audio segments of each parallel converted audio respectively.

[0196] Step 412: The server sorts the audios according to the margin duration of each audio in the target audio set to obtain the audio conversion sorting.

[0197] Step 413: The server performs audio conversion on each audio respectively according to the audio conversion sorting of each audio in the target set to obtain the audio conversion result.

[0198] Step 414: The server merges the audio conversion results of the audio segments after segmenting each parallel converted audio respectively to obtain the merged conversion result of each parallel converted audio respectively.

[0199] Specifically, when performing steps 400 - 414, for the specific steps, refer to the above steps 200 - 203, which will not be elaborated here.

[0200] Traditional audio conversion usually adopts automatic speech recognition technology and natural language processing technology. Through two technical solutions, one is online audio conversion, which is real-time conversion, generally used in scenarios such as voice customer service and robot conversations, with high costs. The other is offline audio conversion, generally used in scenarios such as voice quality inspection and recording analysis. After the voice call is hung up, the audio conversion operation can start. The timeliness of audio conversion is poor, and both methods require a large amount of hardware resources.

[0201] In the embodiments of the present application, when performing audio conversion, according to the priority of the audio, the audio is added to the target audio set in sequence, and the margin duration of each audio in the target audio is determined. For the audio with a duration less than the set duration threshold, it is segmented into multiple audio segments, and the audios are sorted according to the margin duration, and audio conversion is performed on each audio respectively according to the audio conversion sorting, which improves the timeliness of audio conversion. Without increasing hardware resources, based on the offline audio conversion solution, it is close to the efficiency of online audio conversion, improves the timeliness of audio conversion, makes full use of hardware resources, and reduces the investment in hardware resources.

[0202] Based on the same inventive concept, an embodiment of the present application further provides an audio conversion device. Since the principles of the above device and equipment for solving problems are similar to those of an audio conversion method, the implementation of the above device can refer to the implementation of the method, and the repeated parts will not be elaborated again.

[0203] As Figure 5 shown, it is a schematic structural diagram of an audio conversion device provided by an embodiment of the present application, including:

[0204] An acquisition unit 501, configured to respectively acquire the set completion duration and the audio conversion duration corresponding to each audio in the target audio set;

[0205] A setting unit 502, configured to respectively determine the margin duration of each audio according to the set completion duration and the audio conversion duration corresponding to each audio in the target audio set;

[0206] A sorting unit 503, configured to sort the audios according to the margin duration of each audio in the target audio set to obtain an audio conversion sorting;

[0207] A conversion unit 504, configured to perform audio conversion on each audio respectively according to the audio conversion sorting of each audio in the target set.

[0208] In an implementation manner, the acquisition unit 501 is further configured to:

[0209] When it is determined that a preset set update condition is met, execute the following steps:

[0210] Respectively acquire the priority of each audio in the to-be-processed audio set;

[0211] According to the priorities of the audios in the to-be-processed audio set, screen the audios in the to-be-processed audio set to obtain a target audio set, where the target audio set is a set of the screened audios.

[0212] In an implementation manner, the acquisition unit 501 is specifically configured to:

[0213] When it is determined that a set update instruction issued by the user is received, determine that the preset set update condition is met;

[0214] When it is determined that all the audios in the target audio set have completed audio conversion, determine that the preset set update condition is met;

[0215] When it is determined that a preset set update period is reached, determine that the preset set update condition is met.

[0216] In an implementation manner, the acquisition unit 501 is specifically configured to:

[0217] Sort each audio in the set to be processed according to the priority of each audio in the set to be processed, and obtain the audio priority sorting;

[0218] According to the audio priority sorting, for each audio in the set of audios to be processed in turn, perform the following steps:

[0219] Determine the sum of the audio durations of each audio in the target audio set, and obtain the first total audio duration;

[0220] Determine the sum of the audio duration of an audio in the set of audios to be processed and the first total audio duration, and obtain the second total audio duration;

[0221] If the second total audio duration is lower than the set audio duration threshold, add an audio to the target audio set.

[0222] In one implementation, the sorting unit 503 is further configured to:

[0223] Filter out the parallel conversion audios with the margin duration less than the set duration threshold from the target audio set;

[0224] Split the parallel conversion audios to obtain multiple audio segments of each parallel conversion audio respectively;

[0225] Add the multiple audio segments of each parallel conversion audio to the target audio set;

[0226] Obtain the audio conversion duration of the multiple audio segments of each parallel conversion audio;

[0227] According to the audio conversion duration of the multiple audio segments of each parallel conversion audio, determine the margin duration of the multiple audio segments of each parallel conversion audio respectively.

[0228] In one implementation, the conversion unit 504 is further configured to:

[0229] If there are audio segments of at least one parallel conversion audio in the target audio set, then merge the audio conversion results of the audio segments after splitting each parallel conversion audio respectively to obtain the merged conversion results of each parallel conversion audio.

[0230] In a method, apparatus, electronic device, and readable storage medium for audio conversion provided by an embodiment of the present application, the set completion duration and the audio conversion duration corresponding to each audio in the target audio set are respectively obtained; according to the set completion duration and the audio conversion duration corresponding to each audio in the target audio set, the margin duration of each audio is respectively determined; according to the margin durations of the audios in the target audio set, the audios are sorted to obtain an audio conversion sorting; and according to the audio conversion sorting of the audios in the target set, each audio is respectively subjected to audio conversion. In this way, when performing audio conversion, the audios can be sorted according to the margin duration of the audios, and each audio can be respectively converted according to the audio conversion sorting, thereby improving the timeliness of audio conversion.

[0231] Figure 6 FIG. shows a schematic structural diagram of an electronic device 6000. Refer to Figure 6 As shown, the electronic device 6000 includes: a processor 6010 and a memory 6020. Optionally, it may further include a power supply 6030, a display unit 6040, and an input unit 6050.

[0232] The processor 6010 is the control center of the electronic device 6000, connects various components through various interfaces and lines, and executes various functions of the electronic device 6000 by running or executing software programs and / or data stored in the memory 6020, thereby performing overall monitoring of the electronic device 6000.

[0233] In an embodiment of the present application, when the processor 6010 calls a computer program stored in the memory 6020, it executes the method for audio conversion provided by the embodiment shown in Figure 2 the embodiment shown in the figure.

[0234] Optionally, the processor 6010 may include one or more processing units; preferably, the processor 6010 may integrate an application processor and a modem processor, where the application processor mainly processes the operating system, user interface, applications, etc., and the modem processor mainly processes wireless communication. It can be understood that the above modem processor may not be integrated into the processor 6010. In some embodiments, the processor and the memory may be implemented on a single chip, and in some embodiments, they may also be separately implemented on independent chips.

[0235] The memory 6020 may mainly include a program storage area and a data storage area. Among them, the program storage area may store the operating system, various applications, etc.; the data storage area may store data created according to the use of the electronic device 6000, etc. In addition, the memory 6020 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices, etc.

[0236] The electronic device 6000 further includes a power source 6030 (such as a battery) for supplying power to each component. The power source can be logically connected to the processor 6010 through a power management system, so as to manage functions such as charging, discharging, and power consumption through the power management system.

[0237] The display unit 6040 can be used to display information input by the user or information provided to the user, as well as various menus of the electronic device 6000. In the embodiments of the present invention, it is mainly used to display the display interfaces of various applications in the electronic device 6000 and objects such as text and pictures displayed in the display interfaces. The display unit 6040 may include a display panel 6041. The display panel 6041 can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc.

[0238] The input unit 6050 can be used to receive information such as numbers or characters input by the user. The input unit 6050 may include a touch panel 6051 and other input devices 6052. Among them, the touch panel 6051, also known as a touch screen, can collect touch operations of the user thereon or nearby (such as operations of the user using a finger, a stylus, or any suitable object or accessory on or near the touch panel 6051).

[0239] Specifically, the touch panel 6051 can detect the touch operation of the user, detect the signals brought by the touch operation, convert these signals into contact coordinates, send them to the processor 6010, and receive and execute the commands sent by the processor 6010. In addition, the touch panel 6051 can be implemented in multiple types such as resistive, capacitive, infrared, and surface acoustic wave. The other input devices 6052 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, etc.

[0240] Of course, the touch panel 6051 can cover the display panel 6041. After the touch panel 6051 detects a touch operation thereon or nearby, it transmits it to the processor 6010 to determine the type of touch event. Subsequently, the processor 6010 provides a corresponding visual output on the display panel 6041 according to the type of touch event. Although in Figure 6 the touch panel 6051 and the display panel 6041 are implemented as two independent components to realize the input and output functions of the electronic device 6000, in some embodiments, the touch panel 6051 and the display panel 6041 can be integrated to realize the input and output functions of the electronic device 6000.

[0241] The electronic device 6000 may further include one or more sensors, such as a pressure sensor, a gravitational acceleration sensor, a proximity light sensor, etc. Of course, according to the needs in specific applications, the above-mentioned electronic device 6000 may further include other components such as a camera. Since these components are not the key components used in the embodiments of the present application, therefore, in Figure 6 it is not shown and will not be described in detail.

[0242] Those skilled in the art can understand that Figure 6 this is only an example of an electronic device and does not constitute a limitation on the electronic device. It may include more or fewer components than shown in the figure, or combine certain components, or different components.

[0243] In the embodiments of the present application, a readable storage medium stores a computer program. When the computer program is executed by a processor, the communication device can execute each step in the above embodiments.

[0244] For the convenience of description, the above parts are divided into respective modules (or units) according to functions and described separately. Of course, when implementing the present application, the functions of the respective modules (or units) can be implemented in the same or multiple software or hardware.

[0245] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0246] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0247] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the processes Figure 1 or blocks Figure 1 specified in one or more of the processes and / or blocks.

[0248] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes Figure 1 or blocks Figure 1 specified in one or more of the blocks.

[0249] Although the preferred embodiments of the present application have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn of the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present application.

[0250] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.

Claims

1. A method for audio conversion, characterized in that, including: respectively obtaining a set completion duration and an audio conversion duration corresponding to each audio in the target audio set; respectively determining a margin duration for each audio according to the set completion duration and the audio conversion duration corresponding to each audio in the target audio set; screening out parallel conversion audios with a margin duration less than a set duration threshold from the target audio set; splitting the parallel conversion audios to respectively obtain multiple audio segments for each parallel conversion audio; adding the multiple audio segments of each parallel conversion audio to the target audio set; obtaining the audio conversion duration of the multiple audio segments of each parallel conversion audio; respectively determining the margin duration of the multiple audio segments of each parallel conversion audio according to the audio conversion duration of the multiple audio segments of each parallel conversion audio; sorting the audios according to the margin duration of each audio in the target audio set to obtain an audio conversion sorting; performing audio conversion on each audio respectively according to the audio conversion sorting of each audio in the target audio set.

2. The method according to claim 1, wherein further including: when it is determined that a preset set update condition is met, perform the following steps: respectively obtaining the priority of each audio in the audio set to be processed; screening each audio in the audio set to be processed according to the priority of each audio in the audio set to be processed to obtain the target audio set, where the target audio set is a set of screened-out audios.

3. The method according to claim 2, wherein The determination of reaching the preset set update condition includes: when it is determined that a set update instruction issued by the user is received, it is determined that the preset set update condition is met; when it is determined that each audio in the target audio set has completed audio conversion, it is determined that the preset set update condition is met; when it is determined that a preset set update period is reached, it is determined that the preset set update condition is met.

4. The method according to claim 2, wherein The screening of each audio in the audio set to be processed according to the priority of each audio in the audio set to be processed to obtain the target audio set includes: sorting each audio in the audio set to be processed according to the priority of each audio in the audio set to be processed to obtain an audio priority sorting; sequentially for each audio in the audio set to be processed according to the audio priority sorting, perform the following steps: determining the sum of the audio durations of each audio in the target audio set to obtain a first total audio duration; determining the sum of the audio duration of an audio in the audio set to be processed and the first total audio duration to obtain a second total audio duration; if the second total audio duration is lower than a set audio duration threshold, add the one audio to the target audio set.

5. The method according to claim 1, wherein After performing audio conversion on each audio respectively according to the audio conversion sorting of each audio in the target audio set, it further includes: if there are audio segments of at least one parallel conversion audio in the target audio set, respectively merge the audio conversion results of the audio segments after splitting each parallel conversion audio to respectively obtain a merged conversion result for each parallel conversion audio.

6. An audio conversion device, characterized in that, including: an obtaining unit for respectively obtaining a set completion duration and an audio conversion duration corresponding to each audio in the target audio set; A setting unit, configured to respectively determine the margin duration of each audio according to the set completion duration and the audio conversion duration corresponding to each audio in the target audio set; The setting unit is further configured to: screen out parallel conversion audios with a margin duration less than a set duration threshold from the target audio set; split the parallel conversion audios to respectively obtain multiple audio segments of each parallel conversion audio; add the multiple audio segments of each parallel conversion audio to the target audio set; obtain the audio conversion duration of the multiple audio segments of each parallel conversion audio; and respectively determine the margin duration of the multiple audio segments of each parallel conversion audio according to the audio conversion duration of the multiple audio segments of each parallel conversion audio; A sorting unit, configured to sort the audios according to the margin duration of each audio in the target audio set to obtain an audio conversion sorting; A conversion unit, configured to respectively perform audio conversion on each audio according to the audio conversion sorting of each audio in the target audio set.

7. The device according to claim 6, characterized in that, The obtaining unit is further configured to: When it is determined that a preset set update condition is met, perform the following steps: Respectively obtain the priority of each audio in the audio set to be processed; According to the priorities of the audios in the audio set to be processed, screen the audios in the audio set to be processed to obtain the target audio set, where the target audio set is a set of the screened-out audios.

8. An electronic device, characterized in that, It includes a processor and a memory, and the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the method according to any one of claims 1-5 is run.

9. A readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1-5 is run.

Citation Information

Patent Citations

  • Audio scene classification model generation method and device, equipment and storage medium

    CN111653290A

  • Video synthesis method and device, electronic equipment and readable storage medium

    CN113132780A