Audio real-time playback method, device, equipment and readable storage medium

By preprocessing and segmented editing of the audio data file to be played, combined with the output characteristics of the preset audio buffer, real-time audio playback is realized, solving the problem of waiting process in the prior art and reducing the waiting time.

CN116760923BActive Publication Date: 2025-05-06CHINA MERCHANTS BANK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310714665.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-15
Publication Date
2025-05-06
Estimated Expiration
2043-06-15

AI Technical Summary

Technical Problem

When playing the phone recordings between customers and customer service, the existing technology needs to convert the audio into text, desensitization and then convert it to audio, resulting in the process that needs to wait for the entire process to end and the waiting time is long.

Method used

By obtaining the audio data file to be played, preprocessing it, determining the part to be edited, and processing the audio file according to the preset audio buffer segment caches and edits the audio file, synchronously writing the edited data to the output stream to achieve real-time playback.

Benefits of technology

The waiting time for the re-listening recording content is reduced, and the processed part is played in real time during the process of the audio data file being processed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116760923B_ABST
    Figure CN116760923B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device, equipment and readable storage medium for real-time audio playback, the method comprising the steps of: obtaining an audio data file to be played; preprocessing the audio data file, and determining the part to be edited in the audio data file based on the result of the preprocessing; caching the audio file in the audio data file in segments according to a preset audio buffer and the part to be edited, and editing and processing them in sequence; when caching the audio file in the preset audio buffer, synchronously writing the edited data to the output stream, and playing the corresponding audio content in real time according to the output stream. The present application implements preprocessing of the audio data file, determining the part to be edited of the audio data file, and directly processing the part to be edited when caching according to the preset audio buffer, and realizing the actions of caching data and writing to the output stream at the same time, and realizing the effect of real-time audio playback.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of audio processing technology, and in particular to a real-time audio playback method, device, equipment and readable storage medium. Background Art

[0002] When people call the corresponding telephone customer service, in order to ensure that the customer service provides good service, the conversation between the customer and the customer service is usually recorded and played in certain scenarios, for example, to check whether the service provided by the customer service meets the standards.

[0003] When a customer communicates with customer service over the phone, the content of the communication may involve some private information, such as ID number, phone number, etc., which may lead to the disclosure of the customer's personal privacy when the recording is played. Usually, the corresponding audio processing technology is used to convert the recording into text first, desensitize part of the content in the text, and then convert it into audio based on the desensitized text.

[0004] However, in the above method, when converting audio to text and then converting the text to audio, the processing process needs to be processed offline, that is, relevant personnel need to wait for the entire processing flow of the audio data to be completed before playing the audio, and the waiting time is long. Summary of the invention

[0005] In view of this, the present application provides a real-time audio playback method, device, equipment and readable storage medium, aiming to reduce the waiting time for re-listening to the recorded content.

[0006] To achieve the above object, the present application provides a method for real-time audio playback, which comprises the following steps:

[0007] Get the audio data file to be played;

[0008] Preprocessing the audio data file, and determining a portion to be edited in the audio data file according to a result of the preprocessing;

[0009] According to the preset audio buffer and the part to be edited, the audio files in the audio data file are cached in segments, and the audio files are edited in sequence;

[0010] When the preset audio buffer caches the audio file, the edited data is synchronously written into the output stream, and the corresponding audio content is played in real time according to the output stream.

[0011] Exemplarily, the audio data file includes an audio file and a text file obtained by text conversion according to the content of the audio file, and the step of preprocessing the audio data file includes:

[0012] Read the content of the text file and parse it to obtain a text information list of the audio conversation content;

[0013] According to the text information list, analyzing the first audio duration occupied by each word in the audio file and the second audio duration between adjacent words;

[0014] According to the first audio duration and the second audio duration, predict the total audio duration of the audio file after editing; wherein the preprocessing includes parsing the text file to predict the total audio duration after editing.

[0015] Exemplarily, the part to be edited includes a first part and a second part, and the step of determining the part to be edited in the audio data file according to the result of the preprocessing includes:

[0016] Determine, according to the text information list, a first part in the audio data file that needs to be desensitized;

[0017] A second part that needs to be compressed in the audio data file is determined according to the text information list and the total duration of the audio.

[0018] Exemplarily, the step of determining the first part that needs to be desensitized in the audio data file according to the text information list includes:

[0019] Acquire a semantic information list matching the content in the text information list;

[0020] Analyzing the digital word content in the text information list according to the semantic information list, and determining adjacent word content adjacent to the digital word content;

[0021] If the adjacent word content is a number, then analyzing whether the adjacent word of the adjacent word content is a number, until the analyzed content is non-numeric content;

[0022] The words that have been determined to be numbers are the first ones to be desensitized.

[0023] Exemplarily, the step of predicting the total duration of the audio after the audio file is edited based on the first audio duration and the second audio duration includes:

[0024] When the duration of the second audio is greater than a preset duration, predicting a compressed audio duration after compressing the duration of the second audio;

[0025] The total duration of the audio after the audio file is edited is predicted according to the compressed audio duration and the first audio duration.

[0026] Exemplarily, the part to be edited includes a first part and a second part, and the step of caching the audio files in the audio data file in segments according to the preset audio buffer and the part to be edited, and editing them in sequence, includes:

[0027] According to the preset audio buffer, when caching the audio file in the audio data file in segments, determining the first part and the second part involved in the file of the current cache segment;

[0028] The first part is converted into a sine wave audio with a preset fixed frequency, and the second part is cut to implement editing processing of the segmented cache file.

[0029] Exemplarily, when the preset audio buffer caches the audio file, the step of synchronously writing the edited data to the output stream includes:

[0030] When the preset audio buffer caches the audio file, the edited data is synchronously converted into a real-time audio stream in a PCM-encoded WAV format and written into an output stream.

[0031] Exemplarily, to achieve the above purpose, the present application also provides a real-time audio playback device, the device comprising:

[0032] An acquisition module, used to acquire the audio data file to be played;

[0033] A determination module, used for preprocessing the audio data file and determining the part to be edited in the audio data file according to the result of the preprocessing;

[0034] A processing module, used for caching the audio files in the audio data file in segments according to the preset audio buffer and the part to be edited, and performing editing processing on the audio files;

[0035] The playing module is used to synchronously write the edited data into the output stream when the preset audio buffer caches the audio file, and play the corresponding audio content in real time according to the output stream.

[0036] Exemplarily, to achieve the above-mentioned purpose, the present application also provides a real-time audio playback device, which includes: a memory, a processor, and a real-time audio playback program stored in the memory and executable on the processor, wherein the real-time audio playback program is configured to implement the steps of the real-time audio playback method as described above.

[0037] Exemplarily, to achieve the above-mentioned purpose, the present application also provides a computer-readable storage medium, on which a real-time audio playback program is stored, and when the real-time audio playback program is executed by a processor, the steps of the real-time audio playback method as described above are implemented.

[0038] Compared with the related art, in which relevant personnel need to wait for the entire processing flow of the audio data to be completed before playing the video, that is, the waiting time for playing the audio is long, in the present application, an audio data file to be played is obtained; the audio data file is preprocessed, and based on the result of the preprocessing, the part to be edited in the audio data file is determined; according to the preset audio buffer and the part to be edited, the audio files in the audio data file are cached in segments, and edited in turn; when the audio file is cached in the preset audio buffer, the edited data is synchronously written to the output stream, and the corresponding audio content is played in real time according to the output stream. That is to say, by preprocessing the audio to be played and determining the part to be edited in the audio data file based on the result of the preprocessing, the audio data file can be segmented and cached according to the preset audio buffer, and the edited and processed data can be written to the output stream when the preset audio buffer caches the audio data file, that is, the cached file is processed in a targeted manner, and at the same time, the processed part of the file is played synchronously, so that the relevant personnel do not need to wait for the file processing process, but can directly play the processed part in real time during the audio data file processing process, thereby realizing real-time playback of the audio data file that needs to be processed, thereby reducing the waiting time for re-listening to the recorded content. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 This is a flowchart of the first embodiment of the real-time audio playback method of the present application;

[0040] Figure 2 This is a detailed flowchart of step S120 in the first embodiment of the real-time audio playback method of this application;

[0041] Figure 3 A schematic diagram of the preset audio buffer application process of the real-time audio playback method of this application;

[0042] Figure 4 This is a schematic diagram of the structure of the hardware operating environment involved in the embodiment of the present application.

[0043] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0044] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0045] This application provides a real-time audio playback method. Figure 1 , Figure 1 This is a flowchart of the first embodiment of the real-time audio playback method of the present application.

[0046] The embodiment of the present application provides an embodiment of a method for real-time audio playback. It should be noted that although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here. For ease of description, the following omits the steps of the method for real-time audio playback, and the method for real-time audio playback includes:

[0047] Step S110: obtaining the audio data file to be played;

[0048] In this embodiment, the recording of the conversation between the customer service and the customer is mainly voice. In order to take into account the recording efficiency and occupied space, the VOX format is adopted, and the adaptive differential pulse code modulation (ADPCM) encoding is adopted. Among them, because there may be some audio content such as certificate numbers and card numbers in the recording, there may be a risk of leaking customer privacy when listening back. At present, most of the technical solutions for similar audio desensitization are offline processing of audio, generating text, desensitizing the text, and then comparing the generated audio, and do not support the VOX audio format. These solutions have some defects: offline processing and playback, which takes a long time and requires a large storage space; desensitizing the text and then generating audio, it is difficult to ensure the accuracy of the un-desensitized part; VOX audio format is not supported, so it cannot be played directly on the HTML5 page; in the process of handling business for customers, customers may need to wait, and the long silent time in the recording is not processed, which will result in a poor experience when listening back.

[0049] Therefore, in order to avoid the above-mentioned problems, an audio playback method is proposed in this embodiment, which mainly solves the problem that relevant personnel need to wait for a long time for the audio processing process when listening to the audio. Secondly, improvements are made to the poor accuracy after text desensitization, the data after the translated text to audio does not support the VOX audio format, and the long silence time in the recording affects the listening experience.

[0050] Among them, it should be noted that when the relevant personnel want to listen to the corresponding audio data file again, they will perform the corresponding operation according to the corresponding operation interface. The operation instruction will be sent to the system. The system will retrieve the corresponding data according to the instruction content and process the data accordingly. After the data processing is completed, it will be written to the output stream to generate an audio playback output to achieve the effect that the relevant personnel can hear the audio content of the recording.

[0051] When the system receives a request from relevant personnel to listen to the recording again, it will retrieve the corresponding audio data file from the database of the audio data files storing the recording. The audio data files are all provided with unique identification information. For example, a unique audio data file identifier is generated based on the time when the recording was generated or the customer service corresponding to the recording. The recording content can be matched from the database based on the identifier.

[0052] Among them, the audio data file to be played refers to the recording that the relevant personnel want to listen to again. The recording and audio data files mentioned below are the same content, and both refer to the recording data of the telephone communication between the customer and the customer service.

[0053] Step S120: preprocessing the audio data file, and determining the part to be edited in the audio data file according to the result of the preprocessing;

[0054] After obtaining the audio data file, in order to avoid the possibility of customer-related privacy being involved in the recording when listening back to the recording, it is necessary to edit part of the content in the audio data file, wherein the editing process may include desensitization, cutting or compression, etc. The main purpose of the editing process is to edit the sensitive information and useless audio content in the audio data file, and preprocessing is the pre-processing of the audio data file before editing the audio data file. This processing process is mainly used to perform corresponding analysis on the audio data file, so as to determine the part to be edited in the audio data file that needs to be edited, for example, determining the time period that needs to be compressed from the audio data file, or determining part of the audio content that needs to be desensitized from the audio data file, etc.

[0055] It should be noted that the audio data file includes the call recording between the customer and the customer service, and also includes the text file obtained by translating the corresponding call recording. When pre-processing the audio data file, it is necessary to use the audio file to process the audio content corresponding to each node in the audio file, and also use the text file as a reference to determine whether the audio content is sensitive information, or whether the audio content is meaningless silence, etc.

[0056] Among them, preprocessing is equivalent to conducting an overall analysis of the audio data file in advance, and pre-marking and calculating the audio data file, and determining in advance the part of the audio data file to be edited, so as to improve its processing efficiency and accuracy when the audio data file is processed subsequently.

[0057] Step S130: caching the audio files in the audio data file in segments according to the preset audio buffer and the part to be edited, and editing them in sequence;

[0058] The preset audio buffer is a buffer for audio data. Its main function is to cache and read the audio data file to be played, edit the read data accordingly, and write the data to the output stream after processing, so that the audio data file can be played in the form of a file stream according to the output stream.

[0059] In the process of caching the audio file by presetting the audio buffer, the size of the audio file cached each time or the length of the audio file cached each time can be set, for example, caching an audio file of 100 seconds each time.

[0060] Since the size of the audio file cached by the preset audio buffer each time is set, and the actual length of the recorded content is not a fixed value, the length of the call recording between the customer and the customer service staff is different in different situations. Usually, the length of the call recording is greater than the length of the preset cached audio file. At this time, the preset audio buffer cannot cache the audio file through a single caching process, and the preset audio buffer will cache the same audio file multiple times. When the preset audio buffer caches the audio file multiple times and in segments, the node of the cached data each time will be automatically recorded to ensure continuous and segmented caching of the entire audio file.

[0061] Step S140: When the preset audio buffer caches the audio file, the edited data is synchronously written into the output stream, and the corresponding audio content is played in real time according to the output stream.

[0062] When the corresponding segmented audio file is obtained in the cache, the segmented audio needs to be edited and processed. When the segmented data is processed in sequence as mentioned in this embodiment, there is a prerequisite, that is, after the audio file cached in the previous segment is processed and written to the output stream, the preset audio buffer will cache the next segmented audio file and process the next segmented audio file, that is, when the segmented audio file obtained in the previous cache is processed and played in real time through the output stream, the preset audio buffer will cache the next segmented audio file, thereby realizing continuous caching, processing and playback of audio files in the form of file streams.

[0063] Among them, when the corresponding audio file is cached according to the preset audio buffer, the last audio file that has been edited can be synchronously written to the output stream, and the corresponding audio content can be played in real time according to the output stream, that is, the effect of synchronous caching, processing and playback is achieved, avoiding the situation where relevant personnel need to wait for a long time for audio file processing when they want to listen to the recorded content again.

[0064] It should be noted that in the traditional audio-to-text process, the text is desensitized and then converted to audio after the desensitization process. It is necessary to store the corresponding data of the audio and text involved in each step, so as to rewrite or process the corresponding data. However, by presetting the audio buffer, the audio file is cached and processed in segments, and the processed data is played in real time, without occupying additional content, the effect of real-time caching and output is achieved, and the situation where audio files and corresponding text files occupy space is avoided.

[0065] Exemplarily, when the preset audio buffer caches the audio file, the step of synchronously writing the edited data to the output stream includes:

[0066] Step a: When the preset audio buffer caches the audio file, the edited data is synchronously converted into a PCM-encoded WAV format real-time audio stream and written into the output stream.

[0067] When the preset audio buffer caches the audio file, the edited data is synchronously converted into a PCM-encoded WAV format real-time audio stream and written into the output stream, that is, by converting the data format into the WAV format, the required storage space is reduced.

[0068] In the process of caching, processing, and writing the audio file to the output stream by using the preset audio buffer, in order to ensure that the final played audio is consistent with the performance of the original audio file, it is necessary to synchronously obtain the relevant information of the audio file, such as the channel or sampling rate. The audio format is VOX and the sampling rate is f sampling =8000Hz, channel is mono.

[0069] Compared with the related art, in which relevant personnel need to wait for the entire processing flow of the audio data to be completed before playing the audio, that is, the waiting time for playing the audio is long, in the present application, an audio data file to be played is obtained; the audio data file is preprocessed, and based on the result of the preprocessing, the part to be edited in the audio data file is determined; according to the preset audio buffer and the part to be edited, the audio files in the audio data file are cached in segments, and edited in turn; when the audio file is cached in the preset audio buffer, the edited data is synchronously written to the output stream, and the corresponding audio content is played in real time according to the output stream. That is to say, by preprocessing the audio to be played and determining the part to be edited in the audio data file based on the result of the preprocessing, the audio data file can be segmented and cached according to the preset audio buffer, and the edited and processed data can be written to the output stream when the preset audio buffer caches the audio data file, that is, the cached file is processed in a targeted manner, and at the same time, the processed part of the file is played synchronously, so that the relevant personnel do not need to wait for the file processing process, but can directly play the processed part in real time during the audio data file processing process, thereby realizing real-time playback of the audio data file that needs to be processed, thereby reducing the waiting time for re-listening to the recorded content.

[0070] For example, refer to Figure 2 Based on the first embodiment of the real-time audio playback method of the present application, another embodiment is proposed, wherein the step of preprocessing the audio data file comprises:

[0071] Step S210: reading the content of the text file and parsing to obtain a text information list of the audio conversation content;

[0072] When preprocessing an audio data file, the main thing is to perform corresponding analysis on the audio files in the audio data file to determine the part to be edited, and the audio data file includes a text file and an audio file. When preprocessing an audio data file, it is necessary to analyze the audio data file comprehensively based on the text file and the audio file.

[0073] Among them, the text file is converted into a text file based on the audio content of the audio file through the speech-to-text technology, which converts the conversation content between the customer service and the customer in audio format into a text file.

[0074] By reading the contents of the text file and parsing it, a text information list of the corresponding audio conversation content can be obtained. The text information list includes the text information of the conversation content and the time points corresponding to each conversation content in the audio file. For example, at the beginning of the conversation, the customer service will greet the customer through a fixed script. In the middle of the conversation, the customer and the customer service will talk about the corresponding business content. At the end of the conversation, the customer service will say goodbye to the customer through a fixed script. There are certain differences in the content of the conversations at the above different time points. Therefore, when parsing to obtain the text information list, it is also necessary to include the time data corresponding to the time points of the conversation content of the audio file.

[0075] Step S220: Analyze the first audio duration occupied by each word in the audio file and the second audio duration between adjacent words according to the text information list;

[0076] The text information list includes text information (the content of the conversation between the customer and the customer service) and the corresponding time information of the text information in the audio file. For example, the total length of the audio file is 10 minutes. The initial time of the audio file is 0 and its end time is 600 seconds. Within 0-10 seconds, the customer service is greeting the customer. Within 20-300 seconds, the customer is asking the customer service for detailed information of the business. Between 300 and 350 seconds, the customer service handles the corresponding business for the customer through online operations. Between 350 and 580 seconds, the customer asks about another business. Between 580 and 600 seconds, the customer service is saying goodbye to the customer.

[0077] That is, based on the text information list, the speaking intervals in the audio file corresponding to the conversation between the customer and the customer service can be determined. For example, after the customer service explains the specific details of the business, the customer agrees to handle the corresponding business. It takes a certain amount of time for the customer service to handle the business, and no conversation will occur at this time. The call audio and silent audio can be determined based on the existence of corresponding text content (words) in the text information list. Among them, the silent audio is the audio during the waiting time when the customer service and the customer do not have a conversation. The audio has no text content, that is, the corresponding text or words cannot be converted based on this audio segment. If the silent audio occupies too long, it will affect the experience of subsequent relevant personnel when re-listening to the recording, and the silent audio needs to be removed.

[0078] Among them, the overall duration of the audio file can be analyzed based on the first audio duration occupied by each word in the audio file and the second audio duration between adjacent words. The first audio duration corresponds to the total duration of the effective conversation, that is, the total duration of the conversation between the customer and the customer service. The first audio duration needs to be retained, and the second audio duration corresponds to the duration between words. The second audio duration includes the duration of pauses in normal conversations and also includes the duration corresponding to silent audio.

[0079] Step S230: predicting the total duration of the audio after the audio file is edited based on the first audio duration and the second audio duration; wherein the preprocessing includes parsing the text file to predict the total duration of the audio after the editing process.

[0080] By combining the first audio duration and the second audio duration, the total audio duration of the audio file after editing can be predicted in advance, wherein the total duration after the silent audio is removed from the audio file is mainly predicted, so as to determine the audio length of each segment cache of the preset audio buffer based on the total duration, so as to avoid the audio length of each segment cache of the preset audio buffer being too short, which causes the relevant personnel to wait for more time.

[0081] At the same time, the total length of the audio can be used to determine in advance the total length of the audio played by the relevant personnel, so that an audio time bar can be output according to the total length of the audio, and then the relevant personnel can be provided with the function of selecting the corresponding time point on the time bar of the total length of the audio and jumping accordingly.

[0082] In addition, the total audio duration provides a certain basis for the subsequent processing of the audio file. In the total audio duration, the second audio duration occupied by the silent audio and the position of the second audio duration in the total audio duration can be divided out accordingly. That is, the total audio duration is calculated, and all time nodes within the total audio duration are also counted. The time nodes correspond to the dialogue nodes in the call recording.

[0083] Exemplarily, when determining the total duration of the audio, a summation method is adopted to sum the duration of all current words and the interval between words (the duration of the first audio and the duration of the second audio) to obtain: sum , in seconds.

[0084] Among them, the calculated audio byte length is: t sum *f sampling *2.

[0085] Among them, 2 is the final output PCM encoded audio, and each sampling point occupies 16 bits, that is, 2 bytes.

[0086] Exemplarily, the part to be edited includes a first part and a second part, and the step of determining the part to be edited in the audio data file according to the result of the preprocessing includes:

[0087] Step b: determining, according to the text information list, a first part in the audio data file that needs to be desensitized;

[0088] Step c: Determine the second part of the audio data file that needs to be compressed based on the text information list and the total duration of the audio.

[0089] According to the text information list, a part of the audio data in the audio data file that needs to be desensitized is taken as the first part. Further, in combination with the total audio duration, the second part that needs to be compressed in the audio data file can be determined. Both the first part and the second part are parts to be edited. Among them, the first part refers to sensitive information (such as ID numbers or card numbers, etc.). Among them, the second part refers to silent audio (audio with no dialogue content for a long time, or audio corresponding to no word text). For the above two parts, different editing and processing methods need to be used to edit and process them.

[0090] Exemplarily, the step of determining the first part that needs to be desensitized in the audio data file according to the text information list includes:

[0091] Step d: Obtain a semantic information list that matches the content in the text information list;

[0092] When determining the first part that needs to be desensitized in the audio file according to the text information list, since the text content in the text information list is the text content converted from the speech content of the audio file, there are certain conversion errors in this text content, or conversion mistakes, or when converting numbers, the Arabic numerals in the original ID number are converted into Chinese characters. For example, 2 is converted into two. If it is determined whether there is sensitive information in the text information according to "two", it may be due to the fact that its text content is in Chinese characters and the possible digit in the ID number or card number it represents is ignored, resulting in the inability to perform desensitization processing on the sensitive information related to "two". Therefore, when determining the first part to be desensitized according to the text information list, the corresponding semantic information list needs to be obtained first.

[0093] Among them, this semantic information list is mainly used to replace the content representing numbers with Chinese characters when converting different speech to text with the same numbers. For example, the semantic information list includes the situation of converting two to 2. At the same time, it also includes setting the corresponding semantic information list for some special pronunciation numbers. For example, in a phone number, the usual pronunciation of "1" is called "yao". Therefore, in the semantic information list, different Chinese characters need to be corresponding to the number 1, such as "yao" or "yi" or "yi", etc.

[0094] Step e: Analyze the digital word content in the text information list according to the semantic information list, and determine the adjacent word content adjacent to the digital word content;

[0095] Through this semantic information list, some words in the text information list can be replaced with homophones, so as to expand the scope of querying sensitive information from the text information list, so as to ensure that all sensitive information in the text information list is desensitized.

[0096] Among them, mainly based on the relevant content in the above examples, the digital word content in the text information list is analyzed. During the analysis, it is considered to convert part of the digital content expressed in Chinese characters into digital expression again.

[0097] Among them, when analyzing the content of digital words, taking into account that when the card number or ID number is designed in the communication process, the sensitive information should be continuous word content, that is, when explaining the above sensitive information, the number will be said continuously. Therefore, when analyzing the content of digital words, it is necessary to synchronously analyze the adjacent word content adjacent to the digital word content. If the adjacent word content is a number, it can be proved that the digital word content currently being analyzed is sensitive information, and the adjacent word content is also sensitive information.

[0098] Step f: if the content of the adjacent words is a number, then analyzing whether the adjacent words of the adjacent words are numbers, until the analyzed content is non-numeric content;

[0099] Step g: The words that have been determined to be numbers are taken as the first part that needs to be desensitized.

[0100] At the same time, when the content of the adjacent word is a number, it is further analyzed whether the adjacent word of the adjacent word content is a number, and the above process is repeated until the analyzed content is non-numeric content.

[0101] For example, when analyzing the text content of "The card number is 12345678, and the place of origin is Province A", the first number analyzed is 1. At this time, the adjacent words are analyzed, one is "yes" and the other is "2". At this time, it can be determined that 1 and 2 may be sensitive information. Further, the adjacent words "1" and "3" of "2" are analyzed again, and the search is repeated until it is determined that "belongs to" is non-digital content. At this time, the analysis is stopped, and 1-8 are desensitized as sensitive information.

[0102] Among them, when the number is greater than a certain number of digits, it can be treated as sensitive information. For example, if it is greater than 5 digits, it can be treated as sensitive information. If it is less than 5 digits, for example, the number has only 4 digits, such as Room 1900 of a community, this information is used as a contact address and has a low level of privacy, so it can be treated as normal information.

[0103] Exemplarily, the step of predicting the total duration of the audio after the audio file is edited based on the first audio duration and the second audio duration includes:

[0104] Step h: when the duration of the second audio is greater than a preset duration, predicting a compressed audio duration after compressing the duration of the second audio;

[0105] Step i: predicting the total duration of the audio after the audio file is edited based on the duration of the compressed audio and the duration of the first audio.

[0106] When the duration of the second audio is greater than the preset duration, it proves that the silent audio corresponding to the duration of the second audio is too long. If it is not processed, the relevant personnel will need to listen to the silent audio for a longer period of time when re-listening to the call recording, which will affect the re-listening experience of the relevant personnel when re-listening to the call recording. At the same time, customer service will generate some silent audio during the call when handling related business. When re-listening to the audio, it is necessary to determine that the customer service is handling related business for the customer by listening to the silent audio. That is, in this embodiment, the silent audio needs to be processed to reduce its duration, and some silent audio also needs to be retained to reflect the silent audio that may exist when the customer service is handling business during the call between the customer and the customer service.

[0107] Therefore, when the duration of the second audio is longer than the preset duration, it is determined that it needs to be compressed. In the present embodiment, the silent audio is compressed, mainly by cutting part of the silent audio. For example, the duration of the second audio is 15 seconds, and the preset duration is 10 seconds according to the actual setting. At this time, the duration of the second audio is longer than the preset duration, and the duration of the second audio needs to be compressed, and the second audio duration can be cut to an audio of the same length as the preset duration.

[0108] In summary, the compressed audio duration after compressing the second audio duration can be predicted based on the preset duration and the second audio duration, and thus the total audio duration of the audio file after corresponding editing processing can be further predicted based on the compressed audio duration and the first audio duration.

[0109] It should be noted that when predicting the compressed audio length after the second audio length is compressed, it will be marked according to the position of the second audio length in the entire audio file. For example, the second audio length is 15 seconds, and its corresponding silent audio is between 140 seconds and 155 seconds in the audio file. The predicted compressed length is 10 seconds, and the predicted position of the silent audio is between 140 seconds and 150 seconds. It is necessary to compress the original silent audio between 150 seconds and 155 seconds. That is, by pre-searching for the silent audio and marking its position, the audio to be processed can be marked in advance to facilitate subsequent processing.

[0110] Exemplarily, the part to be edited includes a first part and a second part, and the step of caching the audio files in the audio data file in segments according to the preset audio buffer and the part to be edited, and editing them in sequence, includes:

[0111] Step j: according to the preset audio buffer, when caching the audio file in the audio data file in segments, determining the first part and the second part involved in the file of the current cache segment;

[0112] Step k: converting the first part into a sine wave audio with a preset fixed frequency, and cutting the second part to implement editing processing of the segmented cache file.

[0113] When caching audio files in the audio data file according to the preset audio buffer segments, the first part related to sensitive information and the second part involving silent audio will be desensitized and cropped accordingly during the above preprocessing process. According to the above content, when the audio file is preprocessed and the first and second parts are determined, their positions in the audio file will be marked, that is, the time nodes of the first and second parts will be determined during the preprocessing process, so that the preset audio buffer can directly find the corresponding first and second parts, and then quickly complete the processing process.

[0114] Among them, since the preset audio buffer is in the form of segments, the content of the audio file is read in sequence, that is, the content of a segment of the audio file is cached each time, the length of the segment content is fixed, and the starting point and end point of the segment are only related to the fixed cache length when the preset audio buffer is cached each time. Each time the segment is cached, it will not be cached according to the characteristics of a segment of audio. For example, you can refer to Figure 3 ,exist Figure 3 In the preset audio buffer, fixed-length segmented audio is cached from the audio file. The segmented audio includes silent audio, audio to be desensitized, and normal conversation audio. Among them, the silent audio cached on the far left is only the silent audio in the original audio file. If the silent audio is measured according to the cached audio, it will be determined that its length does not exceed the preset length. It can be compressed only according to the cached segmented content. In fact, the total length of the silent audio in the original audio file combined with other adjacent silent audio is greater than the preset duration. Therefore, this part is actually the part that needs to be compressed.

[0115] Therefore, in order to ensure the accuracy of editing and processing through the preset audio buffer, it is necessary to mark the second part (the time node corresponding to the silent audio, and the position and length of the silent audio that needs to be compressed) in advance when calculating the total audio duration in advance through preprocessing.

[0116] When the audio file is edited and processed through the preset audio buffer, the content of the first part is converted into a sine wave audio with a preset fixed frequency, that is, the original normal conversation audio is converted into a beeping sound, so as to avoid the exposure of sensitive information corresponding to the first part.

[0117] When the audio file is edited and processed by using the preset audio buffer, the audio of the second part is compressed and cropped accordingly, thereby achieving compression of the audio file.

[0118] In this embodiment, the content of the text file is read and parsed to obtain a text information list of the audio conversation content; based on the text information list, the first audio duration occupied by each word in the audio file and the second audio duration between each adjacent word are analyzed; based on the first audio duration and the second audio duration, the total audio duration of the audio file after editing is predicted; wherein the preprocessing includes parsing the text file and predicting the total audio duration after editing, that is, by preprocessing the audio data file, the part to be edited is marked in advance, so as to facilitate the accuracy of subsequent editing and improve the efficiency of editing.

[0119] In addition, the present application also provides a real-time audio playback device, the real-time audio playback device comprising:

[0120] An acquisition module, used to acquire the audio data file to be played;

[0121] A determination module, used for preprocessing the audio data file and determining the part to be edited in the audio data file according to the result of the preprocessing;

[0122] A processing module, used for caching the audio files in the audio data file in segments according to the preset audio buffer and the part to be edited, and performing editing processing on the audio files;

[0123] The playing module is used to synchronously write the edited data into the output stream when the preset audio buffer caches the audio file, and play the corresponding audio content in real time according to the output stream.

[0124] Exemplarily, the determination module includes:

[0125] A reading submodule, used for reading the content of the text file and parsing to obtain a text information list of the audio conversation content;

[0126] An analysis submodule, configured to analyze, according to the text information list, the first audio duration occupied by each word in the audio file and the second audio duration between each adjacent word;

[0127] A prediction submodule, configured to predict the total duration of the audio after the audio file is edited based on the first audio duration and the second audio duration; wherein the preprocessing includes parsing the text file to predict the total duration of the audio after the editing process;

[0128] A first determination submodule, configured to determine, according to the text information list, a first part of the audio data file that needs to be desensitized;

[0129] The second determining submodule is used to determine the second part that needs to be compressed in the audio data file according to the text information list and the total audio duration.

[0130] Exemplarily, the first determining submodule includes:

[0131] An acquisition unit, configured to acquire a semantic information list matching the content in the text information list;

[0132] An analyzing unit, configured to analyze the digital word content in the text information list according to the semantic information list, and determine adjacent word content adjacent to the digital word content;

[0133] A judgment unit, configured to analyze whether the adjacent words of the adjacent words are numbers if the adjacent words are numbers, until the analyzed content is non-number content;

[0134] The determination unit is used to take the words determined to be numbers as the first part that needs to be desensitized.

[0135] Exemplarily, the prediction submodule includes:

[0136] A first prediction unit is used to predict a compressed audio duration after compressing the second audio duration when the second audio duration is greater than a preset duration;

[0137] The second prediction unit is used to predict the total audio duration of the audio file after the audio file is edited according to the compressed audio duration and the first audio duration.

[0138] Exemplarily, the processing module includes:

[0139] A third determining submodule is used to determine the first part and the second part involved in the file of the current cache segment when caching the audio file in the audio data file in segments according to the preset audio buffer;

[0140] The processing submodule is used to convert the first part into a sine wave audio with a preset fixed frequency, and cut the second part to implement editing processing of the segmented cache file.

[0141] Exemplarily, the playback module includes:

[0142] The playing submodule is used to convert the edited data into a PCM-encoded WAV format real-time audio stream and write it into the output stream when the preset audio buffer caches the audio file.

[0143] The specific implementation of the real-time audio playback device of the present application is basically the same as the embodiments of the real-time audio playback method described above, and will not be repeated here.

[0144] In addition, the present application also provides a real-time audio playback device. Figure 4 As shown, Figure 4 It is a structural diagram of the hardware operating environment involved in the embodiment of the present application.

[0145] For example, Figure 4 This is a structural diagram of the hardware operating environment of the real-time audio playback device.

[0146] like Figure 4 As shown, the real-time audio playback device may include a processor 401, a communication interface 402, a memory 403 and a communication bus 404, wherein the processor 401, the communication interface 402 and the memory 403 communicate with each other via the communication bus 404, and the memory 403 is used to store computer programs; the processor 401 is used to implement the steps of the real-time audio playback method when executing the program stored in the memory 403.

[0147] The communication bus 404 mentioned in the above audio real-time playback device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus 404 can be divided into an address bus, a data bus, and a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0148] The communication interface 402 is used for communication between the above-mentioned real-time audio playback device and other devices.

[0149] The memory 403 may include a random access memory (RMD) or a non-volatile memory (NM), such as at least one disk memory. Optionally, the memory 403 may also be at least one storage device located away from the processor 401.

[0150] The above-mentioned processor 401 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0151] The specific implementation of the real-time audio playback device of the present application is basically the same as the embodiments of the real-time audio playback method described above, and will not be repeated here.

[0152] In addition, an embodiment of the present application further proposes a computer-readable storage medium, on which a real-time audio playback program is stored. When the real-time audio playback program is executed by a processor, the steps of the real-time audio playback method described above are implemented.

[0153] The specific implementation of the computer-readable storage medium of the present application is basically the same as the embodiments of the above-mentioned real-time audio playback method, and will not be repeated here.

[0154] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or system including the element.

[0155] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0156] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0157] The above are only preferred embodiments of the present application, and are not intended to limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A real-time audio playback method, characterized in that: The real-time audio playback method comprises the following steps: Acquire an audio data file to be played, wherein the audio data file includes an audio file and a text file obtained by text conversion according to the content of the audio file; Preprocessing the audio data file, and determining a portion to be edited in the audio data file according to a result of the preprocessing; The step of preprocessing the audio data file comprises: Read the content of the text file and parse it to obtain a text information list of the audio conversation content; According to the text information list, analyzing the first audio duration occupied by each word in the audio file and the second audio duration between adjacent words; Predicting the total duration of the audio file after the audio file is edited based on the first audio duration and the second audio duration; wherein the preprocessing includes parsing the text file to predict the total duration of the audio file after the audio file is edited; According to the preset audio buffer and the part to be edited, the audio files in the audio data file are cached in segments, and the audio files are edited in sequence; The part to be edited includes a first part and a second part. The steps of caching the audio files in the audio data file in segments according to the preset audio buffer and the part to be edited, and editing them in sequence, include: According to the preset audio buffer, when caching the audio file in the audio data file in segments, determining the first part and the second part involved in the file of the current cache segment; Converting the first part into a sine wave audio with a preset fixed frequency, and cutting the second part to implement editing processing of the segmented cache file; When the preset audio buffer caches the audio file, the edited data is synchronously written into the output stream, and the corresponding audio content is played in real time according to the output stream.

2. The real-time audio playback method according to claim 1, characterized in that: The step of determining the part to be edited in the audio data file according to the result of the preprocessing comprises: Determine, according to the text information list, a first part in the audio data file that needs to be desensitized; A second part that needs to be compressed in the audio data file is determined according to the text information list and the total duration of the audio.

3. The real-time audio playback method according to claim 2, characterized in that: The step of determining the first part that needs to be desensitized in the audio data file according to the text information list includes: Acquire a semantic information list matching the content in the text information list; Analyzing the digital word content in the text information list according to the semantic information list, and determining adjacent word content adjacent to the digital word content; If the adjacent word content is a number, then analyzing whether the adjacent word of the adjacent word content is a number, until the analyzed content is non-numeric content; The words that have been determined to be numbers are the first ones to be desensitized.

4. The real-time audio playback method according to claim 1, characterized in that: The step of predicting the total duration of the audio file after editing the audio file according to the first audio duration and the second audio duration includes: When the duration of the second audio is greater than a preset duration, predicting a compressed audio duration after compressing the duration of the second audio; The total duration of the audio after the audio file is edited is predicted according to the compressed audio duration and the first audio duration.

5. The real-time audio playback method according to claim 1, characterized in that: The step of synchronously writing the edited data into the output stream when the audio file is cached in the preset audio buffer comprises: When the preset audio buffer caches the audio file, the edited data is synchronously converted into a real-time audio stream in a PCM-encoded WAV format and written into an output stream.

6. A real-time audio playback device, characterized in that: The audio real-time playback device comprises: An acquisition module, used for acquiring an audio data file to be played, wherein the audio data file includes an audio file and a text file obtained by text conversion according to the content of the audio file; A determination module, used for preprocessing the audio data file and determining the part to be edited in the audio data file according to the result of the preprocessing; Optionally, the determining module includes: A reading submodule, used for reading the content of the text file and parsing to obtain a text information list of the audio conversation content; An analysis submodule, configured to analyze, according to the text information list, the first audio duration occupied by each word in the audio file and the second audio duration between each adjacent word; A prediction submodule, configured to predict the total duration of the audio after the audio file is edited based on the first audio duration and the second audio duration; wherein the preprocessing includes parsing the text file to predict the total duration of the audio after the editing process; A processing module, used for caching the audio files in the audio data file in segments and performing editing processing on the audio files according to a preset audio buffer and the part to be edited, wherein the part to be edited includes a first part and a second part; Optionally, the processing module includes: A third determining submodule is used to determine the first part and the second part involved in the file of the current cache segment when caching the audio file in the audio data file in segments according to the preset audio buffer; A processing submodule, used for converting the first part into a sine wave audio with a preset fixed frequency, and cutting the second part, so as to implement editing processing of the segmented cache file; The playing module is used to synchronously write the edited data into the output stream when the preset audio buffer caches the audio file, and play the corresponding audio content in real time according to the output stream.

7. A real-time audio playback device, characterized in that: The device comprises: a memory, a processor, and a real-time audio playback program stored in the memory and executable on the processor, wherein the real-time audio playback program is configured to implement the steps of the real-time audio playback method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a real-time audio playback program, and when the real-time audio playback program is executed by the processor, the steps of the real-time audio playback method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Video processing method and equipment

    CN107659538A