Ring-back tone recognition method and device, medium and program product
By adopting a ringback tone recognition model based on long and short-term memory network in the intelligent outgoing call system, and combining with the deep learning deployment framework, real-time recognition and adaptive control of ringback tones are achieved, solving the problem that the reason why the called terminal is hung up in the prior art is not accurately identified, and the outgoing call efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202510692826.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-07-18
AI Technical Summary
The existing intelligent outgoing call system cannot accurately identify the reason for the called terminal's hangup in real time, resulting in the invalid number ringing and being called multiple times, affecting the outgoing call efficiency, especially when dealing with non-human ringback tones such as ringtones.
The ringback tone recognition model based on long and short-term memory network training is adopted, and the ringback tone returned by the called terminal is pushed in real time frame by frame through the open source soft switch of the calling communication terminal for identification. Combined with the deep learning deployment framework optimization model, real-time detection of ringback tones and adaptive control of subsequent out-call processes are achieved.
Real-time accurate recognition of ringback tones during outgoing call is achieved, avoiding repeated invalid outgoing calls, and improving outgoing call efficiency and accuracy, especially when handling non-personal ringback tones such as ringtones, it has a higher recognition accuracy.
Smart Images

Figure CN120343158A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of audio recognition, and in particular to a ringback tone recognition method, device, medium and program product. Background Art
[0002] Outbound calls are widely used as a means of reaching customers. Intelligent outbound call systems reduce the labor cost of outbound calls and improve outbound call efficiency. However, intelligent outbound call systems cannot accurately obtain the reasons for customers' hang-up when they are not connected. They can only obtain the general error code of the SIP protocol layer. They cannot accurately identify whether the customer actively hangs up, no one answers, or is in arrears, turned off, or has no number, resulting in invalid numbers ringing continuously and being dialed multiple times later, which seriously affects the efficiency of outbound calls. Therefore, the industry detects the ringback tone fed back by the called terminal to obtain the reason for the customer's hang-up. Most existing technologies use ASR technology to convert the ringback tone into text after the outbound call ends, and obtain the hang-up reason through keyword matching. It is impossible to achieve real-time ringback tone detection in outbound calls, and the recognition accuracy of non-human voices such as color ringtones is poor; or CNN is used for recognition, but CNN has poor accuracy when processing audio classification tasks. At the same time, the obtained hang-up reason is not fully used in the optimization of the outbound call process. Developing a new real-time ringback tone detection method and using the hang-up reason to optimize the outbound call process has become an urgent research topic for outbound call systems in recent years. Summary of the invention
[0003] The embodiments of the present application provide a ringback tone recognition method, device, medium and program product to record the ringback tone in real time during an outbound call for accurate recognition, and to adaptively control the subsequent outbound call process based on the recognition result.
[0004] According to one aspect of the present application, a ringback tone recognition method is provided, the method comprising:
[0005] During an outbound call from a calling communication terminal, a target ringback tone returned by a called communication terminal is obtained through real-time frame-by-frame push of an open source softswitch in the calling communication terminal;
[0006] Using a ringback tone recognition model to recognize the target ringback tone and determine a ringback tone recognition result; wherein the ringback tone recognition model is obtained based on long short-term memory network training;
[0007] The subsequent outbound calling process of the calling communication terminal is controlled according to the ringback tone recognition result.
[0008] According to one aspect of the present application, a ringback tone recognition device is provided, the device comprising:
[0009] A target ringback tone acquisition module, configured to obtain a target ringback tone returned by a called communication terminal through real-time frame-by-frame push of an open source soft switch in the calling communication terminal during an outgoing call process of the calling communication terminal;
[0010] A ringback tone recognition result determination module, configured to recognize the target ringback tone by using a ringback tone recognition model to determine a ringback tone recognition result; wherein, the ringback tone recognition model is trained based on a long short-term memory network;
[0011] A control module, configured to control a subsequent outgoing call process of the calling communication terminal according to the ringback tone recognition result.
[0012] According to another aspect of the present application, an electronic device is provided, and the electronic device includes:
[0013] At least one processor; and
[0014] A memory communicatively connected to the at least one processor; wherein,
[0015] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the ringback tone recognition method according to any embodiment of the present application.
[0016] According to another aspect of the present application, a computer-readable storage medium is provided, and the computer-readable storage medium stores computer instructions for causing a processor to implement the ringback tone recognition method according to any embodiment of the present application when executed.
[0017] According to another aspect of the present application, a computer program product is provided, and the computer program product includes a computer program, and the computer program implements the ringback tone recognition method according to any embodiment of the present application when executed by a processor.
[0018] The technical solution of the embodiment of the present application can obtain a target ringback tone returned by a called communication terminal through real-time frame-by-frame push of an open source soft switch in the calling communication terminal during an outgoing call process of the calling communication terminal, and can obtain the target ringback tone in real time frame by frame. The target ringback tone is recognized by using a ringback tone recognition model to determine a ringback tone recognition result; thus, the target ringback tone can be recognized more accurately. According to the ringback tone recognition result, the subsequent outgoing call process of the calling communication terminal is controlled, and the subsequent outgoing call process can be adaptively controlled according to the recognition result of the target ringback tone, avoiding waste of resources in the process of repeated and ineffective outgoing calls, and improving the outgoing call efficiency.
[0019] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understandable through the following description of the specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0021] Figure 1 It is a flowchart of a ringback tone recognition method provided for an embodiment of the present application;
[0022] Figure 2 It is a flowchart of a ringback tone recognition method provided for another embodiment of the present application;
[0023] Figure 3 It is a flowchart of a ringback tone recognition method provided for yet another embodiment of the present application;
[0024] Figure 4 It is an architecture diagram of a ringback tone detection device provided for an embodiment of the present application;
[0025] Figure 5 It is a schematic structural diagram of a ringback tone recognition device provided for an embodiment of the present application;
[0026] Figure 6 It is a schematic structural diagram of an electronic device provided for an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] In order to enable those skilled in the art to better understand the solutions of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0028] It should be noted that the terms "first", "second", "third", "fourth", "actual", "preset", etc. in the specification, claims and the above-mentioned drawings of the present application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0029] Figure 1 FIG. is a flowchart of a ringback tone recognition method provided by an embodiment of the present application. The embodiment of the present application is applicable to the situation of recognizing a ringback tone during an automatic outbound call. This method can be executed by a ringback tone recognition device, which can be implemented in the form of hardware and / or software, and the ringback tone recognition device can be configured in an electronic device. As Figure 1 shown, the method includes:
[0030] S110. During the outbound call of the calling communication terminal, obtain the target ringback tone returned by the called communication terminal through real-time frame-by-frame pushing of the open source soft switch in the calling communication terminal.
[0031] Among them, the calling communication terminal can be a communication terminal capable of realizing automatic outbound calls, such as a mobile communication terminal configured with a SIM card. The called communication terminal can be a terminal capable of answering calls to realize communication, such as a mobile communication terminal configured with a SIM card. The outbound call process is the process of the calling communication terminal making a call to other called communication terminals. The open source soft switch is a next-generation voice and multimedia switching platform based on IP technology, which realizes the functions of traditional hardware switches through software. The audio frame is the smallest processing unit of the audio stream, usually containing several milliseconds of sampled data. Each frame contains metadata such as a timestamp and encoding parameters. Each time a frame of the target ringback tone is collected, the target ringback tone is pushed in real time.
[0032] In the embodiment of the present application, during the outgoing call process of the calling communication terminal, the target ringback tone is recorded through the open source soft switch in the calling communication terminal. For each frame of the target ringback tone recorded, the target ringback tone is pushed frame by frame in real time to the ringback tone recognition module to obtain the target ringback tone in real time and perform recognition. Specifically, the device for obtaining the target ringback tone for recognition can be the calling communication terminal or the corresponding target server for recognizing the target ringback tone. By obtaining the target ringback tone pushed frame by frame, the real-time acquisition of the target ringback tone can be realized, the timeliness of the recognition of the target ringback tone is improved, and the subsequent outgoing call process can be controlled in time according to the recognition result of the target ringback tone.
[0033] Specifically, the recording of the target ringback tone can be implemented through the application record_session provided by the open source soft switch. It is necessary to set the recording application record_session in the dial plan. In addition, the open source soft switch provides rich parameters to realize the personalized recording of the ringback tone. By setting the parameter data of the recording application record_session in the dial plan, the directional storage of the recorded audio of the ringback tone can be realized, and the storage path, file name and recording duration of the recorded audio of the ringback tone can be specified.
[0034] S120. Use the ringback tone recognition model to recognize the target ringback tone and determine the ringback tone recognition result; wherein, the ringback tone recognition model is trained based on the long short-term memory network.
[0035] Among them, the ringback tone recognition model is a model trained based on the long short-term memory network. The long short-term memory network is a special recurrent neural network. LSTM dynamically controls the flow of information through three gating structures: the forget gate, the input gate, and the output gate. The gating uses the sigmoid function to generate values between 0 and 1, determining the proportion of information to be retained or discarded. The unique cell state design maintains the long-term transmission of information like a conveyor belt. Compared with the traditional RNN, it can effectively alleviate the problem of gradient disappearance / explosion and is suitable for processing long sequence data. The long short-term memory network can be pre-trained with ringback tone samples to obtain the ringback tone recognition model.
[0036] In the embodiments of the present application, a ringback tone recognition model can be used to recognize the target ringback tone and determine the ringback tone recognition result. Specifically, for each frame of the target ringback tone obtained, the ringback tone recognition model can be used to recognize the obtained target ringback tone, or after the target ringback tone within a preset time period is obtained, the ringback tone recognition model can be used to recognize the target ringback tone within the preset time period. In the embodiments of the present application, the ringback tone recognition result can be in the form of text, such as "the called party is out of service", "the called party is on the phone", etc. By training the ringback tone recognition model based on the long short-term memory network, the recognition accuracy of the ringback tone recognition model can be improved, so as to obtain a more accurate recognition result for the target ringback tone.
[0037] S130. According to the ringback tone recognition result, control the subsequent outbound call process of the calling communication terminal.
[0038] In the embodiments of the present application, the subsequent outbound call process of the calling communication terminal can be controlled according to the ringback tone recognition result, so as to avoid repeated and ineffective outbound call processes and timely and adaptively adjust the subsequent outbound call strategy according to the ringback tone recognition result. For example, when the ringback tone recognition result indicates that the called communication terminal cannot be connected in the short term, the outbound call can be paused first and then made again after a certain interval. When the ringback tone recognition result indicates that the called communication terminal cannot be connected permanently, the call to the called communication terminal can be prohibited permanently.
[0039] The technical solution of the embodiments of the present application, during the outbound call process of the calling communication terminal, the target ringback tone returned by the called communication terminal is obtained through the real-time frame-by-frame push of the open source soft switch in the calling communication terminal, and the target ringback tone can be obtained frame by frame in real time. The ringback tone recognition model is used to recognize the target ringback tone to determine the ringback tone recognition result; thus, the target ringback tone can be recognized more accurately. According to the ringback tone recognition result, the subsequent outbound call process of the calling communication terminal is controlled, and the subsequent outbound call process can be adaptively controlled according to the recognition result of the target ringback tone, avoiding the waste of resources in the process of repeated and ineffective outbound calls and improving the outbound call efficiency.
[0040] Figure 2 It is a flowchart of a ringback tone recognition method provided in another embodiment of the present application. The embodiments of the present application are optimized based on the above embodiments. For the solutions not described in detail in the embodiments of the present application, please refer to the above embodiments. As Figure 2 shown, the method of the embodiments of the present application specifically includes the following steps:
[0041] S210. During the outbound call process of the calling communication terminal, when the open source soft switch of the calling communication terminal records each frame of the target ringback tone, the frame of the target ringback tone is sent in real time based on the pre-set IP address and port number.
[0042] In the embodiment of the present application, the target ringback tone is recorded through the open source soft switch of the calling communication terminal. For each frame of the target ringback tone recorded by the open source soft switch of the calling communication terminal, it is sent in real time based on the pre-set IP address and port number, realizing the real-time sending of each frame of the target ringback tone. Among them, the IP address and port number are devices that need to receive the target ringback tone for processing and are pre-set. It can be the target ringback tone recognition device in the present application. When the target server is used as the target ringback tone recognition device, the IP address and port number are the target server.
[0043] S220. When it is determined that the target ringback tone is received by listening to the IP address and port number, assemble each frame of the target ringback tone in real time to form a target ringback tone in a preset audio format.
[0044] When it is determined that the target ringback tone is received by listening to the IP address and port, assemble each frame of the target ringback tone in real time to form a target ringback tone in a preset audio format. Specifically, the data frames of the target ringback tone can be received based on a UDP receiver, and each data frame is assembled into a wav format audio. Packet loss may occur during the transmission of UDP packets, and silent frames can be added to the frames where the lost packets are located.
[0045] In the embodiment of the present application, assembling each frame of the target ringback tone in real time to form a target ringback tone in a preset audio format includes:
[0046] Check each frame of the received target ringback tone to detect whether a frame loss event occurs;
[0047] If there is a frame loss event, insert silent frames at the frame loss positions to form a complete target ringback tone.
[0048] Exemplarily, check each frame of the received target ringback tone. By performing integrity verification on the target ringback tone, it is detected whether the target ringback tone is complete, and further whether a frame loss event occurs in the target ringback tone. If a frame loss event occurs in the target ringback tone, in order to maintain the integrity of the target ringback tone, silent frames are inserted at the positions of the lost frames, thereby forming a complete target ringback tone, and the target ringback tone is recognized based on the ringback tone recognition model, avoiding the inability to use the ringback tone recognition model for normal recognition due to frame loss of the target ringback tone and format problems.
[0049] S230. Use the ringback tone recognition model to recognize the target ringback tone and determine the ringback tone recognition result; among them, the ringback tone recognition model is trained based on a long short-term memory network.
[0050] Exemplarily, detect the target ringback tone in the above-mentioned preset audio format and output the ringback tone recognition result.
[0051] S240. Control the subsequent outgoing call process of the calling communication terminal according to the recognition result of the ringback tone.
[0052] The embodiment of the present application provides a method for recognizing a ringback tone. Each time a frame of the target ringback tone is recorded by the open source soft switch of the calling communication terminal, the frame of the target ringback tone is sent in real time based on the pre-set IP address and port number; when it is determined that the target ringback tone is received by listening to the IP address and port number, each frame of the target ringback tone is assembled in real time to form a target ringback tone in a preset audio format. The above solution can send the target ringback tone in real time based on the pre-set IP address and port number each time a frame of the target ringback tone is recorded, realizing real-time transmission, and performing real-time processing after real-time transmission, improving the timeliness of recognizing the target ringback tone, and timely determining the subsequent outgoing call control strategy according to the recognition result of the target ringback tone.
[0053] Figure 3 The flowchart of a method for recognizing a ringback tone provided by another embodiment of the present application. The embodiment of the present application is optimized based on the above embodiment, and the solutions not described in detail in the embodiment of the present application can be seen in the above embodiment. As Figure 3 shown, the method of the embodiment of the present application specifically includes the following steps:
[0054] S310. Extract the feature information of the sample ringback tone by using a long short-term memory network.
[0055] Among them, the sample ringback tone is obtained in advance during the outgoing call process of the calling communication terminal. Specifically, the calling communication terminal can be used to make an outgoing call in advance, and during the outgoing call process, the open source soft switch records and stores the ringback tone, and the ringback tones that meet the model training quantity are used as the sample ringback tone. The long short-term memory network is used to extract the feature information of the sample ringback tone. The long short-term memory network introduces a cell state and a gating mechanism, effectively alleviating the problem of gradient disappearance, and can retain the context information in a long sequence when processing an audio sequence, and can extract more representative features.
[0056] In the embodiment of the present application, the determination process of the sample ringback tone includes:
[0057] Obtain the ringback tone transmitted via the gateway through the open source soft switch in the calling communication terminal, and record the ringback tone through the recording application in the open source soft switch, and store the recorded multiple ringback tones in a targeted manner;
[0058] The stored ringback tone is recognized based on an automatic speech recognition interface to obtain a ringback tone text, and the ringback tone text is keyword-matched with a target database to determine a label; or a label formulated by a user for the stored ringback tone is obtained;
[0059] The stored ringback tone and the corresponding label are combined to form a sample ringback tone.
[0060] Exemplarily, during the process of an outgoing call from a calling communication terminal, the ringback tone transmitted via a gateway is obtained through an open-source switch, and the ringback tone is recorded through a recording application in the open-source soft switch. Multiple recorded ringback tones are directionally stored to a specified storage path. The stored ringback tone is recognized based on an automatic speech recognition interface to obtain a ringback tone text, and the ringback tone text is keyword-matched with a target database. The text in the target database that matches the ringback tone text is used as a label to obtain the label corresponding to the sample ringback tone. Or a label formulated by a user for the stored ringback tone is obtained, that is, the user listens to the stored ringback tone and inputs the corresponding text of the ringback tone as a label. The stored ringback tone and the corresponding label are combined to form a sample ringback tone, and thus model training is performed based on the sample ringback tone.
[0061] S320. A fully connected layer is used to integrate the sample ringback tone feature information, and a classification result of the sample ringback tone is obtained through a classification output layer.
[0062] Among them, the fully connected layer is a basic layer structure in a neural network, mainly used to integrate and transform feature information. The local features are synthesized into global features through the fully connected layer to improve the robustness of the model. The fully connected layer is used to integrate the sample ringback tone feature information to form the global feature of the sample ringback tone. The integrated sample ringback tone passes through the classification output layer to obtain the classification result of the sample ringback tone and the predicted classification of the sample ringback tone.
[0063] S330. According to the classification result of the sample ringback tone and the label, a loss function is determined for model training optimization to obtain a ringback tone recognition model.
[0064] Among them, the label is the true classification of the sample ringback tone, and the classification result of the ringback tone recognition model for the sample ringback tone is the predicted classification. The model can be evaluated based on the consistency between the true classification and the predicted classification to determine the accuracy of the model recognition. Specifically, according to the classification result of the sample ringback tone and the label, a loss function is determined for model training optimization until the optimization condition is met to obtain the ringback tone recognition model.
[0065] In the embodiment of the present application, the method further includes:
[0066] Convert the trained and optimized model into a format supported by the deep learning deployment framework to obtain a standard format ringback tone recognition model;
[0067] Performing at least one of layer fusion, precision adjustment, and kernel automatic adjustment on the ringback tone recognition model to optimize the ringback tone recognition model;
[0068] The optimized ringback tone recognition model is deployed to the production environment using a deep learning deployment framework, so that the ringback tone recognition model is called in the production environment to recognize the target ringback tone.
[0069] Convert the trained model into a format supported by the deep learning deployment framework, usually the ONNX standard model format, so that models trained with different deep learning frameworks are compatible with the deep learning deployment framework. After the model is converted, use the deep learning deployment framework to optimize the model, perform layer fusion, precision adjustment, and kernel automatic adjustment on the model, and serialize the model into a high-performance inference engine to accelerate the model's inference process. Use the deep learning deployment framework to deploy the inference engine to the production environment to maximize throughput and improve the model's inference speed and performance. After the model is deployed, use the HTTP / REST service of the deep learning deployment framework to receive and process inference requests. Manage the model, such as monitoring model performance indicators and limiting the model's inference rate. Perform version control on the model, and create a data set to train the model again for newly added types of ringback tones or ringback tone samples whose detection accuracy needs to be improved, complete the model upgrade, and continuously improve the model's recognition accuracy.
[0070] S340: During an outbound call by the calling communication terminal, a target ringback tone returned by the called communication terminal is acquired through real-time frame-by-frame push of an open source softswitch in the calling communication terminal.
[0071] S350. Use a ringback tone recognition model to recognize the target ringback tone and determine a ringback tone recognition result; wherein the ringback tone recognition model is obtained based on long short-term memory network training.
[0072] S360: Control a subsequent outbound calling process of the calling communication terminal according to the ringback tone recognition result.
[0073] In the embodiment of the present application, according to the ringback tone recognition result, controlling the subsequent outbound calling process of the calling communication terminal includes:
[0074] If the ringback tone recognition result reflects that the other party is on a call or has turned off the phone, a hang-up signaling is sent to terminate the outbound call process;
[0075] If the interval time reaches the preset interval time, the outbound call will be made again;
[0076] If the ringback tone recognition result indicates that the called number is an invalid number, send a hang-up command to terminate the outbound call process and prohibit outbound calls to this number.
[0077] Control the outbound call process according to the ringback tone recognition result. For example, if the ringback tone detection result indicates that the called party is on the phone or has powered off, send a hang-up signal to terminate the outbound call process in a timely manner and re-call after a preset interval. If the detection result indicates that the called number is an invalid number, send a hang-up signal to terminate the outbound call process in a timely manner and prohibit calls to this number, reducing the occupancy of outbound call resources and improving the outbound call efficiency.
[0078] The embodiment of the present application provides a ringback tone recognition method, which applies a long short-term memory network to the ringback tone detection task. Compared with the ringback tone detection method based on automatic speech recognition, the present invention can detect the ringback tone in real time during the outbound call process and has a higher recognition accuracy for non-human voice ringback tones such as ringback tones. Compared with the ringback tone detection method based on a convolutional neural network, the present invention uses a long short-term memory network for ringback tone detection. The long short-term memory network can better process audio classification tasks than the convolutional neural network and obtain higher accuracy. Optimize the model using a deep learning deployment framework to obtain a high-performance inference engine and deploy it to the server to open the inference service, improving the efficiency of the ringback tone detection task.
[0079] The present application provides a specific implementation method. The architecture diagram of the ringback tone detection device provided by the embodiment of the present application is as Figure 4 shown. Specifically, it includes:
[0080] 1. Ringback tone acquisition unit: mainly used to receive, record, store, and process data of the ringback tone transmitted back by the called terminal, and is composed of a ringback tone receiver, a ringback tone recorder, a data memory, and a data processor.
[0081] Among them, the ringback tone receiver is used to receive the ringback tone transmitted back by the called terminal, the ringback tone recorder is used to record the ringback tone personalized, the data memory is used to store the ringback tone directionally, and the data processor tags according to the ringback tone audio content.
[0082] The specific processing flow of the ringback tone acquisition unit includes:
[0083] Step 1, ringback tone reception. The ringback tone audio stream is transmitted to the gateway through the operator network and then to the open source soft switch through the gateway. The open source soft switch receives the ringback tone audio stream.
[0084] Step 2, ringback tone recording. The ringback tone recording is implemented through the application record_session provided by the open source soft switch. It is necessary to set the recording application record_session in the dialing plan. In addition, the open source soft switch provides rich parameters to achieve personalized recording of the ringback tone.
[0085] Step 3: Data storage: By setting the parameter data of the recording application record_session in the dial plan, the directional storage of the ringback tone recording audio can be realized, and the storage path, file name and recording duration of the ringback tone recording audio can be specified.
[0086] Step 4: Data processing: Classify the recorded ringback tones according to the audio content to form model training samples. Call the ASR interface to convert the unanswered ringback tone audio under the specified path into text, and then match the text with the target database for keywords. If the match is successful, the ringback tone recording audio is tagged. If the match fails, manual tagging is performed.
[0087] 2. Model training unit: mainly used for training the ringback tone detection model, consisting of a data preprocessor, a feature extractor, a feature integration classifier and a parameter adjustment operator.
[0088] Among them, the data preprocessor is used to preprocess the data before model training, the feature extractor is used to extract features from samples, the feature integration classifier is used to classify the extracted features, and the parameter adjustment operator is used to adjust the model parameters.
[0089] The specific processing flow of the model training unit includes:
[0090] Step 1: Data preprocessing before model training. First, normalize the audio file to the specified length, and truncate the overlong files; for files that are not long enough, use the maximum likelihood estimation method to fill the files to the specified length. Secondly, normalize the audio stream to reduce the impact of sample magnitude differences on the training process;
[0091] Step 2: Extract features from samples. LSTM is used for feature extraction. LSTM introduces cell state and gating mechanism, which effectively alleviates the gradient vanishing problem. It can retain distant context information when processing audio sequences and can better extract features.
[0092] Step 3: Classify the extracted features. Use the fully connected layer to integrate the features extracted by LSTM, and use the SoftMax layer to obtain the probability distribution of the ringback tone category;
[0093] Step 4: Adjust model parameters. Use the cross entropy loss function to calculate the loss between the model prediction value and the label, and back propagate to adjust the model parameters.
[0094] Steps 2 to 4 are used to update the model once, and steps 2 to 4 are repeated to iteratively optimize the model to complete the training process.
[0095] 3. Model inference unit: mainly used to deploy the trained ringback tone detection model and open services for inference calls, consisting of a model converter, a model optimizer, a model deployer, and a model management and server. Among them, the model converter is used to convert the model into an intermediate format, the model optimizer is used to optimize the trained model, the model deployer is used to deploy the model, and the model management and server is used to manage the model and open services for inference calls.
[0096] The specific processing flow of the model inference unit includes:
[0097] Step 1: Model conversion. Convert the trained model into a format supported by the deep learning deployment framework, usually the ONNX standard model format, so that models trained with different deep learning frameworks are compatible with the deep learning deployment framework.
[0098] Step 2: Model optimization. After the model is converted, the deep learning deployment framework is used to optimize the model, perform layer fusion, precision adjustment, and kernel automatic adjustment on the model, and serialize the model into a high-performance inference engine to accelerate the model's inference process.
[0099] Step 3: Model deployment: Using a deep learning deployment framework to deploy the inference engine to a production environment can maximize throughput and improve the model’s inference speed and performance.
[0100] Step 4: Model management and service. After the model is deployed, use the HTTP / REST service of the deep learning deployment framework to receive and process inference requests. Manage the model, such as monitoring model performance indicators and limiting the inference rate of the model. Perform version control on the model, create data sets for training the model again for new types of ringback tones or ringback tone samples with improved detection accuracy, complete model upgrades, and continuously improve the recognition accuracy of the model.
[0101] 4. Ringback tone recognition and outbound call optimization unit: It is mainly used to detect the ringback tone and control the outbound call process in real time according to the detection results. It consists of a ringback tone stream sender, a UDP stream receiver, a model caller and an outbound call controller.
[0102] Among them, the ringback tone sender is used to send the recorded ringback tone to the target server, the UDP receiver is used to receive the ringback tone data frame and assemble it into wav format audio, the model caller is used to call the inference model to detect the ringback tone, and the outbound call controller is used to optimize the outbound call process according to the ringback tone detection results.
[0103] The specific processing flow of the ringback tone recognition and outbound call optimization unit includes:
[0104] Step 1, ringback tone streaming. Embed a streaming module in the open-source soft switch. The streaming module pushes the ringback tone audio frame by frame, and sets "execute_on_answer = rstIP address port number" in the dialing plan to specify the IP address and port number of the target server.
[0105] Step 2, UDP stream reception. Listen to the port specified by the ringback tone streamer, and assemble the received data frames into WAV format audio. Packet loss may occur during the transmission of UDP packets. Fill in silent frames for the frames where the lost packets are located.
[0106] Step 3, model invocation. Invoke the model inference unit service to detect the WAV format ringback tone audio and output the detection result.
[0107] Step 4, outbound call process control. Control the outbound call process according to the ringback tone detection result. For example, if the ringback tone detection result is that the other party is on the phone or the other party has shut down, send a hang-up signal to terminate the outbound call process in time, and make a re-call after an interval; if the detection result is that the other party's number is invalid, send a hang-up signal to terminate the outbound call process in time and block the call to this number, reducing the occupancy of outbound call resources and improving the outbound call efficiency, etc.
[0108] Figure 5 FIG. is a schematic structural diagram of a ringback tone recognition device provided by an embodiment of the present application. This device can execute the ringback tone recognition method provided by any embodiment of the present application, and has corresponding functional modules and beneficial effects for executing the method. As Figure 5 shown, the device includes:
[0109] A target ringback tone acquisition module 410, configured to obtain a target ringback tone returned by a called communication terminal through real-time frame-by-frame push of an open-source soft switch in the calling communication terminal during an outbound call process of the calling communication terminal;
[0110] A ringback tone recognition result determination module 420, configured to recognize the target ringback tone by using a ringback tone recognition model and determine a ringback tone recognition result; wherein, the ringback tone recognition model is trained based on a long short-term memory network;
[0111] A control module 430, configured to control the subsequent outbound call process of the calling communication terminal according to the ringback tone recognition result.
[0112] In an embodiment of the present application, the target ringback tone acquisition module 410 obtains the target ringback tone returned by the called communication terminal through real-time frame-by-frame push of the open-source soft switch in the calling communication terminal, including:
[0113] When each frame of the target ringback tone is recorded in the open-source soft switch of the calling communication terminal, the frame of the target ringback tone is sent in real time based on a pre-set IP address and port number;
[0114] When the monitored IP address and port number determine that the target ringback tone is received, each frame of the target ringback tone is assembled in real time to form a target ringback tone in a preset audio format.
[0115] In the embodiment of the present application, the target ringback tone acquisition module 410 assembles each frame of the target ringback tone in real time to form a target ringback tone in a preset audio format, including:
[0116] Check each frame of the received target ringback tone to detect whether a frame loss event occurs;
[0117] If there is a frame loss event, insert a silent frame at the frame loss position to form a complete target ringback tone.
[0118] In the embodiment of the present application, the device further includes:
[0119] A sample ringback tone information determination module, configured to extract features of the sample ringback tone by using a long short-term memory network to obtain sample ringback tone feature information;
[0120] A classification result determination module, configured to integrate the sample ringback tone feature information by using a fully connected layer and obtain a classification result of the sample ringback tone through a classification output layer;
[0121] A ringback tone recognition model determination module, configured to perform model training optimization according to the classification result of the sample ringback tone and a label to determine a loss function, and obtain a ringback tone recognition model.
[0122] In the embodiment of the present application, the device further includes:
[0123] A directional storage module, configured to obtain a ringback tone transmitted via a gateway through an open source softswitch in the calling communication terminal, record the ringback tone through a recording application in the open source softswitch, and perform directional storage on multiple recorded ringback tones;
[0124] A label determination module, configured to identify a ringback tone text based on an automatic speech recognition interface for the stored ringback tone, perform keyword matching on the ringback tone text with a target database to determine a label; or obtain a label formulated by a user for the stored ringback tone;
[0125] A sample ringback tone determination module, configured to combine the stored ringback tone and the corresponding label to form a sample ringback tone.
[0126] In the embodiment of the present application, the device further includes:
[0127] A format conversion module, configured to convert the trained and optimized model into a format supported by a deep learning deployment framework to obtain a ringback tone recognition model in a standard format;
[0128] An optimization module for performing at least one of layer fusion, accuracy adjustment, and kernel automatic adjustment on the ringback tone recognition model to optimize the ringback tone recognition model.
[0129] A deployment module for deploying the optimized ringback tone recognition model to a production environment using a deep learning deployment framework to call the ringback tone recognition model in the production environment to recognize a target ringback tone.
[0130] In an embodiment of the present application, the control module 430 controls the subsequent outbound call process of the calling communication terminal according to the ringback tone recognition result, including:
[0131] If the ringback tone recognition result indicates that the other party is on a call or the other party has powered off, send a hang-up signaling to terminate the outbound call process;
[0132] If the interval time reaches a preset interval time, make a new outbound call;
[0133] If the ringback tone recognition result indicates that the other party's number is invalid, send a hang-up instruction to terminate the outbound call process and prohibit making an outbound call to this number.
[0134] A ringback tone recognition device provided in an embodiment of the present application can execute a ringback tone recognition method provided in any embodiment of the present application, and has corresponding functional modules and beneficial effects for executing the method.
[0135] Figure 6 The structural schematic diagram of an electronic device 10 that can be used to implement the embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present application described and / or claimed herein.
[0136] Such as Figure 6As shown, the electronic device 10 includes at least one processor 11 and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0137] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless ringback tone identification transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0138] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the ringback tone identification method.
[0139] In some embodiments, the ringback tone identification method can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the ringback tone identification method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the ringback tone identification method by any other suitable means (e.g., by means of firmware).
[0140] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0141] The computer programs for implementing the methods of this application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable ringing tone recognition device, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.
[0142] In the context of this application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0143] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0144] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0145] The computing system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0146] An embodiment of the present invention also provides a computer program product, including a computer program which, when executed by a processor, implements the ringback tone recognition method provided in any embodiment of the present application.
[0147] In the process of implementing the computer program product, computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0148] It should be understood that the various forms of the process shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this application may be executed in parallel, sequentially, or in a different order, as long as the information desired by the technical solution of this application can be achieved. This is not limited herein.
[0149] The above specific embodiments do not constitute a limitation on the protection scope of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the protection scope of this application.
Claims
1. A ringback tone recognition method, characterized in that, The method includes: During the outgoing call process of the calling communication terminal, obtaining the target ringback tone returned by the called communication terminal through real-time frame-by-frame push of the open source soft switch in the calling communication terminal; Using a ringback tone recognition model to recognize the target ringback tone and determining the ringback tone recognition result; wherein, the ringback tone recognition model is trained based on a long short-term memory network; Controlling the subsequent outgoing call process of the calling communication terminal according to the ringback tone recognition result.
2. The method according to claim 1, characterized in that, Obtaining the target ringback tone returned by the called communication terminal through real-time frame-by-frame push of the open source soft switch in the calling communication terminal includes: When each frame of the target ringback tone is recorded by the open source soft switch of the calling communication terminal, sending the frame of the target ringback tone in real time based on the pre-set IP address and port number; When it is determined that the target ringback tone is received by listening to the IP address and port number, assembling each frame of the target ringback tone in real time to form a target ringback tone in a preset audio format.
3. The method according to claim 2, characterized in that, Assembling each frame of the target ringback tone in real time to form a target ringback tone in a preset audio format includes: Checking each frame of the received target ringback tone to detect whether a frame loss event occurs; If there is a frame loss event, inserting a silent frame at the frame loss position to form a complete target ringback tone.
4. The method according to claim 1, wherein The construction process of the ringback tone recognition model includes: Using a long short-term memory network to extract features from the sample ringback tone to obtain sample ringback tone feature information; Using a fully connected layer to integrate the sample ringback tone feature information and obtaining the classification result of the sample ringback tone through a classification output layer; Determining a loss function according to the classification result of the sample ringback tone and the label for model training optimization to obtain a ringback tone recognition model.
5. The method according to claim 4, wherein The determination process of the sample ringback tone includes: Obtaining the ringback tone transmitted via the gateway through the open source soft switch in the calling communication terminal, recording the ringback tone through the recording application in the open source soft switch, and storing the recorded multiple ringback tones in a targeted manner; Recognizing the stored ringback tone based on an automatic speech recognition interface to obtain a ringback tone text, and performing keyword matching on the ringback tone text with a target database to determine the label; or obtaining the label specified by the user for the stored ringback tone; Combining the stored ringback tone and the corresponding label to form a sample ringback tone.
6. The method according to claim 4, characterized in that, The method further includes: Converting the trained and optimized model into a format supported by a deep learning deployment framework to obtain a ringback tone recognition model in a standard format; Performing at least one of layer fusion, accuracy adjustment, and kernel automatic adjustment on the ringback tone recognition model to optimize the ringback tone recognition model; Deploying the optimized ringback tone recognition model to a production environment using a deep learning deployment framework to call the ringback tone recognition model in the production environment to recognize the target ringback tone.
7. The method according to claim 1, wherein Controlling the subsequent outgoing call process of the calling communication terminal according to the ringback tone recognition result includes: If the ringback tone recognition result indicates that the other party is on a call or the other party has shut down, sending a hang-up signaling to terminate the outgoing call process; If the interval time reaches a preset interval time, making an outgoing call again. If the ringback tone recognition result indicates that the called party is an invalid number, send a hang-up command to terminate the outbound call process and prohibit outbound calls to this number.
8. An electronic device, characterized in that, The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the ringback tone recognition method according to any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a processor to execute the ringback tone recognition method according to any one of claims 1-7 when executed.
10. A computer program product, characterized in that, It includes a computer program which, when executed by a processor, implements the ringback tone recognition method according to any one of claims 1-7.