Call character conversion method and device, vehicle and storage medium
By converting mobile phone call recordings into text data and storing them in a classified manner, the problem of low efficiency in finding call content in the prior art is solved, and the effect of fast query and efficient management is achieved.
Patent Information
- Application Number
- CN202510513205.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-25
AI Technical Summary
In the prior art, mobile phone call recording is audio data, resulting in low efficiency in finding call content and inability to quickly query and manage.
By recording and converting audio files into text data in response to the call-on operation, the audio text conversion model of local and servers is processed, and the attribute information of the text data is classified and stored.
It realizes quick query of call content, improves the efficiency and accuracy of finding call content, and supports local and server data synchronization management.
Smart Images

Figure CN120378525A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent technologies, and in particular, to a call text conversion method, apparatus, vehicle, and storage medium. Background Art
[0002] With the development of communication technologies, the popularity of smart phones is getting higher and higher, and the demand for smart phone calls will also be higher and higher. Currently, all mobile phone calls are audio data, and call recordings also save the call audio data. If you want to review it again, you need to play the audio. If there is no recording, you can't find the previous records. Therefore, the efficiency of finding call content is low. Summary of the Invention
[0003] Embodiments of this application provide a call text conversion method, apparatus, vehicle, and storage medium, which can effectively improve the efficiency of finding call content.
[0004] The technical solution of this application is implemented as follows:
[0005] Embodiments of this application provide a call text conversion method, which includes:
[0006] In response to an operation of answering a call, perform call recording to determine an audio file of the call process;
[0007] Based on the audio file, perform conversion processing through a preset audio-to-text conversion model to obtain text data corresponding to the audio file; wherein, the preset audio-to-text conversion model is used to convert an audio file into text data;
[0008] Classify based on the attribute information of the text data to obtain a classification result; and store the text data based on the classification result.
[0009] It can be understood that, in response to an operation of answering a call, performing voice recording on all calls can obtain audio files of all calls. By converting the audio files into text data and classifying and storing the text data; since the audio files are converted into text data, there is no need to listen to the audio files during subsequent queries, but directly query the text data, which can improve the query speed. At the same time, querying the text content of the call according to the classification result can further improve the efficiency of finding call content.
[0010] In the above solution, the preset audio-to-text conversion model includes a first audio-to-text conversion model and a second audio-to-text conversion model;
[0011] The step of, based on the audio file, performing conversion processing through a preset audio-to-text conversion model to obtain text data corresponding to the audio file, includes:
[0012] Perform local conversion processing on the audio file through the first audio-to-text conversion model to obtain the local text data corresponding to the audio file;
[0013] Upload the audio file to the server;
[0014] Receive the server text data corresponding to the audio file sent by the server; wherein, the server text data is obtained by the server through converting the audio file by the second audio-to-text conversion model;
[0015] Based on the local text data and the server text data, determine the text data corresponding to the audio file.
[0016] It can be understood that by converting the audio file locally through the first audio-to-text conversion model to obtain local text data, converting the audio file on the server through the second audio-to-text conversion model to obtain server text data, and then screening the local text data and the server text data to determine the text data corresponding to the audio file, the accuracy of the text data can be improved through different conversion methods and screening.
[0017] In the above solution, after uploading the audio file to the server, the method further includes:
[0018] In the case of not receiving the server text data sent by the server, determine the local text data as the text data corresponding to the audio file.
[0019] It can be understood that in the case of not receiving the server text data sent by the server, determining the local text data as the text data corresponding to the audio file can ensure obtaining the text data corresponding to the audio file.
[0020] In the above solution, the determining the text data corresponding to the audio file based on the local text data and the server text data includes:
[0021] Based on the local text data and the server text data, perform error comparison through the audio file to determine the error value between the local text data and the server text data;
[0022] Based on the error value between the local text data and the server text data, determine the text data corresponding to the audio file.
[0023] It can be understood that by performing error comparison between the local text data and the server text data to realize the screening of the text data, the accuracy of the text data can be improved.
[0024] In the above solution, the attribute information includes the text data name and the text data time;
[0025] Classifying based on the attribute information of the text data to obtain a classification result, including:
[0026] Classifying the text data based on the text data name and / or the text data time to obtain a classification result.
[0027] It can be understood that classifying the text data by the text data name and the text data time facilitates subsequent searching for call content.
[0028] In the above solution, the classifying the text data based on the text data name and / or the text data time to obtain a classification result includes:
[0029] Classifying the text data based on the text data name to obtain a classification result; or,
[0030] Classifying the text data based on the text data time to obtain a classification result; or,
[0031] Classifying the text data based on the text data name and the text data time to obtain a classification result.
[0032] It can be understood that classifying the text data by the text data name and the text data time facilitates subsequent classified storage of the text data and searching for call content.
[0033] In the above solution, storing the text data based on the classification result includes:
[0034] Storing the text data locally based on the classification result;
[0035] Sending the classification result and the text data to the server so that the server stores the text data based on the classification result.
[0036] It can be understood that storing the text data locally and on the server facilitates subsequent obtaining of the text data from the server and synchronously obtaining the text data when the terminal device is replaced.
[0037] An embodiment of the present application provides a call text conversion device, including a determination unit, a conversion unit, a classification unit, and a storage unit; wherein,
[0038] The determination unit is configured to perform call recording in response to an operation of answering a call and determine an audio file of the call process;
[0039] The conversion unit is configured to perform conversion processing on the basis of the audio file through a preset audio-to-text conversion model to obtain text data corresponding to the audio file; wherein, the preset audio-to-text conversion model is used to convert an audio file into text data.
[0040] The classification unit is configured to classify based on the attribute information of the text data to obtain a classification result.
[0041] The storage unit is configured to store the text data based on the classification result.
[0042] An embodiment of the present application provides a vehicle, including:
[0043] A memory for storing executable data instructions;
[0044] A processor, when executing the executable instructions stored in the memory, implements the call text conversion method described above.
[0045] An embodiment of the present application provides a computer-readable storage medium storing executable instructions, which, when executed by a processor, implement the call text conversion method described above.
[0046] An embodiment of the present application provides a call text conversion method, device, vehicle, and storage medium. The call text conversion method includes: in response to an operation of answering a call, performing call recording to determine an audio file of the call process; performing conversion processing on the basis of the audio file through a preset audio-to-text conversion model to obtain text data corresponding to the audio file; wherein, the preset audio-to-text conversion model is used to convert an audio file into text data; classifying based on the attribute information of the text data to obtain a classification result; and storing the text data based on the classification result. By adopting the above solution, on the one hand, in response to an operation of answering a call, voice recording is performed on all calls, so that audio files of all calls can be obtained. By converting the audio files into text data and classifying and storing the text data, since the audio files are converted into text data, when querying later, there is no need to listen to the audio files, but directly query the text data, which can improve the query speed. On the other hand, querying the text content of a call according to the classification result can further improve the efficiency of finding the call content. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 An optional flowchart of a call text conversion method provided by an embodiment of the present application Figure 1 ;
[0048] Figure 2 An optional flowchart of a call text conversion method provided by an embodiment of the present application Figure 2 ;
[0049] Figure 3 An optional process schematic diagram of a call text conversion method provided by an embodiment of the present application Figure 3 ;
[0050] Figure 4 An optional framework schematic diagram of a terminal device provided by an embodiment of the present application;
[0051] Figure 5 An optional process schematic diagram of a call text conversion method provided by an embodiment of the present application Figure 4 ;
[0052] Figure 6 A structural schematic diagram of a call text conversion device provided by an embodiment of the present application;
[0053] Figure 7 A structural schematic diagram of a terminal device provided by an embodiment of the present application. Detailed implementation manners
[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will further describe the specific technical solutions of the present application in detail with reference to the accompanying drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application, but are not intended to limit the scope of the present application.
[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0056] In the following descriptions, references to "some embodiments", "this embodiment", "embodiments of the present application", and examples, etc., describe subsets of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0057] If similar descriptions such as "first / second" appear in the application documents, the following explanations are added. In the following descriptions, the terms "first\second\third" are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0058] An embodiment of the present application provides a call text conversion method, Figure 1 An optional process schematic diagram of a call text conversion method provided by an embodiment of the present application Figure 1 , which will be combined withFigure 1 The steps shown will be described.
[0059] S101. In response to an operation of answering a call, perform call recording to determine an audio file of the call process.
[0060] In an embodiment of the present application, the operation of answering a call refers to that after the terminal device receives an incoming call from another device, the call is answered by touching the screen or keys of the terminal device to implement communication between the terminal device and the other device.
[0061] In some embodiments of the present application, the execution subject of the call text conversion method is the terminal device.
[0062] In some embodiments of the present application, the call text conversion method is applicable to scenarios under communication recording.
[0063] In some embodiments of the present application, when the terminal device starts a call with another device, the call can be recorded by enabling the automatic recording function of the terminal device or by receiving a recording instruction to obtain an audio file during the call process.
[0064] It should be noted that the present application records each call to obtain all call recordings of the terminal device.
[0065] In some embodiments of the present application, the terminal device performs call recording in response to an operation of answering a call; and stops recording in response to an operation of ending the call, so as to obtain an audio file during the call process.
[0066] S102. Based on the audio file, perform conversion processing through a preset audio-to-text conversion model to obtain text data corresponding to the audio file; wherein, the preset audio-to-text conversion model is used to convert the audio file into text data.
[0067] In an embodiment of the present application, the preset audio-to-text conversion model includes a first audio-to-text conversion model and a second audio-to-text conversion model. The first audio-to-text conversion model is an audio-to-text conversion model deployed locally; the second audio-to-text conversion model is an audio-to-text conversion model deployed on the server.
[0068] It should be noted that the first audio-to-text conversion model is different from the first audio-to-text conversion model.
[0069] In some embodiments of the present application, the preset audio-to-text conversion model converts the speech signal into readable text through the collaborative work of acoustic modeling and language modeling.
[0070] In some embodiments of the present application, the process of converting an audio file into text by a preset audio-to-text conversion model includes: audio preprocessing, feature extraction signal processing, an acoustic model, and a language model; among them, the audio preprocessing and feature extraction signal processing include: signal processing and feature extraction. The signal processing includes: audio framing, denoising, and enhancement; where
[0071] Audio framing: The continuous speech is segmented into short-time segments (frames) of 20 - 40 ms, and the frames overlap by 10 - 20 ms through a sliding window, which is convenient for capturing the dynamic features of speech.
[0072] Denoising and enhancement: Eliminate environmental noise through techniques such as high-pass filtering and spectral subtraction to improve the signal-to-noise ratio of the speech signal.
[0073] Feature extraction includes: Mel-frequency cepstral coefficients and other features; among them,
[0074] Mel-frequency cepstral coefficients: Extract the spectral features of the audio and simulate the non-linear perception characteristics of the human ear for sound frequencies.
[0075] Other features: Some models directly process the original waveform using linear predictive coding or deep neural networks.
[0076] In some embodiments of the present application, the acoustic model constructs the relationship between phonemes and acoustic features based on the hidden Markov model and Gaussian mixture model. The language model is a neural network model, which is used to learn the word order rules based on a large amount of text data, correct the output errors of the acoustic model, can map the phoneme sequence output by the acoustic model into words, and optimize semantic coherence in combination with the context.
[0077] In some embodiments of the present application, the terminal device performs local conversion processing on the audio file through the first audio-to-text conversion model to obtain local text data corresponding to the audio file; uploads the audio file to the server; receives the server text data corresponding to the audio file sent by the server; and determines the text data corresponding to the audio file based on the local text data and the server text data.
[0078] In some embodiments of the present application, the terminal device performs local conversion processing on the audio file through the first audio-to-text conversion model to obtain local text data corresponding to the audio file; uploads the audio file to the server; in the case of not receiving the server text data sent by the server, determines the local text data as the text data corresponding to the audio file.
[0079] S103. Classify based on the attribute information of the text data to obtain a classification result; and store the text data based on the classification result.
[0080] In the embodiments of the present application, the attribute information of the text data includes the text data name and the text data time.
[0081] It should be noted that when converting an audio file into text output through a preset audio-to-text conversion model, the generated text data includes timestamps and segmentation information. The timestamp can be regarded as the time of the text data; based on the segmentation information, the name of the text data can be determined.
[0082] In some embodiments of the present application, the terminal device can classify the text data based on the name of the text data and / or the time of the text data to obtain a classification result; and store the text data based on the classification result of the text data.
[0083] In some embodiments of the present application, the terminal device can classify the text data based on the name of the text data to obtain a classification result of the text data; and store the text data in the terminal device and the server according to the classification result of the text data.
[0084] In some embodiments of the present application, the terminal device can classify the text data based on the time of the text data to obtain a classification result of the text data; and store the text data in the terminal device and the server according to the classification result of the text data.
[0085] In some embodiments of the present application, the terminal device can classify the text data based on the name of the text data and the time of the text data to obtain a classification result of the text data; and store the text data in the terminal device and the server according to the classification result of the text data.
[0086] It can be understood that on the one hand, in response to the operation of answering a call, all calls can be voiced to obtain audio files of all calls. By converting the audio files into text data and classifying and storing the text data, since the audio files are converted into text data, there is no need to listen to the audio files during subsequent queries, but directly query the text data, which can improve the query speed; on the other hand, querying the text content of the call according to the classification result can further improve the efficiency of finding the call content.
[0087] In some embodiments of the present application, as Figure 2 shown, S102 can be implemented through S201, S202, S203, and S204 as follows:
[0088] S201. Perform local conversion processing on the audio file through the first audio-to-text conversion model to obtain the local text data corresponding to the audio file.
[0089] In some embodiments of the present application, the first audio-to-text conversion model is an audio-to-text conversion model deployed on the terminal device. The computing power of the first audio-to-text conversion model is smaller than that of the second audio-to-text conversion model.
[0090] In some embodiments of the present application, the terminal device may perform local conversion processing on the audio file through the first audio-to-text conversion model to obtain the local text data corresponding to the audio file.
[0091] In some embodiments of the present application, the terminal device may perform audio preprocessing on the audio file through the first audio-to-text conversion model to obtain the preprocessed audio file; extract features from the preprocessed audio file to obtain an audio signal. Through the acoustic model and the language model, the audio signal corresponding to the audio file is converted into text output to obtain the local text data corresponding to the audio file.
[0092] S202. Upload the audio file to the server.
[0093] In some embodiments of the present application, the terminal device sends the audio file to the server.
[0094] In some embodiments of the present application, when the terminal device is connected to the server, it uploads the audio file to the server so that the server can process the audio file.
[0095] It should be noted that the server may be in the cloud.
[0096] S203. Receive the server text data corresponding to the audio file sent by the server; wherein, the server text data is obtained by the server through converting and processing the audio file by the second audio-to-text conversion model.
[0097] In some embodiments of the present application, after the terminal device uploads the audio file to the server, it may receive the server text data corresponding to the audio file sent by the server.
[0098] In some embodiments of the present application, the server text data is obtained by the server through converting and processing the audio file by the second audio-to-text conversion model.
[0099] In some embodiments of the present application, the server performs audio preprocessing and feature extraction on the audio file through the second audio-to-text conversion model to obtain the audio signal extracted by the server. Through the acoustic model and the language model in the second audio-to-text conversion model, the audio signal corresponding to the audio file is converted into text output to obtain the server text data corresponding to the audio file.
[0100] S204. Determine the text data corresponding to the audio file based on the local text data and the server text data.
[0101] In some embodiments of the present application, after the terminal device obtains the local text data through the first audio-to-text conversion model, if it receives the server text data sent by the server, it may determine the server text data as the text data corresponding to the audio file.
[0102] In some embodiments of the present application, based on local text data and server text data, error comparison is performed through an audio file to determine the error value between the local text data and the server text data; based on the error value between the local text data and the server text data, the text data corresponding to the audio file is determined.
[0103] In some embodiments of the present application, the terminal device can perform error comparison through an audio file based on local text data and server text data to determine the error value between the local text data and the server text data; based on the error value between the local text data and the server text data, the text data with a smaller error value is determined as the text data corresponding to the audio file.
[0104] It can be understood that the local text data is obtained by converting the audio file locally through the first audio-to-text conversion model; the server text data is obtained by converting the audio file on the server through the second audio-to-text conversion model, and then the local text data and the server text data are screened to determine the text data corresponding to the audio file. By different conversion methods and screening, the accuracy of the text data can be improved.
[0105] In some embodiments of the present application, as Figure 3 shown, after S202, S205 is further executed as follows:
[0106] S205. In the case where the server text data sent by the server is not received, the local text data is determined as the text data corresponding to the audio file.
[0107] In some embodiments of the present application, if the terminal device does not receive the server text data sent by the server, the local text data is determined as the text data corresponding to the audio file.
[0108] It should be noted that the terminal device not receiving the server text data sent by the server may be due to a problem with the network connection between the terminal device and the server, resulting in the terminal device being unable to receive any data sent by the server, or it may also be that the server itself has a fault and is unable to convert the audio file uploaded by the terminal device, so no data is sent to the terminal device.
[0109] It can be understood that in the case where the server text data sent by the server is not received, determining the local text data as the text data corresponding to the audio file can ensure obtaining the text data corresponding to the audio file.
[0110] In some embodiments of the present application, classification is performed based on the attribute information of the text data, and the classification results include:
[0111] Classify the text data based on the text data name and / or the text data time to obtain a classification result.
[0112] In some embodiments of the present application, classify the text data based on the text data name to obtain a classification result; or classify the text data based on the text data time to obtain a classification result; or classify the text data based on the text data name and the text data time to obtain a classification result.
[0113] In some embodiments of the present application, the terminal device can classify the text data based on the text data name to obtain a classification result.
[0114] Exemplarily, the text data name of text data 1 is "Call with the manager of XX Company", the text data name of text data 2 is "Call with the president of XX University", and the text data name of text data 3 is "Call with XX classmate". It can be distinguished according to business and life. Text data 1 and text data 2 are classified into one category, that is, work calls, and text data 3 is classified into life calls.
[0115] In some embodiments of the present application, the terminal device can classify the text data based on the text data time to obtain a classification result.
[0116] Exemplarily, the text data time of text data 1 is "2025 / 01 / 12 14:12", the text data time of text data 2 is "2025 / 01 / 12 20:15", and the text data time of text data 3 is "2025 / 01 / 17 09:07". It can be classified in chronological order. Specifically, since the dates of text data 1 and text data 2 are both 2025 / 01 / 12, text data 1 and text data 2 are classified into one category; text data 3 is classified into one category.
[0117] In some embodiments of the present application, the terminal device can classify the text data based on the text data name and the text data time to obtain a classification result.
[0118] Exemplarily, the text data name of text data 1 is "Call with the manager of XX Company", the text data time is "2025 / 01 / 12 14:12", the text data name of text data 2 is "Call with the president of XX University", the text data time is "2025 / 01 / 12 20:15", the text data name of text data 3 is "Call with XX classmate", and the text data time is "2025 / 01 / 17 09:07". It can be classified in chronological order and according to the text data name. Text data 1 and text data 2 are classified into one category; text data 3 is classified into one category.
[0119] It can be understood that classifying text data by text data name and text data time facilitates subsequent classified storage of text data and searching for call content.
[0120] In some embodiments of the present application, based on the classification result, the text data is stored, including:
[0121] Based on the classification result, the text data is stored locally;
[0122] The classification result and the text data are sent to the server so that the server stores the text data based on the classification result.
[0123] In some embodiments of the present application, the terminal device can store the text data locally based on the classification result of the text data. The classification result and the text data of the text data can also be sent to the server so that the server stores the text data based on the classification result.
[0124] It can be understood that storing the text data locally and on the server facilitates subsequent retrieval of the text data from the server and, when the terminal device is replaced, synchronously obtaining the text data.
[0125] In some embodiments of the present application, as Figure 4 shown, the terminal device includes a phone status module 11, an audio recording module 22, an audio conversion module 33, a cloud audio conversion module 44, a text recording module 55, and a classification and saving module 66; among them,
[0126] The phone status module 11 is used to handle the answering and hanging up status of the phone. When the phone is in the connected state, the audio recording module is used for call recording. When in the hanging up state, the audio recording module stops recording and saves the call audio.
[0127] The audio recording module 22 is used to record the audio data of the call and transmit the audio to the audio conversion module. It starts recording the call audio data after receiving the phone connection, stops recording the audio data after the phone is hung up, and saves the current audio. After the audio processing is completed, the currently saved audio file is deleted.
[0128] The audio conversion module 33 is used to convert the saved recording data into text data and delete the saved audio after the conversion. Due to limited local computing power, to ensure the conversion quality, the audio conversion is divided into local conversion and cloud conversion. The recording data is input into both local conversion and cloud conversion simultaneously. If the result from the cloud is not received, the local result is used; if the result from the cloud is received, the cloud result is used. The audio conversion module uses the optimal result through arbitration.
[0129] The cloud audio conversion module 44 converts the audio into text by setting up a server and utilizing cloud computing power to ensure accuracy.
[0130] The text recording module 55 saves the converted text. The text is saved both locally and in the cloud to facilitate convenient viewing on different devices.
[0131] The classification and storage module 66 saves the text by establishing a database according to information such as name and time for easy query. Through classification and storage, it is convenient to search.
[0132] It can be understood that this application can record all call information; the previous recorded information can be quickly viewed, improving the user experience and personalized needs. Using the terminal and cloud audio conversion improves the conversion accuracy.
[0133] In some embodiments of this application, such as Figure 5 shown, the call text conversion method includes:
[0134] S1. Monitor the phone call.
[0135] In some embodiments of this application, before monitoring the phone call, a cloud server is established and an audio conversion system is built.
[0136] S2. Save the audio.
[0137] In some embodiments of this application, monitor the call connection and disconnection status. Start phone recording after the call is answered. Stop phone recording after the call is hung up and save the recorded audio.
[0138] It should be noted that the audio is the audio file during the call process.
[0139] S3. Convert the audio to text.
[0140] In some embodiments of this application, input the recorded audio into the local conversion module and the cloud conversion module for call text conversion to obtain local text data and cloud text data, and arbitrate the local text data and the cloud text data to determine the converted text data.
[0141] S4. Record the text.
[0142] In some embodiments of this application, save the converted result file and upload it to the cloud for storage. Establish a database for the file and establish classification information for easy query.
[0143] It can be understood that this application can record all call information; the previous recorded information can be quickly viewed, improving the user experience and personalized needs. Using the terminal and cloud audio conversion improves the conversion accuracy.
[0144] Some embodiments of this application also provide a call text conversion device, such as Figure 6 shown, Figure 6Schematic structural diagram of a call text conversion device provided by an embodiment of the present application. The call text conversion device 60 includes: a determination unit 601, a conversion unit 602, a classification unit 603, and a storage unit 604; where
[0145] The determination unit 601 is configured to perform call recording in response to an operation of answering a call, and determine an audio file of the call process.
[0146] The conversion unit 602 is configured to perform conversion processing on the basis of the audio file through a preset audio-to-text conversion model to obtain text data corresponding to the audio file; where the preset audio-to-text conversion model is used to convert an audio file into text data.
[0147] The classification unit 603 is configured to classify based on the attribute information of the text data to obtain a classification result.
[0148] The storage unit 604 is configured to store the text data based on the classification result.
[0149] In some embodiments of the present application, the preset audio-to-text conversion model includes a first audio-to-text conversion model and a second audio-to-text conversion model.
[0150] The call text conversion device 6 further includes: a sending unit 605 and a receiving unit 606; where
[0151] The conversion unit 602 is further configured to perform local conversion processing on the audio file through the first audio-to-text conversion model to obtain local text data corresponding to the audio file.
[0152] The sending unit 605 is configured to upload the audio file to the server.
[0153] The receiving unit 606 is configured to receive server text data corresponding to the audio file sent by the server; where the server text data is obtained by the server performing conversion processing on the audio file through the second audio-to-text conversion model.
[0154] The determination unit 601 is further configured to determine the text data corresponding to the audio file based on the local text data and the server text data.
[0155] In some embodiments of the present application, after the determination unit 601 uploads the audio file to the server, in the case where the server text data sent by the server is not received, the local text data is determined as the text data corresponding to the audio file.
[0156] In some embodiments of the present application, the determining unit 601 is further configured to compare the errors between the local text data and the server text data based on the local text data and the server text data through the audio file, and determine the error value between the local text data and the server text data; and determine the text data corresponding to the audio file based on the error value between the local text data and the server text data.
[0157] In some embodiments of the present application, the attribute information includes the text data name and the text data time;
[0158] The classifying unit 603 is further configured to classify the text data based on the text data name and / or the text data time to obtain a classification result.
[0159] In some embodiments of the present application, the classifying unit 603 is further configured to classify the text data based on the text data name to obtain a classification result; or classify the text data based on the text data time to obtain a classification result; or classify the text data based on the text data name and the text data time to obtain a classification result.
[0160] In some embodiments of the present application, the storage unit 604 is further configured to store the text data locally based on the classification result;
[0161] The sending unit 605 is further configured to send the classification result and the text data to the server, so that the server stores the text data based on the classification result.
[0162] Based on the call text conversion method of the above embodiments, an embodiment of the present application further provides a terminal device, as Figure 7 shown. Figure 7 FIG. is a schematic structural diagram of a terminal device provided by an embodiment of the present application. The terminal device 70 includes: a processor 701 and a memory 702. The memory 702 is used to store a computer program; the processor 701 is used to call and run the computer program from the memory 702 to execute the call text conversion method as described in the above embodiments.
[0163] In an embodiment of the present application, the above-mentioned processor 701 may be at least one of an Application Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), a Digital Signal Processing Device (DSPD), a Programmable Logic Device (PLD), a Field Programmable Gate Array (FPGA), a Central Processing Unit (CPU), a controller, a microcontroller, and a microprocessor. It can be understood that for different devices, the electronic devices for implementing the functions of the above-mentioned processor may also be others, and the embodiments of the present application do not make specific limitations.
[0164] The embodiment of the present application provides a computer-readable storage medium storing a computer program, which is used to implement the call text conversion method described in any one of the above embodiments when being executed by a processor.
[0165] Exemplarily, the program instructions corresponding to a call text conversion method in this embodiment may be stored on a storage medium such as an optical disc, a hard disk, or a USB flash drive. When the program instructions corresponding to a call text conversion method in the storage medium are read or executed by an electronic device, the call text conversion method described in any one of the above embodiments can be implemented.
[0166] In addition, in the embodiments of the present application, each functional module may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional module.
[0167] In addition, in the embodiments of the present application, each functional module may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional module.
[0168] When the integrated unit is implemented in the form of a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a vehicle to execute all or part of the steps of the method of this embodiment.
[0169] It should be understood that the "one embodiment" or "an embodiment" or "some embodiments" mentioned throughout the specification means that the specific features, structures, or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, the "in one embodiment" or "in an embodiment" or "in some embodiments" that appear throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the order numbers of the above processes do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The serial numbers of the embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments. The above descriptions of each embodiment tend to emphasize the differences between the embodiments. Their similarities or similarities can be referred to each other. For the sake of brevity, they will not be repeated herein.
[0170] The modules described as separate components above may or may not be physically separated. The components shown as modules may or may not be physical modules; they can be located in one place or distributed to multiple network units; some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0171] In addition, in each embodiment of the present application, all the functional modules can be integrated into one processing unit, or each module can be separately used as a unit, or two or more modules can be integrated into one unit; the above integrated modules can be implemented in the form of hardware or in the form of a combination of hardware and software functional units.
[0172] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments.
[0173] The methods disclosed in several method embodiments provided by the embodiments of the present application can be arbitrarily combined without conflict to obtain new method embodiments.
[0174] The features disclosed in several product embodiments provided by the embodiments of the present application can be arbitrarily combined without conflict to obtain new product embodiments.
[0175] The features disclosed in several method or device embodiments provided by the embodiments of the present application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0176] As mentioned above, it is only the implementation manner of the embodiments of the present application, but the protection scope of the embodiments of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the embodiments of the present application. Therefore, the protection scope of the embodiments of the present application shall be subject to the protection scope of the claimed rights.
Claims
1. A call text conversion method, characterized in that, The described call text conversion method includes: In response to an incoming call operation, perform call recording to determine an audio file of the call process; Based on the audio file, perform conversion processing through a preset audio-to-text conversion model to obtain text data corresponding to the audio file; wherein, the preset audio-to-text conversion model is used to convert an audio file into text data; Classify based on the attribute information of the text data to obtain a classification result; and based on the classification result, store the text data.
2. The call text conversion method according to claim 1, wherein The preset audio-to-text conversion model includes a first audio-to-text conversion model and a second audio-to-text conversion model; The step of, based on the audio file, performing conversion processing through a preset audio-to-text conversion model to obtain text data corresponding to the audio file, includes: Through the first audio-to-text conversion model, perform local conversion processing on the audio file to obtain local text data corresponding to the audio file; Upload the audio file to the server; Receive the server text data corresponding to the audio file sent by the server; wherein, the server text data is obtained by the server performing conversion processing on the audio file through the second audio-to-text conversion model; Based on the local text data and the server text data, determine the text data corresponding to the audio file.
3. The call text conversion method according to claim 2, wherein After uploading the audio file to the server, the method further includes: In the case where the server text data sent by the server is not received, determine the local text data as the text data corresponding to the audio file.
4. The call text conversion method according to claim 2, characterized in that, The step of, based on the local text data and the server text data, determining the text data corresponding to the audio file, includes: Based on the local text data and the server text data, perform error comparison through the audio file to determine the error value between the local text data and the server text data; Based on the error value between the local text data and the server text data, determine the text data corresponding to the audio file.
5. The call text conversion method according to claim 1, wherein The attribute information includes the text data name and the text data time; The step of, based on the attribute information of the text data, classifying to obtain a classification result, includes: Classify the text data based on the text data name, and / or, the text data time, to obtain a classification result.
6. The call text conversion method according to claim 5, wherein The step of, based on the text data name, and / or, the text data time, classifying the text data to obtain a classification result, includes: Classify the text data based on the text data name to obtain a classification result; or, Classify the text data based on the text data time to obtain a classification result; or, Classify the text data based on the text data name and the text data time to obtain a classification result.
7. The call text conversion method according to any one of claims 1-6, characterized in that, The step of, based on the classification result, storing the text data, includes: Based on the classification result, store the text data locally; Send the classification result and the text data to the server so that the server stores the text data based on the classification result.
8. A call text conversion device, characterized in that It includes a determination unit, a conversion unit, a classification unit, and a storage unit; wherein, the determination unit is configured to perform call recording in response to an operation of answering a call and determine an audio file of the call process; the conversion unit is configured to perform conversion processing on the basis of the audio file through a preset audio-to-text conversion model to obtain text data corresponding to the audio file; wherein, the preset audio-to-text conversion model is used to convert an audio file into text data; the classification unit is configured to classify based on the attribute information of the text data to obtain a classification result; the storage unit is configured to store the text data based on the classification result.
9. A vehicle, characterized in that, It includes: a memory for storing executable data instructions; a processor, when executing the executable instructions stored in the memory, implements the call text conversion method according to claims 1-7.
10. A computer-readable storage medium, characterized in that, There are executable instructions stored, which are used to cause the processor to implement the call text conversion method according to any one of claims 1 to 7 when executed.