Voice communication method based on artificial intelligence
By integrating voice signal noise reduction, filtering, and natural language processing technologies, a call information parsing chain is constructed, which solves the shortcomings of information recognition and management in traditional voice communication technology, realizes intelligent and automated information processing and personalized reminders, and improves information processing efficiency and accuracy.
Patent Information
- Application Number
- CN202511213068.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-11-18
AI Technical Summary
Traditional voice communication technology cannot effectively identify and extract key information in a call, requiring users to manually organize and record the information. Furthermore, it lacks intelligent and personalized management of call data, failing to meet the demands of modern communication for efficient and accurate information processing.
By employing speech signal noise reduction, filtering, speech recognition, and natural language processing technologies, a call information parsing chain is constructed. Through word segmentation, part-of-speech tagging, and syntactic analysis, combined with an important information keyword database and weight formulas, structured text annotations are generated, and automated processing is achieved through text-to-speech technology.
It achieves intelligent and automated processing from voice input to information output, improving information processing efficiency and accuracy. The generated voice notes have personalized reminder functions, making it easy for users to quickly retrieve and record important information.
Smart Images

Figure CN120977310A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent voice communication, in particular to a voice communication method based on artificial intelligence. BACKGROUND
[0002] In today's information society, voice communication is one of the main ways of daily communication and work, and its efficiency and accuracy are crucial to improving the work efficiency of individuals and organizations. With the rapid development of artificial intelligence technology, applying AI technology to the field of voice communication to realize intelligent processing and analysis of voice signals has become the consensus of the scientific and industrial communities. Traditional voice communication mainly focuses on clear transmission of voice, but ignores the effective management and utilization of a large amount of valuable information generated during communication. How to quickly identify and record key information in the conversation has become an important challenge to improve communication efficiency and subsequent action accuracy.
[0003] Traditional voice communication technology mainly relies on basic speech coding, decoding and noise reduction technology when processing call content. Although these technologies can ensure clear transmission of voice to some extent, they are relatively weak in understanding and processing voice content when facing complex and variable information in the conversation. When facing a conversation containing a large amount of key information, traditional technology often cannot automatically identify and extract these information, resulting in users needing to manually organize and record after the call ends, which not only consumes time and effort, but also easily misses important information. In addition, traditional technology lacks the ability to analyze and utilize call data, and cannot provide personalized information management and reminder services. Moreover, traditional methods lack intelligence and automation in information processing, and cannot meet the needs of modern communication for efficient and accurate information processing.
[0004] Therefore, it is necessary to develop a voice communication method based on artificial intelligence. SUMMARY
[0005] The purpose of the present application is to make up for the shortcomings of the prior art and provide a voice communication method based on artificial intelligence. The present application integrates voice signal noise reduction, filtering processing, speech recognition and natural language processing technology to build a complete call information analysis chain. After converting audio to text, the natural language processing unit analyzes the grammatical structure and semantic relationship of the sentence through word segmentation, part-of-speech tagging and syntax analysis. By combining a preset important information keyword library and dynamically calculating the priority of keywords through an important information recognition weight formula, the capture of important information is ensured. Finally, the recognition results are removed of redundancy and integrated through a concise text generation formula to generate structured text notes, which are converted into voice signals through text-to-speech technology. The entire process does not require human intervention, realizing intelligent and automated processing from voice input to information output. Compared with traditional methods, the efficiency is significantly improved, and the accuracy and integrity of the information are fundamentally guaranteed.
[0006] The application provides the following technical solutions to solve the above technical problems: a voice communication method based on artificial intelligence, the specific steps of which are as follows: Voice information collection: collecting voice signals of both parties in a call through an audio input device and performing noise reduction processing, while retaining the effective components of the main frequency band of the voice through filtering to complete preprocessing; Important information identification: converting the preprocessed voice signal into a digital signal and obtaining text content through voice recognition technology, performing word segmentation, part-of-speech tagging, and syntax analysis using natural language processing technology, calling a preset important information keyword library, calculating the importance weight of keywords in combination with an important information identification weight formula, and determining important information; Voice note generation: generating concise text content for important information through a concise text generation formula, converting it into a voice signal through text-to-speech technology, determining the intonation parameters, and storing the compressed and encoded voice signal in audio format as a voice note; Intelligent reminder triggering: monitoring the call state, automatically playing the generated voice note through an audio output device after the call ends, and allowing the user to control the playback state through a preset operation; Multi-dimensional label generation and retrieval: generating a time label based on the call end time and a character label based on the call participants using artificial intelligence, and allowing the user to retrieve through the labels.
[0007] Further, in the important information identification, the natural language processing technology is used to perform word segmentation on the text content, cutting continuous text into independent words according to semantics, then performing part-of-speech tagging to determine the part of speech of each word, and then performing syntax analysis to clarify the grammatical relationship between words, determine the structure of subject-predicate-object, and understand the overall meaning of the sentence. At the same time, a preset important information keyword library is called, which covers keywords such as to-do items, anniversaries, and key data, and the importance weight of keywords is calculated in combination with an important information identification weight formula to determine important information.
[0008] Further, in the important information identification, the importance weight of keywords is calculated in combination with an important information identification weight formula, which is as follows: wherein, is the importance weight of keywords, is a preset keyword library matching coefficient, is a keyword matching rate, i.e., the proportion of the number of matched keywords in the total number of keywords in the recognized text, is an information frequency coefficient, i.e., the frequency of this type of information being marked as important information in historical calls.
[0009] Further, in the important information recognition, the keyword importance weight is calculated according to the important information recognition weight formula, and when the calculated keyword importance weight is , the important information is determined.
[0010] Further, in the voice note generation, the time, person, event and data key elements are extracted from the recognized important information, the key elements are preliminarily integrated to form an original text, the original text length is calculated, the redundant information of repeated expressions is marked, the proportion of the marked redundant information in the original text is calculated to determine the information redundancy, the concise text generation formula is set according to the information redundancy, the concise length is calculated through the concise text generation formula, the original text is trimmed according to the concise length, the marked redundant information is removed, and the concise text is generated.
[0011] Further, in the voice note generation, the concise length is calculated through the concise text generation formula, and the concise text generation formula is: , wherein is the concise length, is the original text length, is the concise coefficient, is the information complexity index, is the complexity compensation coefficient.
[0012] Further, in the voice note generation, the voice signal is converted into a voice signal through a text-to-speech technology, and the voice signal collected by the audio input device is divided into multiple small segments, the waveform change rule of each small segment is analyzed, the tone frequency of the small segment is obtained, and the number of tone frequencies is calculated to form a tone frequency number sequence , the original voice average tone is calculated, and the tone parameter is determined through the tone adaptation formula.
[0013] Further, in the voice note generation, the tone parameter is determined through the tone adaptation formula, and the tone adaptation formula is: , wherein is the tone parameter, is the information importance index, is the original voice average tone.
[0014] Further, in the intelligent reminding trigger, the call process is continuously monitored through signal feedback to determine whether the call is in progress or in the end state, and after determining the end of the call, the voice note generated for this call is retrieved from the storage location, and the generated voice note is automatically played through the audio output device. During the playing process, the playing state can be controlled by a preset operation, such as pausing, closing or replaying.
[0015] Further, in the multi-dimensional label generation and retrieval, the artificial intelligence generates a time label based on the end time of the call as the original data, and forms a standardized time label by sorting, and extracts the information of the call participants from the call record, and directly uses the call participant information as a character label, and binds the generated time label and character label with the voice note.
[0016] Compared with the prior art, the voice communication method based on artificial intelligence has the following beneficial effects: I. The present application integrates voice signal noise reduction, filtering processing, speech recognition and natural language processing technology, and constructs a complete call information analysis chain. After converting the audio into text, the natural language processing unit analyzes the syntax structure and semantic relationship of the sentence through word segmentation, part-of-speech tagging and syntax analysis, combines a preset important information keyword library, dynamically calculates the priority of the keywords through an important information recognition weight formula, ensures the capture of important information, and finally, through a concise text generation formula, the recognition result is removed and the elements are integrated. Generate structured text notes, and convert them into voice signals through text-to-speech technology. The entire process does not require human intervention, realizes intelligent and automatic processing from voice input to information output, significantly improves the efficiency of the traditional method, and fundamentally guarantees the accuracy and integrity of the information.
[0017] II. The present application analyzes the tone frequency characteristics of the original voice, dynamically adjusts the tone parameters of the voice note using a tone adaptation formula, makes the reminder voice more targeted and emphasized, enhances user perception, generates standardized time labels based on the end time of the call, extracts call participant information to generate character labels, and binds the voice note with the time label and character label. Users can accurately retrieve and find voice notes through labels to quickly locate voice notes.
[0018] Other advantages, objects, and features of the present application will be set forth in part in the following specification, and in part will become apparent to those skilled in the art from a consideration of the following description, or can be learned from practice of the present application. BRIEF DESCRIPTION OF DRAWINGS
[0019] In order to make the technical solutions in the embodiments of the present application or the prior art clearer, the accompanying drawings needed in the embodiments or prior art description will be briefly introduced below. Obviously, the accompanying drawings in the following description only show some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort on the basis of these drawings.
[0020] Figure 1 A framework diagram of a voice communication method based on artificial intelligence; Figure 2 A flowchart of a voice communication method based on artificial intelligence. DETAILED DESCRIPTION
[0021] In order to further illustrate the technical means and effects adopted by the present application to achieve the predetermined application purposes, the specific embodiments, structures, features and effects according to the present application will be described in detail below with reference to the accompanying drawings and preferred embodiments.
[0022] Embodiment one: voice information collection: the salesperson talks with the business client through the mobile phone, the mobile phone built-in audio input device synchronously collects the conversation voice signal of the salesperson and the procurement client, and carries out noise reduction processing on the voice signal, filters out the keyboard clicking sound, air conditioner running sound and other noises in the office background, at the same time, through filtering processing, the effective voice components of 300-3400Hz (main frequency band of human voice) are reserved, and the preprocessing of the voice signal is completed.
[0023] Important information recognition: the preprocessed voice signal is converted into a digital signal, a text content (such as "this time, I need a batch of A type equipment, which should be delivered to the warehouse before this Friday, and the down payment should be paid first") is generated through voice recognition technology, and the text is segmented into independent words through natural language processing technology, and the part of speech of each word is marked, the grammatical relationship is combed through syntax analysis, the overall meaning of the sentence is understood, and the preset business class important information keyword library (covering keywords such as "order quantity", "delivery date", "payment ratio", "equipment model", etc.) is called, the importance weight of the keywords is calculated through the important information recognition weight formula, and the important information recognition weight formula is: , wherein, is the importance weight of the keyword, is the preset keyword matching coefficient, is the keyword matching rate, that is, the proportion of the number of matched keywords in the total number of text words, The information frequency coefficient, i.e. the frequency of such information being marked as important information in historical conversations, is calculated to obtain keyword importance weight ≥ 0.5, and it is determined that "a batch of A-type equipment", "to be delivered to the warehouse before this Friday" and "a part of the deposit" are important information.
[0024] Voice note generation: from the recognized important information, the key elements (time: this Friday; person: procurement customer; event: delivery of order; data: a batch of A-type equipment, a part of the deposit) are extracted and preliminarily integrated into the original text "the customer needs a batch of A-type equipment, requires delivery to the warehouse before this Friday, and the payment is a part of the deposit", the system detects that there is no redundant information of repeated expression in the original text, sets the corresponding simplification coefficient, calculates the simplified length through the concise text generation formula, and the concise text generation formula is: , wherein, is the simplified length, is the original text length, is the simplification coefficient, is the information complexity index, is the complexity compensation coefficient, the original text is cut according to the simplified length (since there is no redundancy, it is directly retained), and the concise text "the customer needs a batch of A-type equipment, delivers to the warehouse before this Friday, and pays a part of the deposit" is generated. The concise text is converted into a voice signal through a text-to-speech technology, and the collected customer voice is divided into multiple small segments, the tone frequency of each small segment is analyzed according to the waveform change rule, and the number of tone frequencies is obtained, and a tone frequency number sequence is formed, the average tone of the original voice is calculated , and the tone parameter is determined through the tone adaptation formula, and the tone adaptation formula is: , wherein is the tone parameter, is the information importance index, is the average tone of the original voice, and the generated voice signal is compressed and encoded to be stored as a voice note in an audio format, as shown in Figure 1 .
[0025] Intelligent reminder triggering: the conversation state is monitored in real time through the signal feedback of the conversation process, and the voice note generated for this conversation is retrieved from the storage location after the conversation ends, and the generated voice note "the customer needs a batch of A-type equipment, delivers to the warehouse before this Friday, and pays a part of the deposit" is automatically played by the mobile phone loudspeaker. If the user does not hear clearly, the user can control the playing state through the preset operation to pause, close or replay.
[0026] Multi-dimensional label generation and retrieval: artificial intelligence obtains the call end time as the original data for generating time labels, and organizes and forms standard time labels. At the same time, the information of the call participants is extracted from the call record, and the call participant information is directly used as the character label. The generated time label and character label are associated and bound with the voice note. The user can subsequently search for the voice note quickly by searching for the time label or the character label.
[0027] In summary, in the business customer call scenario, the audio input device is used to collect and preprocess the speech signals of both parties, the natural language processing technology is used in combination with the important information keyword library to identify important information related to the order, and then a concise text is generated and converted into a voice note with an appropriate tone. The voice note is automatically played after the call ends and supports user control of the playback state. At the same time, time and character labels are generated to associate the voice note for retrieval, effectively assisting sales personnel in recording and following up on key information of business orders.
[0028] Example two: voice information collection: the child communicates with the mother through the WeChat voice call function, and the built-in audio input device of the mobile phone synchronously collects the speech signals of both parties, and performs noise reduction processing on the speech signals to filter out the noise in the mother's environment, and retains the effective components of the main frequency band of the speech through filtering to complete the preprocessing of the speech signals.
[0029] Important information identification: the preprocessed voice is converted into a digital signal, and a text is generated through voice recognition (such as "next weekend is your dad's birthday, let's have a family dinner at the old home restaurant, remember to buy a cake, and also invite your cousin"). Natural language processing technology is used to tokenize the text, which is divided into independent words and then tagged with part-of-speech. Through syntactic analysis, the grammatical relationship is combed, the overall meaning of the sentence is understood, and a preset important information keyword library for the family (covering keywords such as "birthday", "dinner time", "location", "required items", and "invited personnel") is called. Through the important information identification weight formula: , the importance weight of the keywords is calculated, and the calculated importance weight of the keywords is all ≥0.5, which determines that "next weekend is the father's birthday", "dinner at the old home restaurant in the evening", "buy a cake", and "invite a cousin" are important information.
[0030] Voice note generation: the key elements (time: next weekend evening; characters: father, cousin; event: birthday dinner; item: cake) are extracted from the identified important information, and the original text "next weekend is the father's birthday, the whole family will have dinner at the old home restaurant in the evening, and a cake needs to be bought, and the cousin needs to be invited" is preliminarily integrated. It is detected that there are repeated expressions (low redundancy) in the original text, the corresponding simplification coefficient is set, and the concise text generation formula is used: , calculate the compact length, trim according to the compact length, remove redundancy to generate concise text "next weekend father's birthday, go to the old home restaurant for dinner, buy a cake and invite cousin", and then convert it into a voice signal through text-to-speech technology, and divide the mother's voice into multiple small segments, analyze the waveform change rule of each small segment, obtain the tone frequency of the small segment, and count the number of tone frequencies , form a sequence of tone frequency numbers , calculate the average tone of the original voice , and determine the tone parameters through the tone adaptation formula: , and the generated voice signal is stored as a voice note in audio format after compression encoding.
[0031] Intelligent reminder trigger: After monitoring the end of the call through the signal feedback of the call process, the voice note generated for this call is retrieved from the storage location, and the phone speaker automatically plays the voice note "next weekend father's birthday, go to the old home restaurant for dinner, buy a cake and invite cousin". If the user needs to do something else, they can pause the playback through the preset operation, and then listen to it again through the preset operation, as shown in Figure 2 .
[0032] Multi-dimensional label generation and retrieval: artificial intelligence obtains the call end time as the original data for generating time labels, and organizes them into standardized time labels. At the same time, the information of the call participants is extracted from the call record, which is directly used as the character label. The generated time label and character label are associated and bound with the voice note, and the user can quickly retrieve the voice note by searching for the time label or character label.
[0033] In summary, in the family and friend call scenario, the voice signal is collected and preprocessed through the audio input device, the important information such as birthday dinner is recognized through natural language processing technology and family important information keyword library, the concise text is generated and converted into a voice note that meets the tone parameters, the playback is automatically played after the call ends and the user can control the playback, and the time and character labels associated with the voice note are generated for easy retrieval, which helps users to record and handle important information related to family affairs in a timely manner.
[0034] The above is only a preferred embodiment of the present application, and does not limit the present application in any form. Although the present application has been disclosed as above with a preferred embodiment, it is not intended to limit the present application. Any person skilled in the art can make some changes or modifications to the above disclosed technical content without departing from the scope of the technical solution of the present application, and any equivalent embodiments with equivalent changes and modifications are still within the scope of the technical solution of the present application.
Claims
1. A voice communication method based on artificial intelligence, characterized in that, The specific steps of this method are as follows: Voice information acquisition: The voice signals of both parties in the call are acquired through the audio input device and noise reduction is performed. At the same time, the effective components of the main frequency bands of the voice are retained through filtering to complete the preprocessing. Important information identification: The preprocessed speech signal is converted into a digital signal, and the text content is obtained through speech recognition technology. Natural language processing technology is used for word segmentation, part-of-speech tagging and syntactic analysis. At the same time, a preset important information keyword database is called up, and the importance weight of the keywords is calculated by combining the important information identification weight formula to determine the important information. Voice annotation generation: For important information, concise text content is generated using a concise text generation formula, then converted into a speech signal using text-to-speech technology, and the intonation parameters are determined. The generated speech signal is compressed and encoded and stored in audio format as a voice annotation. Intelligent reminder trigger: Monitors call status, and after the call ends, automatically plays the generated voice notes through the audio output device, and the user can control the playback status through preset operations; Multidimensional tag generation and retrieval: Artificial intelligence generates time tags based on the call end time and person tags based on the participants in the call, which users can then use for retrieval.
2. The voice communication method based on artificial intelligence according to claim 1, characterized in that, In the important information identification process, natural language processing technology is used to segment the text content into words, dividing the continuous text into independent words according to semantics. Then, part-of-speech tagging is performed to clarify the part of speech of each word. Next, syntactic analysis is used to sort out the grammatical relationships between words, determine the subject-verb-object and attributive-adverbial-complement structures, and understand the overall meaning of the sentence. At the same time, a preset important information keyword database is called, which covers keywords for to-do items, anniversaries, and key data. Combined with the important information identification weight formula, the importance weight of the keywords is calculated to determine the important information.
3. The voice communication method based on artificial intelligence according to claim 2, characterized in that, In the identification of important information, the importance weight of keywords is calculated using the important information identification weight formula, which is as follows: ,in, As keyword importance weight, The matching coefficient is the preset keyword database. Keyword matching rate, This is the information frequency coefficient.
4. The voice communication method based on artificial intelligence according to claim 3, characterized in that, In the identification of important information, the importance weight of keywords is calculated using the important information identification weight formula. When the calculated keyword importance weight... At that time, it was determined to be important information.
5. The voice communication method based on artificial intelligence according to claim 1, characterized in that, In the voice annotation generation process, key elements such as time, people, events, and data are extracted from the identified important information. These key elements are initially integrated to form the original text, and the length of the original text is calculated. At the same time, redundant information with repeated expressions is marked, and the proportion of marked redundant information in the original text is statistically analyzed to determine the information redundancy. Based on the information redundancy, a simplification coefficient is set in the concise text generation formula. The simplified length is calculated using the concise text generation formula, and the original text is trimmed according to the simplified length to remove the marked redundant information and generate concise text.
6. The voice communication method based on artificial intelligence according to claim 4, characterized in that, In the voice annotation generation process, the concise text length is calculated using a concise text generation formula, which is as follows: ,in, To reduce the length, The original text length. For simplification coefficients, Information complexity index This is the complexity compensation coefficient.
7. The voice communication method based on artificial intelligence according to claim 1, characterized in that, In the voice annotation generation process, text-to-speech technology is used to convert the text into a speech signal. The speech signal collected by the audio input device is divided into multiple segments, and the waveform variation pattern of each segment is analyzed to obtain the pitch frequency of each segment. The number of pitch frequencies is then counted. This forms a sequence of pitch frequency quantities. Calculate the average pitch of the original speech And determine the intonation parameters through intonation adaptation formula. .
8. The voice communication method based on artificial intelligence according to claim 6, characterized in that, In the voice annotation generation process, intonation parameters are determined using an intonation adaptation formula. Its intonation adaptation formula is: ,in, For intonation parameters, As an information importance index, The average pitch of the original speech.
9. The voice communication method based on artificial intelligence according to claim 1, characterized in that, During the intelligent reminder triggering, the call status is continuously monitored through signal feedback of the call process to determine whether the call is in progress or has ended. After the call ends, the voice notes generated for this call are retrieved from the storage location and automatically played through the audio output device. During playback, the playback status can be controlled by preset operations to pause, close, or replay.
10. The voice communication method based on artificial intelligence according to claim 1, characterized in that, In the multidimensional tag generation and retrieval process, artificial intelligence uses the end time of the call as the raw data for generating time tags, and organizes it into standardized time tags. At the same time, it extracts the information of the call participants from the call records, uses the call participant information directly as person tags, and associates and binds the generated time tags and person tags with voice notes.