Business auxiliary method and system for army telephone traffic duty

Through the multi-person voice source separation of voice information, key information refinement of call content and multi-dimensional evaluation of call relationships during military service duty, the problem of insufficient effectiveness and reliability of telephone duty services in the prior art is solved, and efficient and accurate military service communication is achieved.

CN120220727AActive Publication Date: 2025-06-27SHAANXI LINGFANG TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510363270.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-06-27
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

In the prior art, the effectiveness and reliability of military telephone duty services are poor, and the quality of military telephone communications cannot be guaranteed, which affects subsequent service management.

Method used

By collecting voice information, separating multiple voice sources, identifying the subject objects of both parties in the call, dividing call categories, refining key information, building a call information timeline, identifying call associations, performing multi-dimensional evaluations, and integrating evaluation indicators to evaluate the duty quality of telephone personnel.

Benefits of technology

It improves the effectiveness and reliability of telephone duty services, ensures the quality of military telephone communications, enhances the effective flow of information communications, and improves the efficiency and accuracy of telephone duty services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220727A_ABST
    Figure CN120220727A_ABST
Patent Text Reader

Abstract

The invention discloses a business auxiliary method and system for army telephone traffic duty, and relates to the technical field of voice data analysis, and the method comprises the steps: carrying out the multi-person sound source separation processing of voice information, and separating the sound sources of multiple persons speaking at the same time in a call, so as to accurately analyze the content spoken by each person. And a reliable basis is provided for subsequent voice analysis, text analysis and the like. The text content of each call is recognized, key information is extracted from the text content, and then the key information is extracted according to the occurrence condition of keywords. According to the method, correlation among multiple calls on a call information time axis is identified, and multi-dimensional evaluation is carried out, so that the effectiveness and reliability of a telephone traffic duty service are improved, the quality of army telephone traffic communication is ensured, evaluation indexes of all correlated call events on the call information time axis are integrated, and the duty quality of telephone traffic personnel is evaluated; the efficiency and the accuracy of telephone traffic service duty are improved, and effective circulation of information communication on the troops is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of voice data analysis, and particularly to a service assistance method and system for military telephone duty. Background Art

[0002] With the rapid development of modern military communication technology, as a key link in information transmission, the efficiency and accuracy of military telephone duty are directly related to the smooth operation of the command system and the successful completion of combat missions. To improve the effectiveness of telephone duty, it is particularly important to develop a service assistance plan. The plan aims to optimize the telephone operation process by introducing advanced communication technologies and intelligent means, improve the accuracy of telephone operators' understanding of call content, and at the same time speed up the telephone response speed and the transmission efficiency of conveyed tasks. In addition, the plan will also consider factors such as tone and intonation in the call voice to comprehensively improve the service quality and customer satisfaction of telephone duty. Through the application of these series of technical means, military telephone duty will be more efficient and intelligent, providing solid communication support for military operations.

[0003] In the prior art, for the completion of operations such as the transfer of duty telephones of telephone operators, it often requires manual evaluation, resulting in poor effectiveness and reliability of telephone duty operations, unable to guarantee the quality of military telephone communication, and being unfavorable for subsequent service management.

[0004] Therefore, how to improve the effectiveness and reliability of telephone duty operations is a technical problem to be solved currently. Summary of the Invention

[0005] The purpose of the present invention is to solve the problem of poor effectiveness and reliability of telephone duty operations in the prior art, and to propose a service assistance method for military telephone duty, which includes:

[0006] Collect the voice information generated during the calls of military telephone operators during duty, perform multi-person sound source separation processing on the voice information to obtain single-person sound source voice, and identify the two main subject objects in the call through the single-person sound source voice;

[0007] Divide the call categories according to the call request subject, identify the text content of each call, extract key information from the text content, and arrange the key information through time stamps to obtain a call information timeline;

[0008] Identify the associations between multiple calls on the call information timeline and record them as associated call events, and perform multi-dimensional evaluation on each associated call event to obtain evaluation indicators;

[0009] Integrate the evaluation indicators of all associated call events on the call information timeline, evaluate the duty quality of telephone operators, and thus assist in the improvement and optimization of the duty operations of telephone operators.

[0010] In some embodiments of the present application, before performing multi-person sound source separation processing on voice information, the method further includes

[0011] preprocessing the voice information, where the preprocessing includes noise reduction processing and voice activity detection;

[0012] Split the voice signal of the voice information into multiple partial voice signals, calculate the energy distribution and fundamental frequency sequence of each partial voice signal, and count the short-term energy, zero-crossing rate, and peak number. Compare the short-term energy, zero-crossing rate, and peak number with their respective preset thresholds to obtain three deviation amounts, and synthesize the three deviation amounts to obtain the overall deviation amount for each partial voice signal;

[0013] Extract the voice features on each partial voice signal and input them into a preset overlap detection model. The overlap detection model outputs an overlap probability. Set the probability threshold corresponding to each partial voice signal according to the overall deviation amount, and compare the overlap probability of each partial voice signal with the probability threshold to detect the overlapping partial voice signals in the voice signal, and perform multi-person sound source separation processing on the overlapping partial voice signals.

[0014] In some embodiments of the present application, performing multi-person sound source separation processing on the overlapping partial voice signals includes

[0015] Using the method of blind source separation to implement multi-person sound source separation processing on the overlapping partial voice signals, and the method of blind source separation includes independent component analysis and non-negative matrix factorization.

[0016] In some embodiments of the present application, identifying the two main subject objects in the call by single-person sound source speech includes

[0017] Compare the voiceprint feature of the single-person sound source speech with the pre-stored voiceprint library of the operator to identify the operator main body and the opposite person main body. The two main subject objects include both the operator main body and the opposite person main body.

[0018] In some embodiments of the present application, dividing the call category according to the call request subject and identifying the text content of each call includes

[0019] The call categories include two types: outgoing calls and incoming calls;

[0020] When the call request subject is the operator main body, this call is an outgoing call;

[0021] When the call request subject is the opposite person main body, this call is an incoming call;

[0022] Use speech recognition technology to convert each call into preliminary text content, analyze the text difficulty features of the preliminary text content, count all the text difficulty features of the preliminary text content, and evaluate the text recognition difficulty of the text content based on all the text difficulty features;

[0023] Set the context window size of the NLP context understanding model based on the text recognition difficulty of the text content, and implement the specific text recognition of the context of the call, so as to recognize the text content of each call.

[0024] In some embodiments of the present application, key information is extracted from the text content, including,

[0025] Perform part-of-speech tagging on the text content, calculate the TF and IDF of the vocabulary in the text content to screen out keywords, and perform semantic analysis on the text content through the keywords, so as to extract key information.

[0026] In some embodiments of the present application, the associations between multiple calls on the call information timeline are recognized, including,

[0027] Construct association rules based on the two-party subject objects, time windows, key information similarities, outgoing call and incoming call categories of different calls. Use calls as nodes and the strength of the association rules between different calls as edges to construct a call information association graph. Represent outgoing calls and incoming calls with different colors, and identify the associations between multiple calls through graph theory algorithms.

[0028] In some embodiments of the present application, multi-dimensional evaluation is performed on each associated call event to obtain evaluation indicators, including,

[0029] Collect the manual record information of the associated call events during the duty of the operator, compare the matching situation between the manual record information and the text content and key information of the calls of the associated call events, and generate the first evaluation indicator;

[0030] Analyze the tone of each call voice in the associated call event to generate the second evaluation indicator;

[0031] Analyze the task transmission situation of the operator in the associated call event to generate the third evaluation indicator.

[0032] Correspondingly, the present application also provides a business assistance system for military call duty, including,

[0033] The first module is used to collect the voice information generated during the call of the military operator on duty, perform multi-person sound source separation processing on the voice information to obtain single-person sound source voice, and recognize the two-party subject objects in the call through the single-person sound source voice;

[0034] The second module is used to divide call categories according to the call request subject, identify the text content of each call, extract key information from the text content, and arrange the key information through timestamps to obtain the call information timeline;

[0035] The third module is used to identify the associations between multiple calls on the call information timeline and record them as associated call events, and conduct multi-dimensional evaluations on each associated call event to obtain evaluation indicators;

[0036] The fourth module is used to integrate the evaluation indicators of all associated call events on the call information timeline, evaluate the duty quality of the telephone operators, and assist in the improvement and optimization of the duty operations of the telephone operators.

[0037] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0038] 1. Perform multi-person sound source separation processing on voice information, separate the sound sources of multiple people speaking simultaneously in a call, so as to accurately analyze the content spoken by each person, providing a reliable basis for subsequent voice analysis and text analysis, etc. Identify the text content of each call, extract key information from the text content, consider the context of each call to identify accurate text content, and then extract key information based on the occurrence of keywords to better reflect the key points in the call and reduce redundant content.

[0039] 2. Identify the associations between multiple calls on the call information timeline and conduct multi-dimensional evaluations, thereby unifying the completion standards of the telephone operator's duty operations, improving the effectiveness and reliability of the telephone operator's duty operations, ensuring the quality of military telephone communications, integrating the evaluation indicators of all associated call events on the call information timeline, evaluating the duty quality of the telephone operators, thereby assisting in the optimization and improvement of the telephone operator's operations, improving the efficiency and accuracy of the telephone operator's duty operations, and ensuring the effective flow of information communication in the military. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a schematic flow chart of a method for assisting the duty of military telephone operators proposed by the present invention;

[0041] Figure 2 It is a schematic structural diagram of a system for assisting the duty of military telephone operators proposed by the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0042] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.

[0043] Refer to Figure 1, A business assistance method for military telephone duty, including the following steps,

[0044] Step S101, collect the voice information generated during the call of military telephone operators during duty, perform multi-person sound source separation processing on the voice information to obtain single-person sound source voice, and identify the two main objects in the call through the single-person sound source voice recognition.

[0045] In this embodiment, military telephone duty refers to telephone operators in the military who are responsible for communication liaison. At the specified time and place, according to the responsibilities and task requirements, they carry out communication guarantee work such as telephone connection transfer and information transmission. This is an indispensable part of the military communication system and is of great significance for ensuring smooth command and timely and accurate information transmission. Therefore, monitoring the duty quality of telephone operators is crucial for military communication. During the duty of military telephone operators, the voice information generated during the call is collected in real time through special recording equipment or software. These devices should ensure high-quality recording to reduce the difficulty and error of subsequent processing. The recording data should be stored in digital format for subsequent analysis and processing. During the call, due to the influence of people around, there may be a situation where multiple people speak at the same time, which is recorded in the call and has a great impact on the call voice recognition and needs to be processed. Advanced audio processing technologies, such as the Blind Source Separation (BSS) algorithm, are used to separate the collected multi-person call voice information. Common methods include Independent Component Analysis (ICA) and deep learning-based sound source separation models. These technologies can analyze the different sound source characteristics in the audio signal and separate the mixed voice signal into single-person sound source voice.

[0046] In some embodiments of the present application, before performing multi-person sound source separation processing on the voice information, the method further includes,

[0047] Preprocess the voice information, and the preprocessing includes noise reduction processing and voice activity detection;

[0048] Split the voice signal of the voice information into multiple partial voice signals, calculate the energy distribution and fundamental frequency sequence of each partial voice signal, and count the short-time energy, zero-crossing rate and peak number. Compare the short-time energy, zero-crossing rate and peak number with their respective preset thresholds to obtain three deviation amounts, and synthesize the three deviation amounts to obtain the overall deviation amount under each partial voice signal;

[0049] Extract the speech features on each part of the speech signal and input them into a preset overlap detection model. The overlap detection model outputs the overlap probability. Set the probability threshold corresponding to each part of the speech signal according to the overall deviation amount, and compare the overlap probability of each part of the speech signal with the probability threshold to detect the overlapping part of the speech signal in the speech signal, and perform multi-person sound source separation processing on the overlapping part of the speech signal.

[0050] In this embodiment, first, it is necessary to detect which signal segments in the speech information belong to the overlapping part of the speech signal, so as to perform targeted separation on this part. Short-time energy: reflects the amplitude change of the speech signal, and the energy value increases significantly when multiple people are speaking. Zero-crossing rate: measures the frequency at which the signal passes through zero, and the zero-crossing rate fluctuates greatly due to frequency superposition when multiple people are speaking. YIN algorithm: estimates the fundamental frequency period by analyzing the autocorrelation function of the signal, and multiple fundamental frequency peaks can be detected when multiple people are speaking, and the number of peaks is the number of fundamental frequency peaks. The preset thresholds of the short-time energy, zero-crossing rate, and the number of peaks can be set by historical experience. The three deviation amounts reflect the possibility of overlap from three angles, generating an overall deviation amount, thereby setting the probability threshold. Deep learning overlap detection, training an LSTM or Transformer model (overlap detection model), inputting MFCC features or spectrograms, and outputting the overlap probability of each frame.

[0051] In some embodiments of the present application, multi-person sound source separation processing is performed on the overlapping part of the speech signal, including

[0052] Adopt the method of blind source separation to realize the multi-person sound source separation processing of the overlapping part of the speech signal. The method of blind source separation includes independent component analysis and non-negative matrix factorization.

[0053] In this embodiment, the implementation process of independent component analysis (ICA)

[0054] 1. Preprocessing

[0055] Frame division and windowing: Divide the speech signal into frames (such as a frame length of 25 ms and a frame shift of 10 ms), and add a Hamming window.

[0056] Fourier transform: Perform a short-time Fourier transform (STFT) on each frame of the signal to obtain the frequency domain representation.

[0057] 2. Centering and whitening

[0058] Centering: Subtract the signal mean to make the data zero-mean.

[0059] Whitening: Remove the correlation between signals through PCA to obtain the whitened signal.

[0060] 3. Independent component estimation

[0061] Objective function: Maximize non-Gaussianity (usually using negative entropy as a measure).

[0062] Optimization algorithm: Use the fast ICA algorithm (such as FastICA) to solve the separation matrix through fixed-point iteration.

[0063] 4. Post-processing

[0064] Amplitude adjustment: Scale the amplitude of the separated signal.

[0065] Signal sorting: Sort the separated signals by correlation or fundamental frequency matching.

[0066] Non-negative matrix factorization is achieved through matrix factorization, iterative optimization, and signal reconstruction.

[0067] In some embodiments of the present application, the two main subjects in the call are identified through single-speaker voice recognition, including

[0068] Comparing the voiceprint features of the single-speaker voice with the pre-stored voiceprint library of the operator to identify the operator subject and the opposite-party subject. The two main subject objects include the operator subject and the opposite-party subject.

[0069] In this embodiment, the two parties of the call connected by the operator include the operator subject and the opposite-party subject. Identifying the operator facilitates the quality evaluation of the subsequent operator's duty operations.

[0070] Step S102: Divide the call category according to the call request subject, identify the text content of each call, extract the key information from the text content, and arrange the key information through timestamps to obtain the call information timeline.

[0071] In this embodiment, the call category is divided according to the call request subject. Based on the operator, the call category is divided. One is that the operator answers the incoming call, which is an incoming call, and the other is that the operator dials a call to other points, which is an outgoing call. Key information is extracted from these two types of calls, and a call information timeline is established to facilitate the association between calls.

[0072] In some embodiments of the present application, the call category is divided according to the call request subject, and the text content of each call is identified, including

[0073] The call category includes two types: outgoing calls and incoming calls;

[0074] When the call request subject is the operator subject, the call is an outgoing call;

[0075] When the call request subject is the opposite-party subject, the call is an incoming call;

[0076] Use speech recognition technology to convert each call into preliminary text content, analyze the text difficulty features of the preliminary text content, count all the text difficulty features on the preliminary text content, and evaluate the text recognition difficulty of the text content based on all the text difficulty features;

[0077] Set the context window size of the NLP context understanding model based on the text recognition difficulty of the text content, and implement the specific text recognition of the context of the call, so as to recognize the text content of each call.

[0078] In this embodiment, ordinary speech-to-text conversion cannot consider the context content association of the call, resulting in poor accuracy of text content recognition. First, perform preliminary recognition through speech-to-text technology, set the text difficulty situation, and thus set the context window size of the NLP context understanding model to help the NLP context understanding model better specifically recognize the text content. Use speech recognition technology (such as ASR, Automatic Speech Recognition) to convert the speech content of the call into text. This step can use existing speech recognition services or open-source tools, such as Google Speech-to-Text, IBM Watson Speech to Text, etc. The text difficulty features include the ratio of the number of word types (i.e., the number of different words) to the total number of words, the average sentence length (counted in words or characters), the number of nested clauses, the length of the conversation, the frequency of topic switching, etc. Normalize these text difficulty features, and evaluate the text recognition difficulty (The difficulty of text recognition) of the text content based on all the text difficulty features, which is achieved through the following formula.

[0079]

[0080] Among them, is the text recognition difficulty of the i1-th text content, n is the number of text difficulty features of the i1-th text content, is the weight of the i2-th text difficulty feature, is the magnitude of the i2-th text difficulty feature of the i1-th text content, are respectively the maximum and minimum values in, is the first constant of the i1-th text content, which can be determined by the length of the text content, represents the correction of the average value of all text difficulty features by the average of the maximum and minimum values in, is to balance the magnitude of the correction function.

[0081] In some embodiments of the present application, key information is extracted from the text content, including

[0082] Part-of-Speech Tagging (POS Tagging) is performed on the text content, and the TF and IDF of words in the text content are calculated to screen out keywords, and semantic analysis is performed on the text content through the keywords, so as to extract key information.

[0083] In this embodiment, Part-of-Speech Tagging (POS Tagging) is a basic task in NLP, which assigns a part-of-speech tag to each word in the text, such as noun, verb, adjective, etc. Using the part-of-speech tagging information, key entities and actions in the text can be identified more accurately. Keyword extraction is to identify the most representative words or phrases from the text, and these words or phrases can usually summarize the main content of the text. When extracting key information, keyword extraction can help quickly locate the core content of the text. TF-IDF method: By calculating the frequency of a word in the text (TF) and the inverse document frequency (IDF) in the entire corpus, the importance of the word is evaluated, so as to extract keywords. Semantic analysis is a key step in understanding the meaning of the text, which involves a deep understanding of words, phrases and sentences in the text. When extracting key information, semantic analysis can help identify implicit information and complex relationships in the text. Dependency syntactic analysis: Analyze the dependency relationships between words in a sentence, such as subject-predicate relationship, verb-object relationship, etc., so as to understand the meaning of the sentence more accurately. Sentiment analysis: Identify the sentiment tendency in the text, such as positive, negative or neutral, which helps to understand the attitude and emotion of the caller. Event extraction: Identify events and their related elements from the text, such as event type, participants, time, location, etc., so as to extract key information in the call.

[0084] Step S103, identify the associations between multiple calls on the call information timeline, and record them as associated call events, and perform multi-dimensional evaluation on each associated call event to obtain evaluation indicators.

[0085] In this embodiment, the association between multiple calls means the calls involved in the operator's information transmission. For example, when the operator answers a call and is required to convey a certain instruction to a certain communication point, the requirements are identified through the text content, including the situation of the conveying party, time, specific content of the conveyance, etc. After the operator understands the required content, the operator may immediately make a call to contact the conveying party. From the time the operator answers the call to the time the operator makes a call to the conveying party until the conveyance is completed, the calls involved in this process are the associated calls of a certain associated call event.

[0086] In some embodiments of the present application, identifying the associations between multiple calls on the call information timeline includes

[0087] Based on the two parties of different calls, time window, similarity of key information, outgoing call and incoming call categories, association rules are constructed. Taking calls as nodes and the strength of association rules between different calls as edges, a call information association graph is constructed. Outgoing calls and incoming calls are represented by different colors, and the associations between multiple calls are identified through graph theory algorithms.

[0088] In this embodiment, on the call information timeline, by analyzing features such as call content, identities of both parties of the call, and call time, the associations between multiple calls are identified. For example, one call may be a reply or supplement to another call, or multiple calls may revolve around the same event or task. These associations can be identified through rule matching, graph theory algorithms, or machine learning models. The association rules are combinations of similarity of key information, adjacent time, subject object situation, etc., and outgoing call and incoming call categories. For example, when an incoming call ends, if the operator receives a transfer request, an outgoing call operation should be carried out immediately. Each call is regarded as a node, and the association between calls is regarded as an edge. According to features such as call content, identities of both parties of the call, and call time, a call information graph is constructed. Graph theory algorithms (such as community detection, shortest path algorithm, etc.) are used to further analyze the call information graph. These algorithms can help us identify potential associations and patterns between calls. For example, through the community detection algorithm, we can identify groups of closely connected nodes in the call information graph, and these groups of nodes may represent multiple calls that revolve around the same event or task.

[0089] The similarity of key information includes the following:

[0090] Topic recognition: First, perform topic analysis on the call content to extract the main discussion points or topics of each call. For example, topic modeling techniques (such as LDA) can be used to identify potential topics in the text.

[0091] Keyword extraction: Through keyword extraction algorithms (such as TF-IDF, TextRank, etc.), extract the core words or phrases in each call, and these words or phrases can often reflect the main content of the call.

[0092] Semantic analysis: Use natural language processing techniques for semantic analysis to understand the deep meaning in the call content. For example, through dependency syntax analysis, identify the subject-predicate-object structure in the sentence, and then understand the core information of the call.

[0093] Call time analysis, time series analysis: Arrange the calls in chronological order and analyze the time intervals and sequences between calls. If multiple calls occur continuously in a short period of time, or follow a certain specific time pattern (such as daily, weekly, etc.), then there may be an association between these calls.

[0094] In some embodiments of the present application, each associated call event is evaluated from multiple dimensions to obtain evaluation indicators, including

[0095] Collect the manual record information of the associated call events during the duty of the operator, compare the matching situation between the manual record information and the text content and key information of the calls of the associated call events, and generate the first evaluation indicator;

[0096] Analyze the tone of each call voice in the associated call event to generate the second evaluation indicator;

[0097] Analyze the task transmission situation of the operator in the associated call event to generate the third evaluation indicator.

[0098] In this embodiment, during the duty of the operator, the content of the transfer instruction or requirement he heard will be recorded, and the matching situation between the manual record information and the text content and key information will be converted into the first evaluation indicator. Essentially, it is to see the degree of consistency between the content recorded by the operator and the actual content. The tone of the operator's call voice can reflect the service attitude and other content, which is converted into the second evaluation indicator. According to the key information and the transfer situation of the operator (the text content of the call made by the operator), the task transmission situation of the operator in the associated call event is analyzed to obtain the third evaluation indicator.

[0099] It can be understood that other dimensions that can evaluate the completion quality of the operator's duty can also be included here, all of which belong to the protection scope of the present application.

[0100] Step S104: Integrate the evaluation indicators of all associated call events on the call information timeline, evaluate the duty quality of the operator, and use this to assist in the improvement and optimization of the operator's duty service.

[0101] In this embodiment, integrate the evaluation indicators of all associated call events on the call information timeline (evaluate and integrate the evaluation indicators in units of time and associated call events), and conduct a comprehensive evaluation of the duty quality of the operator (The evaluation of duty performance). It can be achieved through the following formula:

[0102]

[0103] where Te is the comprehensive duty quality of the operator, m is the number of associated call events, β1 j 、β2 j 、β3 j are the respective evaluation weights of the first evaluation indicator, the second evaluation indicator, and the third evaluation indicator under the jth associated call event, respectively, and W1 j 、W2 j 、W3j They are the magnitudes of the first evaluation index, the second evaluation index, and the third evaluation index respectively under the j-th associated call event, and t j is the completion time of the j-th associated call event, and k2 j is the second constant corresponding to the j-th associated call event, represents the correction of the sum of the first evaluation index, the second evaluation index, and the third evaluation index by the completion time of the associated call event.

[0104] According to the evaluation results, improve and optimize the duty operations of the telephone operators. For example, for the problem of slow response speed, the training of telephone operators can be strengthened to improve their reaction speed; for the problem of information accuracy, the information verification mechanism can be improved; for the problem of communication efficiency, the call process and information transmission method can be optimized; for the problem of service attitude, the professional quality training of telephone operators can be strengthened; for the problem of poor task completion, task tracking and supervision can be strengthened.

[0105] Correspondingly, the present application also provides a business assistance system for military telephone duty, such as Figure 2 shown, including,

[0106] The first module is used to collect the voice information generated during the calls of military telephone operators during duty, perform multi-person sound source separation processing on the voice information to obtain single-person sound source voices, and identify the two main subject objects in the call through the single-person sound source voices;

[0107] The second module is used to divide the call categories according to the call request subject, identify the text content of each call, extract key information from the text content, and arrange the key information through time stamps to obtain the call information time axis;

[0108] The third module is used to identify the associations between multiple calls on the call information time axis and record them as associated call events, and perform multi-dimensional evaluations on each associated call event to obtain evaluation indicators;

[0109] The fourth module is used to integrate the evaluation indicators of all associated call events on the call information time axis, evaluate the duty quality of telephone operators, and thus assist in the improvement and optimization of the duty operations of telephone operators.

[0110] Compared with the prior art, the beneficial effects of the present invention are:

[0111] 1. Perform multi-person sound source separation on voice information to separate the sound sources of multiple people speaking simultaneously during a call, so as to accurately analyze the content spoken by each person, providing a reliable basis for subsequent voice analysis, text analysis, etc. Identify the text content of each call, extract key information from the text content, consider the context of each call to identify accurate text content, and then extract key information based on the occurrence of keywords to better reflect the key points in the call and reduce redundant content.

[0112] 2. Identify the associations between multiple calls on the call information timeline and conduct multi-dimensional evaluations, thereby unifying the completion standards of the operator's duty operations, improving the effectiveness and reliability of the call duty operations, ensuring the quality of the military call communication, integrating the evaluation indicators of all associated call events on the call information timeline, evaluating the duty quality of the call personnel, thereby assisting in the optimization and improvement of the operator's business, improving the efficiency and accuracy of the call duty operations, and ensuring the effective circulation of information communication in the military.

[0113] Through the description of the above implementation manners, those skilled in the art can clearly understand that the present invention can be implemented through hardware or by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various implementation scenarios of the present invention.

[0114] Those skilled in the art can understand that the drawings are only schematic diagrams of a preferred implementation scenario, and the modules or processes in the drawings are not necessarily essential for implementing the present invention.

[0115] Those skilled in the art can understand that the modules in the system in the implementation scenario can be distributed in the system of the implementation scenario according to the description of the implementation scenario, or can be correspondingly changed to be located in one or more systems different from the present implementation scenario. The modules in the above implementation scenario can be combined into one module, or can be further split into multiple sub-modules.

[0116] The above is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent replacements or changes, and should be covered by the protection scope of the present invention.

Claims

1. A service assistance method for military telephone duty, characterized in that: include, Collect the voice information generated by the military operators during their on-duty conversations, perform multi-person sound source separation on the voice information, obtain the single-person sound source speech, and identify the two main subjects in the conversation through the single-person sound source speech; Classify calls according to the call request body, identify the text content of each call, extract key information from the text content, arrange the key information by timestamp, and obtain the call information timeline; Identify the associations between multiple calls on the call information timeline and record them as associated call events, and perform multi-dimensional evaluation on each associated call event to obtain evaluation indicators; Integrate the evaluation indicators of all related call events on the call information timeline to evaluate the duty quality of operators, so as to assist the improvement and optimization of operators' duty business.

2. The service assistance method for military telephone duty according to claim 1, characterized in that: Before performing multi-person sound source separation processing on the voice information, the method further includes preprocessing the voice information, the preprocessing including noise reduction processing and voice activity detection; The speech signal of the speech information is split into multiple partial speech signals, the energy distribution and fundamental frequency sequence of each partial speech signal are calculated, and the short-time energy, zero-crossing rate and peak number are counted, and the short-time energy, zero-crossing rate and peak number are compared with their respective preset thresholds to obtain three deviations, and the overall deviation of each partial speech signal is obtained by combining the three deviations; The speech features of each partial speech signal are extracted and input into the preset overlap detection model. The overlap detection model outputs the overlap probability. The probability threshold corresponding to each partial speech signal is set according to the overall deviation. The overlap probability of each partial speech signal and the probability threshold are compared to detect the overlapping partial speech signals in the speech signal, and the overlapping partial speech signals are subjected to multi-person sound source separation processing.

3. The service assistance method for military telephone duty according to claim 2, characterized in that: Perform multi-person sound source separation on overlapping speech signals, including: The method of blind source separation is used to realize the separation of multiple voice sources of overlapping speech signals. The method of blind source separation includes independent component analysis and non-negative matrix decomposition.

4. The service assistance method for military telephone duty according to claim 1, characterized in that: Recognize the two parties in a call through the voice of a single person, including: The voiceprint features of the single-person sound source speech are compared with the voiceprint library pre-stored by the operator to identify the operator subject and the opposite person subject, and the two-party subject objects include the operator subject and the opposite person subject.

5. The service assistance method for military telephone duty according to claim 4 is characterized in that: Classify calls according to the call request body and identify the text content of each call, including: Call categories include outgoing calls and incoming calls; When the call requesting subject is the operator subject, the call is an outgoing call; When the call requesting subject is the opposite person subject, the call is an incoming call; Use speech recognition technology to convert each call into preliminary text content, analyze the text difficulty features of the preliminary text content, count all the text difficulty features on the preliminary text content, and evaluate the text recognition difficulty of the text content based on all the text difficulty features; The context window size of the NLP context understanding model is set based on the difficulty of text recognition based on the text content, and the specific text recognition of the call context is implemented to identify the text content of each call.

6. The service assistance method for military telephone duty according to claim 1, characterized in that: Extract key information from the text content, including: Perform part-of-speech tagging on the text content, calculate the TF and IDF of the words in the text content to filter out keywords, and perform semantic analysis on the text content through keywords to extract key information.

7. The service assistance method for military telephone duty according to claim 5, characterized in that: Identify the relationship between multiple calls on the call information timeline, include, Association rules are constructed based on the subject objects of both parties of different calls, time windows, similarity of key information, and categories of outgoing calls and incoming calls. A call information association graph is constructed with calls as nodes and the association rule strengths of different calls as edges. Outgoing calls and incoming calls are represented by different colors, and the associations between multiple calls are identified through graph theory algorithms.

8. The service assistance method for military telephone duty according to claim 1, characterized in that: Perform multi-dimensional evaluation on each associated call event to obtain evaluation indicators, including: Collecting manual recording information of call events related to the operator's duty process, comparing the manual recording information with the text content and key information of the call related to the call event, and generating a first evaluation index; Analyze the tone of each call voice in the associated call event to generate a second evaluation index; Analyze the task communication of the operator in the related call events to generate the third evaluation index.

9. A business assistance system for military telephone duty, characterized in that: include, The first module is used to collect the voice information generated by the military operators during their on-duty conversations, perform multi-person sound source separation on the voice information, obtain single-person sound source speech, and identify the two main subjects in the conversation through the single-person sound source speech; The second module is used to classify calls according to the call request body, identify the text content of each call, extract key information from the text content, arrange the key information by timestamp, and obtain the call information timeline; The third module is used to identify the associations between multiple calls on the call information timeline and record them as associated call events, and to perform multi-dimensional evaluation on each associated call event to obtain evaluation indicators; The fourth module is used to integrate the evaluation indicators of all related call events on the call information timeline, evaluate the duty quality of the operators, and assist in improving and optimizing the duty business of the operators.

Citation Information

Patent Citations

  • Intelligent evaluation method and system for service quality of telephone customer service

    CN116828109A

  • Audio processing method and system

    CN117524259A

  • Conversation reminding method and system of customer service system

    CN118803134A