A method and system for assisting the operation of a military operator

By performing multi-person voice source separation and multi-dimensional evaluation on the voice information of the military's call center duty, the issues of effectiveness and reliability of call center duty services were resolved, and efficient and accurate communication support was achieved.

CN120220727BActive Publication Date: 2025-11-04SHAANXI LINGFANG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510363270.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-11-04
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

In existing technologies, the effectiveness and reliability of military telephone service are poor, resulting in poor communication quality and affecting subsequent business management.

Method used

By collecting voice information and performing multi-person voice source separation processing, the main subjects of the call are identified, the call categories are divided, key information is extracted, a call information timeline is constructed, and multi-dimensional evaluation is carried out. Evaluation indicators are integrated to optimize the quality of duty performance.

Benefits of technology

It improved the effectiveness and reliability of call center duty operations, ensured the quality of troop communications, improved the efficiency and accuracy of information transmission, and standardized duty procedures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220727B_ABST
    Figure CN120220727B_ABST
Patent Text Reader

Abstract

The application discloses a kind of army traffic service's business auxiliary method and system, it is related to voice data analysis technical field, including the multi-voice source separation processing of voice information, the voice source of the simultaneous speech of multiple people in call is separated, to accurately analyze the content that each person said, provide reliable basis for subsequent voice analysis and text analysis etc..The text content of each call is identified, and the key information is extracted from the text content.According to the situation of key word, the key information is extracted.The association between multiple calls on the call information time axis is identified, and multidimensional evaluation is carried out, to improve the effectiveness and reliability of traffic service business, to ensure the quality of army traffic communication, to integrate all the evaluation indexes of associated call events on the call information time axis, to evaluate the quality of traffic personnel's service, to improve the efficiency and accuracy of traffic service, to ensure the effective circulation of army information communication.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of voice data analysis technology, and in particular to a business support method and system for military telephone operations. Background Technology

[0002] With the rapid development of modern military communication technology, troop telephone operations, as a crucial link in information transmission, directly impact the smooth operation of command systems and the successful completion of combat missions. Therefore, developing a business support solution is particularly important to improve the effectiveness of telephone operations. This solution aims to optimize telephone operation processes by introducing advanced communication technologies and intelligent methods, improving the accuracy of telephone operators' understanding of call content, and accelerating telephone response speed and the efficiency of task transmission. Furthermore, the solution will consider factors such as tone and intonation in the voice to comprehensively enhance the service quality and customer satisfaction of telephone operations. Through the application of these technologies, troop telephone operations will become more efficient and intelligent, providing solid communication support for military operations.

[0003] In existing technologies, the completion of tasks such as call transfer for duty personnel often requires human evaluation, which results in poor effectiveness and reliability of duty operations, fails to guarantee the quality of military communications, and is detrimental to subsequent business management.

[0004] Therefore, how to improve the effectiveness and reliability of call center operations is a technical problem that needs to be solved. Summary of the Invention

[0005] The purpose of this invention is to address the problem of poor effectiveness and reliability of existing telephone call duty services, and to propose a service assistance method for military telephone call duty, which includes:

[0006] The system collects voice information generated during calls by military telephone operators while on duty, performs multi-person voice source separation processing on the voice information to obtain single-person voice source speech, and identifies the two main subjects in the call through single-person voice source speech.

[0007] Calls are categorized based on the call requester, the text content of each call is identified, key information is extracted from the text content, and the key information is arranged by timestamps to obtain a call information timeline.

[0008] Identify the relationships between multiple calls on the call information timeline and record them as related call events. Evaluate each related call event from multiple dimensions to obtain evaluation metrics.

[0009] The evaluation metrics of all related call events on the call information timeline are integrated to evaluate the performance quality of call center agents, thereby assisting in the improvement and optimization of their work.

[0010] In some embodiments of this application, before performing multi-source voice separation processing on the speech information, the method further includes,

[0011] The speech information is preprocessed, including noise reduction and speech activity detection.

[0012] The speech signal of the speech information is divided into multiple parts of the speech signal. The energy distribution and fundamental frequency sequence of each part of the speech signal are calculated, and the short-time energy, zero-crossing rate and peak count are statistically analyzed. The short-time energy, zero-crossing rate and peak count are compared with their respective preset thresholds to obtain three deviations. The overall deviation of each part of the speech signal is obtained by combining the three deviations.

[0013] The speech features of each part of the speech signal are extracted and input into the preset overlap detection model. The overlap detection model outputs the overlap probability. The probability threshold corresponding to each part of the speech signal is set according to the overall deviation. The overlap probability of each part of the speech signal is compared with the probability threshold to detect the overlapping part of the speech signal. The overlapping part of the speech signal is then processed for multi-person sound source separation.

[0014] In some embodiments of this application, multi-source sound separation processing is performed on the overlapping portion of the speech signal, including...

[0015] Blind source separation is used to separate multiple speech sources in overlapping speech signals. The blind source separation method includes independent component analysis and nonnegative matrix factorization.

[0016] In some embodiments of this application, the two main subjects in a call are identified through single-person voice source speech recognition, including:

[0017] The voiceprint features of a single voice source are compared with the voiceprint database pre-stored by the operator to identify the operator and the other party. The two parties are the operator and the other party.

[0018] In some embodiments of this application, call categories are divided based on the call requesting party, and the text content of each call is identified, including:

[0019] Call categories include outgoing calls and incoming calls;

[0020] When the requesting party for the call is an operator, the call is an outgoing call.

[0021] When the requesting party in a call is a recipient, the call is an access call.

[0022] Each call is converted into preliminary text content using speech recognition technology. The text difficulty features of the preliminary text content are analyzed, and all text difficulty features on the preliminary text content are statistically analyzed. The text recognition difficulty of the text content is evaluated based on all text difficulty features.

[0023] The difficulty of text recognition based on text content is set by adjusting the context window size of the NLP context understanding model and realizing the specific text recognition of the context of the call, thereby recognizing the text content of each call.

[0024] In some embodiments of this application, key information is extracted from the text content, including:

[0025] The text content is tagged with parts of speech, and the TF and IDF of words in the text content are calculated to filter out keywords. The text content is then semantically analyzed using keywords to extract key information.

[0026] In some embodiments of this application, the correlation between multiple calls on the call information timeline is identified, including:

[0027] Association rules are constructed based on the two parties involved in different calls, time windows, similarity of key information, and the categories of outgoing and incoming calls. Calls are used as nodes, and the strength of association rules for different calls are used as edges to construct a call information association graph. Outgoing and incoming calls are represented by different colors, and graph theory algorithms are used to identify the associations between multiple calls.

[0028] In some embodiments of this application, each associated call event is evaluated from multiple dimensions to obtain evaluation metrics, including...

[0029] Collect manually recorded information of related call events during the call center's duty process, compare the matching of the manually recorded information with the text content and key information of the related call events, and generate the first evaluation index.

[0030] Analyze the tone of voice in each voice message in the associated call events to generate a second evaluation metric;

[0031] Analyze the task delivery of operators in related call events to generate a third evaluation indicator.

[0032] Correspondingly, this application also provides a business support system for military telephone operations, including,

[0033] The first module is used to collect voice information generated during the calls of the military telephone operators on duty, perform multi-person voice source separation processing on the voice information to obtain single-person voice source speech, and identify the two main subjects in the call through single-person voice source speech.

[0034] The second module is used to classify calls according to the call request subject, identify the text content of each call, extract key information from the text content, and arrange the key information by timestamp to obtain the call information timeline.

[0035] The third module is used to identify the relationships between multiple calls on the call information timeline and record them as related call events. Each related call event is evaluated in multiple dimensions to obtain evaluation indicators.

[0036] The fourth module is used to integrate evaluation metrics of all related call events on the call information timeline to evaluate the performance quality of call center agents, thereby assisting in the improvement and optimization of their work.

[0037] Compared with the prior art, the beneficial effects of this invention are as follows:

[0038] 1. Perform multi-source voice separation processing on voice information to separate the voice sources of multiple people speaking simultaneously in a call. This allows for accurate analysis of each person's speech, providing a reliable foundation for subsequent voice and text analysis. Identify the text content of each call, extract key information from the text, consider the context of each call to identify precise text content, and further extract key information based on keyword occurrences to better highlight the key points of the call and reduce redundant content.

[0039] 2. Identify the relationships between multiple calls on the call information timeline and conduct multi-dimensional evaluations to unify the completion standards of operators' duties, improve the effectiveness and reliability of call service, ensure the quality of military call communication, integrate the evaluation indicators of all related call events on the call information timeline, evaluate the duty quality of call operators, thereby assisting operators in optimizing and improving their work, improving the efficiency and accuracy of call service, and ensuring the effective flow of information communication in the military. Attached Figure Description

[0040] Figure 1 This is a flowchart illustrating a business support method for military telephone duty proposed in this invention;

[0041] Figure 2 This is a schematic diagram of the structure of a military communications duty support system proposed in this invention. Detailed Implementation

[0042] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0043] Reference Figure 1A method for assisting military telephone operators includes the following steps:

[0044] Step S101: Collect voice information generated during the call by the military telephone operator during duty, perform multi-person voice source separation processing on the voice information to obtain single-person voice source speech, and identify the two main subjects in the call through single-person voice source speech.

[0045] In this embodiment, "military telephone operator duty" refers to the communication support work performed by telephone operators in the military, at designated times and locations, according to their duties and tasks, including telephone call transfer and information transmission. This is an indispensable part of the military communication system and is of great significance for ensuring smooth command and timely and accurate information transmission. Therefore, monitoring the quality of telephone operators' duty is crucial for military communications. During the duty process, the voice information generated during the call is collected in real time using specialized recording equipment or software. These devices should ensure high-quality recording to reduce the difficulty and error of subsequent processing. The recording data should be stored in a digital format for easy subsequent analysis and processing. During the call, due to the influence of people around, multiple people may speak simultaneously, which is recorded and has a significant impact on the speech recognition. This requires processing, using advanced audio processing technologies, such as the Blind Source Separation (BSS) algorithm, to separate the collected multi-person call voice information. Commonly used methods include Independent Component Analysis (ICA) and deep learning-based sound source separation models. These technologies can analyze the characteristics of different sound sources in audio signals and separate mixed speech signals into individual voice sources.

[0046] In some embodiments of this application, before performing multi-source voice separation processing on the speech information, the method further includes,

[0047] The speech information is preprocessed, including noise reduction and speech activity detection.

[0048] The speech signal of the speech information is divided into multiple parts of the speech signal. The energy distribution and fundamental frequency sequence of each part of the speech signal are calculated, and the short-time energy, zero-crossing rate and peak count are statistically analyzed. The short-time energy, zero-crossing rate and peak count are compared with their respective preset thresholds to obtain three deviations. The overall deviation of each part of the speech signal is obtained by combining the three deviations.

[0049] The speech features of each part of the speech signal are extracted and input into the preset overlap detection model. The overlap detection model outputs the overlap probability. The probability threshold corresponding to each part of the speech signal is set according to the overall deviation. The overlap probability of each part of the speech signal is compared with the probability threshold to detect the overlapping part of the speech signal. The overlapping part of the speech signal is then processed for multi-person sound source separation.

[0050] In this embodiment, the first step is to detect which signal segments in the speech information belong to the overlapping speech signal, and then perform targeted separation of this part. Short-time energy: reflects the amplitude change of the speech signal; the energy value increases significantly when multiple people are speaking. Zero-crossing rate: measures the frequency at which the signal crosses zero; the zero-crossing rate fluctuates more due to frequency superposition when multiple people are speaking. YIN algorithm: estimates the fundamental frequency period by analyzing the signal's autocorrelation function; multiple fundamental frequency peaks can be detected when multiple people are speaking, and the number of peaks is equal to the number of fundamental frequency peaks. The preset thresholds for short-time energy, zero-crossing rate, and number of peaks can be set based on historical experience. These three deviations reflect the possibility of overlap from three perspectives, generating an overall deviation value to set the probability threshold. Deep learning overlap detection is used to train an LSTM or Transformer model (overlap detection model), inputting MFCC features or a spectrogram, and outputting the overlap probability for each frame.

[0051] In some embodiments of this application, multi-source sound separation processing is performed on the overlapping portion of the speech signal, including...

[0052] Blind source separation is used to separate multiple speech sources in overlapping speech signals. The blind source separation method includes independent component analysis and nonnegative matrix factorization.

[0053] In this embodiment, the independent component analysis (ICA) process is described.

[0054] 1. Preprocessing

[0055] Framing and Windowing: The audio signal is framed (e.g., 25ms frame length, 10ms frame shift) and a Hamming window is added.

[0056] Fourier Transform: Perform a Short Time Fourier Transform (STFT) on each frame of the signal to obtain a frequency domain representation.

[0057] 2. Centralization and Whitening

[0058] Centralization: Subtract the signal mean to make the data mean zero.

[0059] Whitening: The correlation between signals is removed by PCA to obtain the whitened signal.

[0060] 3. Independent component estimation

[0061] Objective function: Maximize non-Gaussianity (negative entropy is commonly used as a metric).

[0062] Optimization algorithm: Use a fast ICA algorithm (such as FastICA) to solve the separation matrix through fixed-point iteration.

[0063] 4. Post-processing

[0064] Amplitude adjustment: Amplitude scaling is applied to the separated signal.

[0065] Signal sorting: Sort separated signals by correlation or fundamental frequency matching.

[0066] Nonnegative matrix factorization is achieved through matrix factorization, iterative optimization, and signal reconstruction.

[0067] In some embodiments of this application, the two main subjects in a call are identified through single-person voice source speech recognition, including:

[0068] The voiceprint features of a single voice source are compared with the voiceprint database pre-stored by the operator to identify the operator and the other party. The two parties are the operator and the other party.

[0069] In this embodiment, the two parties in a call connected by the operator include the operator and the other party, which helps to identify the operator and facilitates the quality evaluation of the operator's subsequent duties.

[0070] Step S102: Classify calls according to the call request subject, identify the text content of each call, extract key information from the text content, and arrange the key information by timestamp to obtain the call information timeline.

[0071] In this embodiment, calls are categorized based on the requesting party. Using the operator as a foundation, calls are classified into two types: one where the operator answers an incoming call (incoming call), and another where the operator makes a call to another location (outgoing call). Key information is extracted from these two types of calls, and a call information timeline is established to facilitate the correlation between calls.

[0072] In some embodiments of this application, call categories are divided based on the call requesting party, and the text content of each call is identified, including:

[0073] Call categories include outgoing calls and incoming calls;

[0074] When the requesting party for the call is an operator, the call is an outgoing call.

[0075] When the requesting party in a call is a recipient, the call is an access call.

[0076] Each call is converted into preliminary text content using speech recognition technology. The text difficulty features of the preliminary text content are analyzed, and all text difficulty features on the preliminary text content are statistically analyzed. The text recognition difficulty of the text content is evaluated based on all text difficulty features.

[0077] The difficulty of text recognition based on text content is set by adjusting the context window size of the NLP context understanding model and realizing the specific text recognition of the context of the call, thereby recognizing the text content of each call.

[0078] In this embodiment, ordinary speech-to-text conversion cannot consider the contextual relationships of the conversation, resulting in poor text recognition accuracy. First, speech-to-text technology is used for preliminary recognition to set text difficulty levels, thereby setting the context window size of the NLP context understanding model and helping it better recognize the specific text content. Then, speech recognition technology (such as ASR, Automatic Speech Recognition) is used to convert the speech content of the conversation into text. This step can use existing speech recognition services or open-source tools, such as Google Speech-to-Text and IBM Watson Speech to Text. Text difficulty features include the ratio of the number of word types (i.e., the number of different words) to the total number of words, average sentence length (in words or characters), the number of nested clauses, the length of the conversation, and the frequency of topic changes. These text difficulty features are normalized, and the text recognition difficulty is evaluated based on all text difficulty features using the following formula.

[0079]

[0080] in, Let be the text recognition difficulty of the i-th text content, and n be the number of text difficulty features of the i-th text content. The weight of the i2th text difficulty feature. Let i be the magnitude of the difficulty feature of the i2th text content of the i1th text content. They are respectively The maximum and minimum values ​​in the range. The first constant of the i1th text content can be determined by the length of the text content. express The average of the maximum and minimum values ​​in the data is used to correct for the average of all text difficulty features. This is to balance the size of the correction function.

[0081] In some embodiments of this application, key information is extracted from the text content, including:

[0082] The text content is tagged with parts of speech, and the TF and IDF of words in the text content are calculated to filter out keywords. The text content is then semantically analyzed using keywords to extract key information.

[0083] In this embodiment, Part-of-Speech Tagging (POS Tagging) is a fundamental task in NLP, assigning a part-of-speech tag, such as noun, verb, or adjective, to each word in the text. Using POS tagging information, key entities and actions in the text can be identified more accurately. Keyword extraction involves identifying the most representative words or phrases from the text, which typically summarize the main content. Keyword extraction helps quickly locate the core content of the text when extracting key information. The TF-IDF method assesses the importance of words by calculating their frequency (TF) in the text and their inverse document frequency (IDF) across the entire corpus, thereby extracting keywords. Semantic analysis is a crucial step in understanding the meaning of text, involving a deep understanding of words, phrases, and sentences. Semantic analysis helps identify implicit information and complex relationships in the text when extracting key information. Dependency parsing analyzes the dependency relationships between words in a sentence, such as subject-verb and verb-object relationships, to more accurately understand the meaning of the sentence. Sentiment analysis: Identifying the sentiment tendency in the text, such as positive, negative, or neutral, helps in understanding the speaker's attitude and emotions. Event extraction: Identifying events and related elements from the text, such as event type, participants, time, and location, to extract key information from the call.

[0084] Step S103: Identify the relationships between multiple calls on the call information timeline and record them as related call events. Evaluate each related call event in multiple dimensions to obtain evaluation indicators.

[0085] In this embodiment, the association between multiple calls means the calls involved in the information transmission by the operator. For example, when an operator enters a call, they are asked to transmit a certain instruction to a certain communication point. The requirement is identified through the text content, including the situation of the transmitting party, the time, and the specific content to be transmitted. After understanding the requirement, the operator may immediately make a call to contact the transmitting party. From the time the operator enters the call to the time the operator makes the call to the transmitting party until the transmission is completed, the calls involved in this process are all related calls of a certain related call event.

[0086] In some embodiments of this application, the correlation between multiple calls on the call information timeline is identified, including:

[0087] Association rules are constructed based on the two parties involved in different calls, time windows, similarity of key information, and the categories of outgoing and incoming calls. Calls are used as nodes, and the strength of association rules for different calls are used as edges to construct a call information association graph. Outgoing and incoming calls are represented by different colors, and graph theory algorithms are used to identify the associations between multiple calls.

[0088] In this embodiment, the relationships between multiple calls are identified by analyzing features such as call content, the identities of the callers, and call time on the call information timeline. For example, one call may be a response to or supplement to another call, or multiple calls may revolve around the same event or task. These relationships can be identified using rule matching, graph theory algorithms, or machine learning models. Relationship rules are combinations of similarities in key information, adjacent times, and subject information, categorizing calls as outgoing and incoming. For example, if an operator receives a relay request after an incoming call ends, they should immediately initiate an outgoing call. Each call is treated as a node, and the relationships between calls are considered edges. A call information graph is constructed based on features such as call content, the identities of the callers, and call time. Graph theory algorithms (such as community detection and shortest path algorithms) are used to further analyze the call information graph. These algorithms can help us identify potential relationships and patterns between calls. For example, through community detection algorithms, we can identify closely connected groups of nodes in the call information graph, which may represent multiple calls revolving around the same event or task.

[0089] Similarities in key information include the following:

[0090] Topic identification: First, the call content is analyzed to extract the main discussion points or topics of each call. For example, topic modeling techniques (such as LDA) can be used to identify latent topics in the text.

[0091] Keyword extraction: Using keyword extraction algorithms (such as TF-IDF, TextRank, etc.), the core words or phrases in each call are extracted. These words or phrases often reflect the main content of the call.

[0092] Semantic analysis: Utilizing natural language processing techniques to perform semantic analysis and understand the deeper meaning within the content of a conversation. For example, through dependency parsing, subject-verb-object structures in sentences can be identified, thereby understanding the core information of the conversation.

[0093] Call time analysis, time series analysis: Arrange calls in chronological order and analyze the time intervals and sequence between calls. If multiple calls occur consecutively within a short period of time, or follow a specific time pattern (such as daily, weekly, etc.), then these calls may be related.

[0094] In some embodiments of this application, each associated call event is evaluated from multiple dimensions to obtain evaluation metrics, including...

[0095] Collect manually recorded information of related call events during the call center's duty process, compare the matching of the manually recorded information with the text content and key information of the related call events, and generate the first evaluation index.

[0096] Analyze the tone of voice in each voice message in the associated call events to generate a second evaluation metric;

[0097] Analyze the task delivery of operators in related call events to generate a third evaluation indicator.

[0098] In this embodiment, during their shifts, operators record the instructions or requests they hear. The matching of manually recorded information with the text content and key information is converted into a first evaluation indicator, essentially assessing the degree to which the operator's recorded content matches the actual content. The tone of the operator's voice reflects service attitude and is converted into a second evaluation indicator. Based on key information and the operator's relaying of information (the text content of the call made by the operator), the third evaluation indicator is derived by analyzing the operator's task delivery in related call events.

[0099] It is understood that this may also include other dimensions that can evaluate the quality of call center operators' performance, all of which fall within the scope of protection of this application.

[0100] Step S104: Integrate the evaluation indicators of all related call events on the call information timeline to evaluate the duty quality of call center operators, thereby assisting in the improvement and optimization of call center operators' duty performance.

[0101] In this embodiment, the evaluation metrics of all related call events on the call information timeline are integrated (the evaluation metrics are integrated in units of time and related call events) to conduct a comprehensive evaluation of the duty performance of call agents. This can be achieved through the following formula:

[0102]

[0103] Where Te represents the overall performance quality of call center operators, m represents the number of related call events, and β1 j β2 j β3 j These are the evaluation weights of the first, second, and third evaluation indicators for the j-th associated call event, respectively, W1. j W2 j W3j Let t represent the magnitudes of the first, second, and third evaluation metrics for the j-th associated call event. j Let k2 be the completion time of the j-th associated call event. j Let j be the second constant corresponding to the j-th associated call event. This indicates the correction of the sum of the first, second, and third evaluation indicators based on the completion time of the associated call event.

[0104] Based on the evaluation results, improvements and optimizations were made to the call center staff's on-duty operations. For example, to address the issue of slow response time, call center staff training could be strengthened to improve their reaction speed; to address the issue of information accuracy, the information verification mechanism could be improved; to address the issue of communication efficiency, the call flow and information transmission methods could be optimized; to address the issue of service attitude, call center staff professional ethics training could be strengthened; and to address the issue of poor task completion, task tracking and supervision could be strengthened.

[0105] Correspondingly, this application also provides a business support system for military telephone operations, such as... Figure 2 As shown, including,

[0106] The first module is used to collect voice information generated during the calls of the military telephone operators on duty, perform multi-person voice source separation processing on the voice information to obtain single-person voice source speech, and identify the two main subjects in the call through single-person voice source speech.

[0107] The second module is used to classify calls according to the call request subject, identify the text content of each call, extract key information from the text content, and arrange the key information by timestamp to obtain the call information timeline.

[0108] The third module is used to identify the relationships between multiple calls on the call information timeline and record them as related call events. Each related call event is evaluated in multiple dimensions to obtain evaluation indicators.

[0109] The fourth module is used to integrate evaluation metrics of all related call events on the call information timeline to evaluate the performance quality of call center agents, thereby assisting in the improvement and optimization of their work.

[0110] Compared with the prior art, the beneficial effects of this invention are as follows:

[0111] 1. Perform multi-source voice separation processing on voice information to separate the voice sources of multiple people speaking simultaneously in a call. This allows for accurate analysis of each person's speech, providing a reliable foundation for subsequent voice and text analysis. Identify the text content of each call, extract key information from the text, consider the context of each call to identify precise text content, and further extract key information based on keyword occurrences to better highlight the key points of the call and reduce redundant content.

[0112] 2. Identify the relationships between multiple calls on the call information timeline and conduct multi-dimensional evaluations to unify the completion standards of operators' duties, improve the effectiveness and reliability of call service, ensure the quality of military call communication, integrate the evaluation indicators of all related call events on the call information timeline, evaluate the duty quality of call operators, thereby assisting operators in optimizing and improving their work, improving the efficiency and accuracy of call service, and ensuring the effective flow of information communication in the military.

[0113] Through the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented in hardware or by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) and includes several instructions to cause a computer device (such as a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0114] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of a preferred embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing the present invention.

[0115] Those skilled in the art will understand that the modules in the system of the implementation scenario can be distributed throughout the system of the implementation scenario as described, or they can be modified to reside in one or more systems different from this implementation scenario. The modules of the above-mentioned implementation scenario can be merged into one module, or they can be further divided into multiple sub-modules.

[0116] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for assisting military telephone operator duties, characterized in that, include, The system collects voice information generated during calls by military telephone operators while on duty, performs multi-person voice source separation processing on the voice information to obtain single-person voice source speech, and identifies the two main subjects in the call through single-person voice source speech. Calls are categorized based on the call requester, the text content of each call is identified, key information is extracted from the text content, and the key information is arranged by timestamps to obtain a call information timeline. Identify the relationships between multiple calls on the call information timeline and record them as related call events. Evaluate each related call event from multiple dimensions to obtain evaluation metrics. The evaluation metrics of all related call events on the call information timeline are integrated to evaluate the performance quality of call center agents, thereby assisting in the improvement and optimization of their work. Among them, single-person voice source speech is the voice source speech of the main object after separating the surrounding voice source speech; Before performing multi-source sound separation processing on the speech information, the method further includes, The speech information is preprocessed, including noise reduction and speech activity detection. The speech signal of the speech information is divided into multiple parts of the speech signal. The energy distribution and fundamental frequency sequence of each part of the speech signal are calculated, and the short-time energy, zero-crossing rate and peak count are statistically analyzed. The short-time energy, zero-crossing rate and peak count are compared with their respective preset thresholds to obtain three deviations. The overall deviation of each part of the speech signal is obtained by combining the three deviations. The speech features of each part of the speech signal are extracted and input into the preset overlap detection model. The overlap detection model outputs the overlap probability. The probability threshold corresponding to each part of the speech signal is set according to the overall deviation. The overlap probability of each part of the speech signal is compared with the probability threshold to detect the overlapping part of the speech signal. The overlapping part of the speech signal is then processed for multi-person sound source separation.

2. The operational support method for military telephone operations according to claim 1, characterized in that, Multi-source speech separation processing is performed on the overlapping speech signals, including, Blind source separation is used to separate multiple speech sources in overlapping speech signals. The blind source separation method includes independent component analysis and nonnegative matrix factorization.

3. The operational support method for military telephone operations according to claim 1, characterized in that, The system identifies the main subjects in a single-speaker voice call, including... The voiceprint features of a single voice source are compared with the voiceprint database pre-stored by the operator to identify the operator and the other party. The two parties are the operator and the other party.

4. The operational support method for military telephone operations according to claim 3, characterized in that, Calls are categorized based on the caller, and the text content of each call is identified, including: Call categories include outgoing calls and incoming calls; When the requesting party for the call is an operator, the call is an outgoing call. When the requesting party in a call is a recipient, the call is an access call. Each call is converted into preliminary text content using speech recognition technology. The text difficulty features of the preliminary text content are analyzed, and all text difficulty features on the preliminary text content are statistically analyzed. The text recognition difficulty of the text content is evaluated based on all text difficulty features. The difficulty of text recognition based on text content is set by adjusting the context window size of the NLP context understanding model and realizing the specific text recognition of the context of the call, thereby recognizing the text content of each call.

5. The operational support method for military telephone operations according to claim 1, characterized in that, Extract key information from the text content, including: The text content is tagged with parts of speech, and the TF and IDF of words in the text content are calculated to filter out keywords. The text content is then semantically analyzed using keywords to extract key information.

6. The operational support method for military telephone operations according to claim 4, characterized in that, Identify the relationships between multiple calls on the call information timeline. include, Association rules are constructed based on the two parties involved in different calls, time windows, similarity of key information, and the categories of outgoing and incoming calls. Calls are used as nodes, and the strength of association rules for different calls are used as edges to construct a call information association graph. Outgoing and incoming calls are represented by different colors, and graph theory algorithms are used to identify the associations between multiple calls.

7. The operational support method for military telephone operations according to claim 1, characterized in that, Each related call event is evaluated from multiple dimensions to obtain evaluation metrics, including: Collect manually recorded information of related call events during the call center's duty process, compare the matching of the manually recorded information with the text content and key information of the related call events, and generate the first evaluation index. Analyze the tone of voice in each voice message in the associated call events to generate a second evaluation metric; Analyze the task delivery of operators in related call events to generate a third evaluation indicator.

8. A business support system for military telephone operations, characterized in that, The system, applied to the operational support method for military telephone operations as described in any one of claims 1-7, comprises: The first module is used to collect voice information generated during the calls of the military telephone operators on duty, perform multi-person voice source separation processing on the voice information to obtain single-person voice source speech, and identify the two main subjects in the call through single-person voice source speech. The second module is used to classify calls according to the call request subject, identify the text content of each call, extract key information from the text content, and arrange the key information by timestamp to obtain the call information timeline. The third module is used to identify the relationships between multiple calls on the call information timeline and record them as related call events. Each related call event is evaluated in multiple dimensions to obtain evaluation indicators. The fourth module is used to integrate evaluation metrics of all related call events on the call information timeline to evaluate the performance quality of call center agents, thereby assisting in the improvement and optimization of their work.

Citation Information

Patent Citations

  • Intelligent evaluation method and system for service quality of telephone customer service

    CN116828109A

  • Conversation reminding method and system of customer service system

    CN118803134A