Intelligent speech analysis system based on multi-agent
The multi-agent intelligent voice analysis system solves the problem of not being able to obtain information about a single subject from long conversations in business activities, and achieves high-precision extraction and analysis of conversation content, supporting the real-time and reliable nature of business decisions and providing valuable business insights.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HANGZHOU HUMPBACK WHALE TECHNOLOGY CO LTD
- Filing Date
- 2026-01-05
- Publication Date
- 2026-05-01
AI Technical Summary
In business activities, it is impossible to accurately extract valuable information from individual subjects in long dialogue texts, resulting in information being discarded or lost in recording, making it impossible to calculate its value.
The intelligent voice analysis system employs a multi-agent architecture, including modules for multimodal data acquisition and conversion, intelligent semantic unit segmentation, subject identification and association integration, industry-customized analysis, and visualization decision support. Through speech recognition, semantic boundary detection, subject aggregation, and industry-specific analysis, it achieves high-precision extraction and analysis of dialogue content.
It achieves sentence-level precision information extraction from dialogue text, enabling customized analysis solutions based on the needs of different business entities, improving the real-time performance and reliability of business decisions, reducing computational latency, and providing valuable customer insights and employee behavior analysis.
Smart Images

Figure CN121459795B_ABST
Abstract
Description
Multi-Agent-Based Intelligent Voice Analysis System Technical Field
[0001] This invention relates to the field of speech analysis technology, specifically to a multi-agent-based intelligent speech analysis system. Background Technology
[0002] Currently, business activities refer to a company's purchasing, sales, exchange, and bank lending activities. In addition to strategy and decision-making, business capabilities should also include long-term forecasting capabilities. Business activities are an important part of economic activities and are important factors in commodity flow and social stability. Business activities reflect economic strength. In current business activities, in common scenarios such as sales, meetings, interviews, and medical visits, dialogues often carry complex and high-value information. Voice analysis is a technology that processes voice data to extract information. Its core value lies in improving service efficiency, optimizing customer experience, and mining business value.
[0003] In traditional business activities, dialogues usually revolve around one or more subjects, such as customers, interviewees, or patients. Because it is impossible to accurately obtain valuable information about a single subject from long dialogue texts, this information is usually discarded or recorded by individuals with losses, resulting in an unquantifiable loss of value. Summary of the Invention
[0004] This invention provides a multi-agent-based intelligent voice analysis system, which can effectively solve the problem mentioned in the background art that the valuable information of a single subject cannot be accurately obtained in long-term dialogue texts. This information is usually discarded or recorded by individuals with loss, resulting in the inability to statistically measure the value loss.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a multi-agent-based intelligent voice analysis system, characterized in that: multi-agent collaborative technology is applied to the processing of dialogue content in sales, meeting and interview scenarios to achieve information extraction and analysis with sentence-level precision;
[0006] It includes a multimodal data acquisition and conversion module, a semantic unit intelligent segmentation module, a subject identification and association integration module, an industry-customized analysis module, a visualization decision support module, and an operation improvement and algorithm optimization module;
[0007] The multimodal data acquisition and conversion module acquires high-quality, structured initial text data, including a data acquisition unit, a speech recognition unit, a structured processing unit, and a data preprocessing unit.
[0008] According to the above technical solution, the data acquisition unit is used in commercial scenarios to collect audio streams in real time by deploying smart terminal devices;
[0009] The speech recognition unit uses an advanced speech model to perform multilingual speech recognition and converts the audio stream into structured dialogue text.
[0010] The structured processing unit records key metadata simultaneously while converting the text. The key metadata includes timestamps and speaker IDs.
[0011] The data preprocessing unit performs noise reduction, segmentation, and format standardization on the original audio stream. Combining the ASR results with voiceprint information, it initially constructs the association between speaker ID, text, and time to generate an original dialogue dataset containing speaker ID, text, and timestamp.
[0012] According to the above technical solution, the semantic unit intelligent segmentation module generates independent minimum paragraphs with complete meaning from continuous and lengthy dialogue text through semantic boundary detection, ensuring that each paragraph contains complete semantic expression.
[0013] According to the above technical solution, the subject identification and association integration module re-aggregates semantic units that are scattered at different time points but belong to the same subject, forming a complete dialogue record for each subject, generating an associated dialogue text for each identified subject, and integrating all related speech segments together to provide a data foundation for subsequent individualized analysis.
[0014] It includes multidimensional fusion clustering units and algorithm optimization units.
[0015] According to the above technical solution, the multidimensional fusion clustering unit combines multiple types of information for dynamic clustering, rather than relying on a single technology. Specifically, it includes voiceprint ID, title recognition, and temporal continuity.
[0016] According to the above technical solution, the industry-customized analysis module extracts key information and quantitative indicators from the aggregated main dialogue based on the needs of a specific industry, including a demand mapping unit, a deep analysis unit, and a confidence calibration unit.
[0017] The demand mapping unit is a configurable industry analysis adapter provided by the system. Based on the industry template selected by the user, the system automatically loads the corresponding analysis dimensions and indicators.
[0018] According to the above technical solution, the visualization decision support module presents the analysis results in an intuitive and multi-dimensional visualization form, specifically generating visualization outputs in three dimensions: macro trend charts, micro insight tables, and early warning dashboards.
[0019] The macro trend chart uses visual charts to show the proportion of discussion time for different topics in the dialogue and their relationship over time, intuitively showing the flow and focus of topics;
[0020] The micro-level insight table lists the top 5 most frequently used technical terms and keywords, along with their associated business entities, providing specific data support.
[0021] The early warning dashboard automatically triggers red warning icons for detected abnormal semantic patterns by monitoring and analyzing results in real time, enabling immediate risk warnings during and after events.
[0022] According to the above technical solution, the operation improvement and algorithm optimization module includes internal operation improvement and improvement strategies for determining different operation scenarios. The internal operation improvement includes ASR and multi-Agent collaboration optimization, semantic segmentation optimization, subject aggregation optimization, and industry analysis adapter optimization.
[0023] According to the above technical solution, the ASR and multi-Agent collaborative optimization adopts incremental speech recognition and multi-Agent parallel processing, processes long audio segments to reduce latency, and uses the semantic segments before and after to correct the current recognition result.
[0024] According to the above technical solution, determining the improvement strategy for different operating scenarios means that different improvement strategies need to be determined in different operating scenarios. The specific operating scenarios and improvement strategies are as follows:
[0025] In scenarios where multiple people speak rapidly and alternately, the specific improvement measures are: voiceprint + semantic continuity matching, and parallel ASR of overlapping audio segments;
[0026] In scenarios with long periods of invalid speech / blank segments, the specific improvement measures are: use speech activity detection to remove invalid segments;
[0027] In scenarios involving industry template updates or new metrics, specific improvement measures include: dynamically loading analysis templates and hot-updating analysis agents;
[0028] In scenarios involving abnormal semantics or sudden events, specific improvement measures include: triggering immediate alerts for abnormal keywords and activating the anomaly analysis agent.
[0029] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0030] 1. By utilizing a large language model to process the ASR text of business activities, the text is first subdivided into dialogue segments based on features such as semantic continuity. Then, segments related to the same subject are aggregated into related dialogue texts of that subject. Information can be extracted and processed from these related dialogue texts of different business entities according to their requirements, resulting in valuable customer insights and employee behavior analysis. During the analysis process, information can be extracted at the word level from the dialogue text, and quantifiable analysis results can be obtained from the entire text. This enables the full mining of the business value contained in each recording, helping to build the enterprise's resource barriers and core competitiveness.
[0031] It enables more flexible processing of ASR text to accurately identify related dialogue segments between different subjects and discard worthless text segments. For the extracted important segments, different analysis solutions can be customized for business entities in different fields to meet the needs of diverse business activities and help users build information barriers.
[0032] 2. With the support of high-accuracy ASR technology and high-availability large language model capabilities, this invention enables businesses to accurately extract the information they need from the audio of business activities in order to improve customer service and gain insights into employee behavior. This allows multi-agent collaborative technology to be applied to the processing of dialogue content in sales, meetings and interview scenarios, achieving information extraction and analysis with sentence-level accuracy. In actual operation, it can adapt to different scenarios and ensure the accuracy of recognition, aggregation and analysis.
[0033] Furthermore, multi-agent collaboration and incremental processing reduce computational latency for long dialogues. Dynamic industry adapters and online learning capabilities enable the system to automatically optimize analysis strategies as business metrics change, provide timely warnings for abnormal semantics, and enhance the real-time performance and reliability of business decisions. Attached Figure Description
[0034] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.
[0035] In the attached diagram:
[0036] Figure 1 is a structural block diagram of the analysis system of the present invention. Detailed Implementation
[0037] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0038] Example: As shown in Figure 1, the present invention provides a technical solution, an intelligent voice analysis system based on multi-agent, which applies multi-agent collaborative technology to the processing of dialogue content in sales, meetings and interview scenarios, and achieves information extraction and analysis with sentence-level precision;
[0039] It includes a multimodal data acquisition and conversion module, a semantic unit intelligent segmentation module, a subject identification and association integration module, an industry-customized analysis module, a visualization decision support module, and an operation improvement and algorithm optimization module;
[0040] The multimodal data acquisition and conversion module acquires high-quality, structured initial text data, including a data acquisition unit, a speech recognition unit, a structured processing unit, and a data preprocessing unit.
[0041] Based on the above technical solution, the data acquisition unit is used in commercial scenarios to collect audio streams in real time by deploying smart terminal devices. Commercial scenarios include physical stores and conference rooms, and smart terminal devices include microphones, voice recorders, and cameras.
[0042] The speech recognition unit uses an advanced speech model for multilingual speech recognition. The speech model uses the Transformer architecture to convert audio streams into structured dialogue text.
[0043] The structured processing unit records key metadata simultaneously while converting text. Key metadata includes timestamps and speaker IDs. Timestamps refer to the start and end times of each statement, accurate to the millisecond level, and speaker IDs are used to distinguish different speakers through voiceprint recognition technology.
[0044] The data preprocessing unit performs noise reduction, segmentation, and format standardization on the original audio stream to improve the efficiency of subsequent processing. Combining the results of automatic speech recognition with voiceprint information, it initially constructs the relationship between speaker ID, text, and time to generate an original dialogue dataset containing speaker ID, text, and timestamp, laying the foundation for subsequent in-depth analysis.
[0045] Based on the above technical solution, the semantic unit intelligent segmentation module generates independent minimum paragraphs with complete meaning from continuous and lengthy dialogue text through semantic boundary detection, ensuring that each paragraph contains complete semantic expression;
[0046] Semantic boundary detection uses natural language processing technology to analyze the semantic coherence and topic transition points in text, identify natural semantic boundaries, and solve the problem of semantic boundary ambiguity. It calculates the semantic change score of each word position and accurately determines paragraph boundaries by comparing it with a preset threshold.
[0047] Specifically, long conversations are divided into a series of smallest semantic units to ensure that each unit expresses a relatively independent and complete idea and information point, thus avoiding information fragmentation and semantic breaks.
[0048] Based on the above technical solution, the subject identification and association integration module re-aggregates semantic units that are scattered at different time points but belong to the same subject, forming a complete dialogue record for each subject, generating an associated dialogue text for each identified subject, and integrating all related speech segments together to provide a data foundation for subsequent individualized analysis.
[0049] It includes multidimensional fusion clustering units and algorithm optimization units.
[0050] Based on the above technical solutions, the multidimensional fusion clustering unit combines multiple types of information for dynamic clustering, rather than relying on a single technology. Specifically, it includes voiceprint ID, title recognition, and temporal continuity.
[0051] Voiceprint ID serves as the primary basis for biometrics; title recognition identifies the names and titles used in conversations; and temporal continuity refers to the high probability that the same speaker will speak consecutively within a short period of time.
[0052] The algorithm optimization unit uses the subject matching probability formula to comprehensively calculate the voiceprint matching degree, title matching degree and time proximity, and obtains a comprehensive probability to determine whether different paragraphs belong to the same subject, thus solving the problems of title change and voiceprint recognition error.
[0053] Finally, a related dialogue text is generated for each identified subject, integrating all related speech segments to provide a data foundation for subsequent individualized analysis.
[0054] Based on the above technical solution, the industry-customized analysis module extracts key information and quantitative indicators from the aggregated main dialogue according to the needs of a specific industry, including a demand mapping unit, a deep analysis unit, and a confidence calibration unit.
[0055] The demand mapping unit is a configurable industry analysis adapter provided by the system. Based on the industry template selected by the user, which includes retail, finance, and manufacturing, the system automatically loads the corresponding analysis dimensions and indicators. Specifically, the retail industry focuses on the sensitivity of average order value, while the financial industry focuses on risk warnings.
[0056] The deep analysis unit performs deep analysis based on a template by analyzing the Agent, as detailed below:
[0057] For the financial industry, the focus is on extracting risk warning statements. A specialized FinBERT model is used to identify and extract risk warning statements containing keywords related to volatility and leverage.
[0058] For the manufacturing industry, the frequency of mention of statistical technical parameters, specifically the number of repetitions and frequency of mention of technical parameters with a tolerance of ±0.1mm;
[0059] The confidence calibration unit assigns a confidence score to each output analysis result. The analysis results include extracted viewpoints and statistically determined frequencies to enhance the reliability and interpretability of the analysis results. The confidence score is based on the completeness of the statement and the support of data, while the support of data refers to whether there is supporting data.
[0060] Based on the above technical solution, the visualization decision support module presents the analysis results in an intuitive and multi-dimensional visualization form to help managers quickly understand and make decisions. Specifically, it generates visualization outputs in three dimensions: macro trend charts, micro insight tables, and early warning dashboards.
[0061] The macro trend chart is a visualization chart using the Sankey diagram to show the proportion of discussion time for different topics in the dialogue and their relationship over time. It intuitively shows the flow and focus of topics, including product features, prices and services.
[0062] The micro-level insight table lists the top 5 most frequently used technical terms and keywords, along with their associated business entities, providing specific data support.
[0063] The early warning dashboard automatically triggers red warning icons for detected abnormal semantic patterns by monitoring and analyzing results in real time, enabling immediate risk warnings during and after events.
[0064] Based on the above technical solution, the operation improvement and algorithm optimization module includes internal operation improvement and improvement strategies for different operation scenarios. The internal operation improvement includes ASR and multi-Agent collaboration optimization, semantic segmentation optimization, subject aggregation optimization, and industry analysis adapter optimization.
[0065] Based on the above technical solutions, the ASR and multi-Agent collaboration optimization adopts incremental speech recognition and multi-Agent parallel processing, processes long audio segments to reduce latency, solves the problem of ASR recognition delay and omission that may be caused by long dialogues or multiple people speaking at the same time, and uses the semantic segments before and after to correct the current recognition result, thereby improving recognition efficiency and accuracy.
[0066] Semantic paragraph segmentation optimization introduces semantic boundary detection formulas and clustering algorithm formulas to improve paragraph segmentation accuracy and solve the problem of fuzzy semantic boundaries, which leads to paragraphs that are too long or too short.
[0067] The specific formula for paragraph boundary detection is as follows:
[0068] ;
[0069] in, , indicating the first Whether a word marks a paragraph boundary is indicated by 1 (yes) or 0 (no). For words The semantic change score at the location. The preset judgment threshold;
[0070] The specific formula for clustering paragraphs with fuzzy boundaries is as follows:
[0071] in, , indicating paragraph and Overall similarity between them;
[0072] Paragraph and The semantic similarity is specifically based on cosine similarity of the embedded vectors;
[0073] Paragraph and The continuity similarity over time is specifically a decay function based on location distance;
[0074] It is an adjustable weight hyperparameter used to balance the relative importance of semantic similarity and temporal continuity in the overall similarity;
[0075] Subject aggregation optimization uses the subject matching probability formula, combined with voiceprint and semantic features, to avoid misjudgment of the same subject due to changes in title and voiceprint noise, thereby reducing misjudgment;
[0076] The specific formula for the subject matching probability is as follows:
[0077]
[0078] in, Indicates sample and The overall similarity or attribution probability of belonging to the same semantic unit, specifically the same speaker paragraph or the same topic block;
[0079] These represent the normalized similarity in three dimensions: linguistic features, title semantics, and temporal continuity, respectively.
[0080] For learnable weights, satisfying ;
[0081] Industry analysis adapter optimization addresses the issue of significant differences in terminology and indicators across different industries by dynamically loading industry keyword libraries and weight parameters, adapting to different industry needs and supporting rapid adaptation to new scenarios.
[0082] Keyword weights are calculated as follows:
[0083] ;
[0084] in, Indicates the first The overall weight of each word (or term, keyword), For words Normalized term frequencies, specifically TF and TF-IDF normalized values. For words The semantic or contextual influence score is specifically based on attention weights, centrality metrics, and domain importance measures.
[0085] The weight coefficients are non-negative and satisfy the normalization constraint: .
[0086] Based on the above technical solution, it is necessary to determine improvement strategies for different operating scenarios. Specific operating scenarios and improvement strategies are as follows:
[0087] In scenarios where multiple people speak rapidly and alternately, the specific improvement measures are: voiceprint + semantic continuity matching, and parallel ASR of overlapping audio segments;
[0088] In scenarios with long periods of invalid speech / blank segments, the specific improvement measures are: using Voice Activity Detection (VAD) to remove invalid segments;
[0089] In scenarios involving industry template updates or new metrics, specific improvement measures include: dynamically loading analysis templates and hot-updating analysis agents;
[0090] In scenarios involving abnormal semantics or sudden events, specific improvement measures include: triggering immediate alerts for abnormal keywords and activating the anomaly analysis agent.
[0091] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A multi-agent-based intelligent voice analysis system, characterized in that: This invention applies multi-agent collaborative technology to dialogue content processing in sales, meetings, and interview scenarios, achieving sentence-level precision information extraction and analysis. It includes a multimodal data acquisition and conversion module, a semantic unit intelligent segmentation module, a subject identification and association integration module, an industry-customized analysis module, a visualization decision support module, and an operation improvement and algorithm optimization module. The multimodal data acquisition and conversion module acquires structured initial text data, including a data acquisition unit, a speech recognition unit, a structured processing unit, and a data preprocessing unit. The data acquisition unit, in a commercial scenario, uses intelligent terminal devices to collect audio streams in real time. The speech recognition unit employs an advanced large-scale speech model for multilingual speech recognition, converting the audio stream into structured dialogue text. The structured processing unit simultaneously records key metadata, including timestamps and speaker IDs, while converting the text. The data preprocessing unit performs noise reduction, segmentation, and format standardization on the original audio stream. Combining ASR results and voiceprint information, it initially constructs the association between speaker ID, text, and time to generate an original dialogue dataset containing speaker ID, text, and timestamps. The semantic unit intelligent segmentation module generates independent minimum segments with complete meaning from continuous and lengthy dialogue text through semantic boundary detection, ensuring that each segment contains complete semantic expression. The subject identification and association integration module re-aggregates semantic units scattered at different time points but belonging to the same subject, forming a complete dialogue record for each subject. It generates an associated dialogue text for each identified subject and integrates all related speech segments to provide a data foundation for subsequent individualized analysis. It includes a multi-dimensional fusion clustering unit and an algorithm optimization unit; the multi-dimensional fusion clustering unit combines multiple types of information for dynamic clustering, rather than relying on a single technology, specifically including voiceprint ID, title recognition, and temporal continuity; the algorithm optimization unit uses the subject matching probability formula to comprehensively calculate the voiceprint matching degree, title matching degree, and temporal proximity, and obtains a comprehensive probability to determine whether different segments belong to the same subject.
2. The intelligent voice analysis system based on multiple agents according to claim 1, characterized in that: The industry-customized analysis module extracts key information and quantitative indicators from the aggregated main dialogues based on the needs of a specific industry. It includes a demand mapping unit, a deep analysis unit, and a confidence calibration unit. The demand mapping unit is provided by the system through a configurable industry analysis adapter. Based on the industry template selected by the user, the system automatically loads the corresponding analysis dimensions and indicators.
3. The intelligent voice analysis system based on multiple agents according to claim 1, characterized in that: The visualization decision support module presents the analysis results in an intuitive and multi-dimensional visual format, specifically generating three dimensions of visualization output: macro trend charts, micro insight tables, and early warning dashboards. The macro trend chart uses visual charts to show the proportion of discussion time for different topics in the dialogue and their evolution over time, intuitively displaying the flow and focus of topics. The micro insight table lists the top 5 high-frequency technical terms, keywords, and their associated business entities, providing specific data support. The early warning dashboard monitors the analysis results in real time and automatically triggers red warning icons for detected abnormal semantic patterns, achieving immediate risk warnings during and after the event.
4. The intelligent voice analysis system based on multiple agents according to claim 1, characterized in that: The operation improvement and algorithm optimization module includes internal operation improvement and improvement strategies for different operation scenarios. The internal operation improvement includes ASR and multi-Agent collaboration optimization, semantic segmentation optimization, subject aggregation optimization, and industry analysis adapter optimization.
5. The intelligent voice analysis system based on multiple agents according to claim 4, characterized in that: The ASR and multi-Agent collaborative optimization adopts incremental speech recognition and multi-Agent parallel processing, processes long audio segments to reduce latency, and uses preceding and following semantic segments to correct the current recognition result.
6. The intelligent voice analysis system based on multiple agents according to claim 4, characterized in that: The improvement strategies for different operating scenarios are as follows: Different improvement strategies need to be determined for different operating scenarios. Specific operating scenarios and improvement strategies are as follows: In a scenario where multiple people rapidly alternate speaking, the specific improvement measures are: voiceprint + semantic continuity matching, and parallel ASR for overlapping audio segments; In a scenario with long periods of invalid speech / blank segments, the specific improvement measures are: using voice activity detection to remove invalid segments; In a scenario with industry template updates or new metrics, the specific improvement measures are: dynamically loading analysis templates and hot-updating the analysis agent; In a scenario with abnormal semantics or sudden events, the specific improvement measures are: triggering immediate alerts for abnormal keywords and activating the anomaly analysis agent.
Citation Information
Patent Citations
Film and television play table book extraction method and device, storage medium and computer equipment
CN121121616A