Intelligent voice comparison and identification method, device, equipment and medium

By constructing voice comparison tasks, processing voice information, extracting voice feature data, and calling voiceprint comparison model for comparison processing, the problem of incoherent and low degree of automation of voice comparison assisted identification in the prior art is solved, and efficient and accurate voice comparison and report generation is achieved.

CN119993205APending Publication Date: 2025-05-13PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510146059.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art has incoherent processes, low automation, insufficient processing efficiency and comparison accuracy in the voice comparison aid identification process, especially when processing multi-source voice data processing, data inconsistency and processing delay are prone to problems.

Method used

By constructing a voice comparison task, processing voice information, extracting voice feature data, and calling voiceprint comparison model for comparison processing, generating comparison results and generating an appraisal report, we realize efficient processing of voice comparison assisted identification.

Benefits of technology

It realizes efficient processing of voice comparison assisted identification, reduces manual intervention, improves the accuracy of voice comparison and the efficiency of report generation, and is suitable for multiple voice data source scenarios, improving the intelligence level of business processes and the reliability of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993205A_ABST
    Figure CN119993205A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, the field of medical health and the field of financial science and technology, and discloses an intelligent voice comparison and identification method, which comprises the steps of constructing a voice comparison task, processing voice information to generate processed voice data, extracting voice feature data from the processed voice data, and sending the voice feature data to a server; and calling the voiceprint comparison model to carry out comparison processing to generate a comparison result, generating an identification report based on the comparison result, and sending the identification report to the target processing end. According to the invention, through automation of voice task construction, voice data processing, feature extraction, comparison processing and report generation processes, efficient processing of voice comparison auxiliary identification is realized, manual intervention is reduced, the accuracy of voice comparison and the efficiency of report generation are improved, the method is suitable for various voice data source scenes, and the user experience is improved. Therefore, the intelligent level of the business process and the reliability of data processing are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence technology, medical health and financial technology, and in particular to an intelligent voice comparison and identification method, device, equipment and storage medium. Background Art

[0002] In the technical field of voice matching and assisted identification, existing technologies generally rely on manual operations, and the degree of process automation is low. For example, the format conversion, editing, data uploading and comparison result collation of voice files all need to be processed manually. This operation mode not only increases the workload, but also easily leads to human errors, making it difficult to meet the business needs of large-scale voice data processing and rapid response.

[0003] In addition, the existing technology is not efficient enough in the preprocessing and comparison of voice data, especially when processing voice files of various formats and source channels, data inconsistency and processing delays are prone to occur. Especially in the analysis of voice data in the field of medical health, the patient's voice information may come from a variety of scenarios, such as telephone consultations, clinic recordings, remote online consultations, etc. These voice data vary greatly in format, sound quality, channel, etc., and require complex preprocessing operations to ensure data consistency and reliability. However, the processing efficiency of these multi-source voice data in the existing technology is low, which can easily lead to delays in the analysis of patients' voice features, affect the accuracy of diagnosis results, and thus have a negative impact on the quality of medical services.

[0004] Similarly, in identity authentication scenarios in the financial sector, the sources of voice data are equally diverse, including customer phone recordings, online recording uploads, and self-service voice input. The format and sound quality of these recording data may be limited by factors such as device differences, recording environment, and channel transmission, making it difficult for existing voice matching models to cope with them. This limitation will cause the accuracy of identity authentication to decrease in high-risk business scenarios, such as account unlocking, transaction authorization, and complaint verification, which not only increases business processing time, but may also bring challenges to customer experience and business compliance.

[0005] The applicability of current voiceprint comparison models is also limited, especially in voice data comparisons of non-telephone channels, where the accuracy of the comparison results is low. Since voice files may be compressed, there may be channel differences or audio quality degradation, misjudgment is prone to occur when comparing voiceprint models, which affects the accuracy and stability of key business scenarios such as customer identity verification and complaint verification.

[0006] In addition, the existing technology of identification report generation and feedback process is too dependent on manual operation, and the steps of report formatting, packaging and feedback confirmation are complicated and inefficient. Such report generation method is easy to affect the timely transmission of information and business processing efficiency in regulatory feedback in the financial field and diagnostic reports in the medical and health field, increasing management costs and error risks. Summary of the invention

[0007] The main purpose of the present invention is to provide an intelligent voice comparison and identification method, device, equipment and storage medium, aiming to solve the technical problems of the prior art in the process of voice comparison assisted identification, such as inconsistent process, low degree of automation, and insufficient processing efficiency and comparison accuracy.

[0008] To achieve the above object, the present invention provides an intelligent voice comparison and identification method, comprising:

[0009] Construct a speech comparison task including speech information, speech processing strategy and comparison strategy;

[0010] Processing the voice information based on the voice processing strategy to generate processed voice information;

[0011] Extracting speech feature data from the processed speech information;

[0012] Based on the comparison strategy of the voice comparison task, calling the voiceprint comparison model to perform comparison processing on the voice feature data to generate a comparison result;

[0013] An identification report is generated based on the comparison result, and the identification report is sent to a target processing end.

[0014] Furthermore, in order to achieve the above-mentioned object, the present invention provides an intelligent voice comparison and identification device, comprising:

[0015] A task construction module, used to construct a speech comparison task including speech information, speech processing strategy and comparison strategy;

[0016] A speech preprocessing module, used for processing the speech information based on the speech processing strategy to generate processed speech information;

[0017] A feature extraction module, used to extract voice feature data from the processed voice information;

[0018] A comparison processing module, used to call a voiceprint comparison model to perform comparison processing on the voice feature data based on the comparison strategy of the voice comparison task, and generate a comparison result;

[0019] The report generation and sending module is used to generate an identification report based on the comparison result and send the identification report to the target processing end.

[0020] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer device, which includes a memory, a processor, and an intelligent voice comparison and identification program stored in the memory and executable on the processor, and the intelligent voice comparison and identification program, when executed by the processor, implements the steps of the intelligent voice comparison and identification method as described above.

[0021] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which an intelligent voice comparison and identification program is stored, and when the intelligent voice comparison and identification program is executed by a processor, the steps of the intelligent voice comparison and identification method as described above are implemented.

[0022] Beneficial effects: The present invention relates to the fields of artificial intelligence technology, medical health and financial technology, and discloses an intelligent voice comparison and identification method, including: constructing a voice comparison task, processing voice information to generate processed voice data, extracting voice feature data from the processed voice data, calling a voiceprint comparison model to perform comparison processing to generate a comparison result, and generating an identification report based on the comparison result and sending it to a target processing end. The present invention realizes efficient processing of voice comparison-assisted identification by automating the processes of voice task construction, voice data processing, feature extraction, comparison processing and report generation, reduces manual intervention, improves the accuracy of voice comparison and the efficiency of report generation, and is applicable to a variety of voice data source scenarios, thereby improving the intelligence level of business processes and the reliability of data processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:

[0024] Figure 1 A schematic diagram of an application environment of an intelligent voice comparison and identification method in an embodiment of the present invention;

[0025] Figure 2 A schematic diagram of a flow chart of an embodiment of an intelligent voice comparison and identification method of the present invention;

[0026] Figure 3 A schematic diagram of functional modules of a preferred embodiment of the intelligent voice comparison and identification device of the present invention;

[0027] Figure 4 A schematic diagram of the structure of a computer device in one embodiment of the present invention;

[0028] Figure 5 FIG. 4 is another schematic diagram of the structure of a computer device in one embodiment of the present invention. DETAILED DESCRIPTION

[0029] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.

[0030] The intelligent voice comparison and identification method provided by the embodiment of the present invention can be applied in the following aspects: Figure 1 In an application environment, the user end communicates with the server end through a network. The server end can construct a voice comparison task through the user end, process the voice information to generate processed voice data, extract voice feature data from the processed voice data, call the voiceprint comparison model to perform comparison processing to generate comparison results, generate an identification report based on the comparison results and send it to the target processing end. The present invention realizes efficient processing of voice comparison-assisted identification by automating the processes of voice task construction, voice data processing, feature extraction, comparison processing and report generation, reduces manual intervention, improves the accuracy of voice comparison and the efficiency of report generation, and is suitable for a variety of voice data source scenarios, thereby improving the intelligence level of business processes and the reliability of data processing. Among them, the user end can be but is not limited to various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server end can be implemented with an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.

[0031] See also Figure 2 , Figure 2 This is a flow chart of an embodiment of the intelligent voice comparison and identification method provided by the present invention. It should be noted that although the logical order is shown in the flow chart, in some cases, the steps shown or described may be performed in a different order than that here.

[0032] like Figure 2 As shown, the intelligent voice comparison and identification method proposed by the present invention comprises the following steps:

[0033] S10, constructing a speech comparison task including speech information, speech processing strategy and comparison strategy;

[0034] In this embodiment, receiving voice information is the basis for constructing the voice comparison task. The voice information can be audio files from various sources, including telephone voice, recording files, conference voice, online voice input, etc. The file format may be different types such as wav, mp3, m4a, etc. The system needs to have multi-format compatibility when receiving voice information to ensure that voice data from various sources can be recognized and imported into the task system.

[0035] In the implementation process, the data receiving interface module can be built to realize the automatic reception and recognition of multi-format voice files, and the metadata of the files (such as format, size, duration, source channel, etc.) can be recorded and stored after reception. This step can use the RESTful interface or FTP upload method to manage voice information in combination with the database.

[0036] Parsing voice information is to analyze and identify the basic properties of the received audio files, including the format type, source channel, file length and sound quality of the voice. The purpose of parsing is to provide support for subsequent voice processing strategies and ensure that different types of voice files can be processed appropriately in subsequent processing.

[0037] In the implementation process, the audio parsing library (such as FFmpeg, SoX, etc.) can be used to automatically analyze the format and attributes of the voice file, and the parsing results can be output as a structured data format and stored in the task system for subsequent use. For the source channel, the source of the file can be recorded in the form of metadata tags.

[0038] The purpose of generating a speech comparison task ID is to uniquely identify the processing flow of each speech comparison task, ensuring that each task can be effectively managed and tracked in the system. The task ID usually contains information such as task ID, task name, creation time, task status, etc., and is associated with voice information, speech processing strategy, and comparison strategy.

[0039] During the implementation process, a unique task identifier can be automatically generated through the task management module, and the task identifier can be associated and stored with relevant voice information, processing strategies, and comparison strategies to ensure the full-process traceability and manageability of the task.

[0040] The speech processing strategy is to select appropriate preprocessing methods for different types of speech files, including format conversion, noise reduction, audio editing, etc. The purpose of the associated speech processing strategy is to ensure that when the task is executed, the optimal processing solution is automatically selected according to the attributes and source channels of the speech file, so as to improve the processing efficiency and accuracy of the speech data.

[0041] During the implementation process, multiple voice processing strategy templates can be preset through the configuration management module, and the appropriate strategy can be automatically matched according to the attributes of the voice information. For format conversion, the audio conversion tool can be called; for noise reduction processing, the audio filtering algorithm can be used; for audio editing, it can be achieved through time period marking and segmented storage.

[0042] The matching strategy is to configure the call parameters and matching thresholds of the voiceprint matching model according to different matching scenarios and business requirements. The selection of the matching strategy directly affects the accuracy of the matching results and the applicability of the model. Therefore, when constructing a voice matching task, the matching strategy needs to be associated with the task.

[0043] During the implementation process, the policy management module can be used to configure the comparison policy template, which includes the parameter configuration of model call, similarity threshold setting, comparison result classification standard, etc. When the system performs the comparison task, it will automatically call the voiceprint comparison model according to the associated comparison policy and output the comparison result.

[0044] Example description: In the field of healthcare, the voice data provided by patients during multiple visits may come from different channels, such as telephone consultations, hospital recording equipment, or online consultation platforms. After receiving these voice data, the system needs to parse and record voice files of different formats according to the source channel and format type of the data, such as identifying whether it comes from the patient's telephone recording, clinic recording, or self-service voice input device. To improve the accuracy of data processing, the system can associate preset voice processing strategies, such as giving priority to noise reduction processing for clinic recordings and performing format conversion operations for telephone recordings. In the process of data sharing between different hospitals, comparison strategies can also be associated, such as matching files according to the patient's voice characteristics to avoid confusion in patient identity information. This task management method ensures the automated reception, parsing, processing, and comparison of voice data, making the management of patient files more efficient.

[0045] In the process of customer identity authentication in the financial field, voice data may come from telephone customer service recordings, online recordings uploaded by customers, or voice self-service systems. After receiving the customer's voice file, it is necessary to parse the source channel of the voice file, such as identifying whether the voice file is uploaded by the customer through telephone recording or APP recording, and recording attribute information such as format type and audio length. In actual business processing, different voice files may be processed differently. For example, format conversion and noise reduction processing are performed first for telephone recordings, while feature extraction can be directly performed for APP recordings. In addition, in the complaint verification scenario, different comparison strategies can be configured for different complaint types, such as setting stricter comparison thresholds for high-risk complaints and setting more relaxed parameters for comparison tasks for general identity verification. This task management-based operation method can effectively reduce manual participation, improve the efficiency and accuracy of identity authentication, and reduce risks in financial business.

[0046] By constructing voice comparison tasks, the reception, analysis, processing strategy association and comparison strategy configuration of voice information are automatically completed, realizing the full process management of voice data from reception to comparison task generation, improving the processing efficiency of voice data and the degree of automation of comparison tasks, reducing manual intervention, and improving the accuracy of comparison results and the system's task management capabilities.

[0047] S20, processing the voice information based on the voice processing strategy to generate processed voice information;

[0048] In this embodiment, the detection of the voice file format type is a basic step in voice processing. By identifying the format type of the voice file, it is determined whether a format conversion operation needs to be performed. Common voice file formats include wav, mp3, m4a, etc. Files of different formats may affect the recognition accuracy during the comparison process, so it is necessary to select an appropriate processing method according to the format type.

[0049] In the implementation process, the format type of the voice file can be automatically identified through audio analysis tools (such as FFmpeg), and the format type information can be recorded in the task system as a processing basis. If the detected file format does not meet the system's preset format requirements, the system will automatically trigger the format conversion operation.

[0050] Format conversion refers to converting voice file formats that do not meet the comparison requirements into a unified format to ensure consistency and accuracy in the subsequent comparison process. For example, converting voice files in mp3 or m4a format into wav format with a sampling rate of 8kHz to meet the input requirements of the voiceprint comparison model.

[0051] During the implementation process, the audio processing tool can be called to realize the automatic format conversion operation and perform integrity check on the converted file to ensure that the converted file can be normally loaded into the subsequent feature extraction module.

[0052] Voice files may contain a variety of sound information, such as background music, environmental noise, multi-person conversations, etc. The purpose of separating voice segments is to extract clear human voice segments to improve the accuracy of comparison. This step usually includes human voice recognition and separation, and removes non-human voice parts from the audio.

[0053] In the implementation process, a machine learning-based speech separation model or audio filtering algorithm can be used to separate the human voice segment from the speech file and save it as a separate speech file. For speech files in complex environments, the audio can be filtered multiple times based on the separation results to improve the quality of human voice extraction.

[0054] Environmental noise in voice files may affect the accuracy of comparison results. Denoising refers to reducing the background noise in voice files to an acceptable range through audio filtering and noise reduction algorithms, thereby improving the extraction effect of voice features.

[0055] In the implementation process, denoising can be performed for different types of noise based on traditional frequency domain filtering algorithms (such as low-pass filters) or denoising models based on deep learning. For telephone voice, the focus can be on processing line noise; for recording files, the focus can be on removing environmental noise.

[0056] After completing the basic processing, the quality of the voice file needs to be tested to ensure that the processed voice file meets the requirements of subsequent comparison. Voice quality testing usually includes signal-to-noise ratio analysis, audio duration check, audio integrity check, etc. to ensure the clarity and integrity of the voice file.

[0057] During the implementation process, the signal processing tool can be used to analyze the signal-to-noise ratio of the processed voice file to ensure that the voice is clear and intelligible, and automatically check the duration and integrity of the audio. If it is detected that the quality of the voice file does not meet the standard, it can trigger reprocessing or notify the operator for manual intervention.

[0058] Example description: In the field of healthcare, patient consultation recordings may contain environmental noise, repeated information, or discontinuous voice fragments. To ensure the consistency of recorded data during patient file management and comparison, appropriate voice processing strategies can be selected for recording files from different sources. For example, for recording files from the clinic, denoising and voice fragment separation can be performed first; for recording files of online consultations, format conversion and quality inspection can be performed directly. By standardizing the recording files, the patient's voice data can be ensured to be more accurate and reliable in the subsequent analysis and comparison process, thereby improving the management efficiency of medical and health data.

[0059] In the customer complaint handling scenario in the financial field, customer recording files may come from telephone customer service, online recordings, or self-service voice devices. These recording files have different formats and source channels, and may contain noise, silent segments, or incomplete information. After receiving these voice files, you can select the appropriate processing strategy based on the source channel of the file. For example, for telephone customer service recording files, you can perform format conversion, denoising, and silent segment detection; for online recording files, you can focus on audio segment separation and quality detection. Through standardized processing of customer voice data, ensure that during customer identity authentication and complaint verification, the voice comparison results are more accurate, reduce the risk of misjudgment, and improve customer complaint handling efficiency.

[0060] By performing multi-step preprocessing operations on voice information according to the voice processing strategy, including format conversion, voice segment separation, denoising and quality detection, the clarity and consistency of voice files during feature extraction and comparison are improved, standardized processing of voice data is achieved, comparison errors caused by audio quality problems are reduced, and the accuracy and efficiency of voice comparison are improved.

[0061] S30, extracting voice feature data from the processed voice information;

[0062] In this embodiment, spectrum analysis is a basic step to extract key features from speech information. By analyzing the frequency distribution of speech, the frequency characteristics of speech can be identified, which are crucial for distinguishing different voice individuals. Spectrum analysis usually converts speech signals into a representation in the frequency domain to facilitate the subsequent extraction of speech feature data.

[0063] In the implementation process, Fourier transform (such as fast Fourier transform FFT) can be used to perform spectrum analysis on the speech signal to convert the time domain speech data into frequency domain data. At the same time, the sliding window technology can be used to segment the speech data to ensure that stable frequency features can be extracted in each time period.

[0064] Duration features and energy features are important indicators further extracted from frequency features. Duration features represent the duration of a speech segment and can be used to determine the continuity and rhythm of speech; energy features represent the volume changes of a speech segment and can be used to identify changes in the strength of speech.

[0065] In the implementation process, the duration feature can be determined by calculating the duration of each frequency segment, and the energy feature can be determined by accumulating the energy value of each frequency segment. At the same time, these features can be normalized to enable consistent comparison and matching between different speech segments.

[0066] The feature vector is the core data structure used for voice matching. By converting duration features and energy features into feature vectors, voice data can be represented as high-dimensional data points, which can be used to match the feature vector of the voiceprint template.

[0067] In the implementation process, a fixed-length vector can be used to represent the feature data of each speech segment. The dimension of the feature vector can be adjusted according to the requirements of the comparison model, and the data in the vector can be standardized to reduce errors caused by factors such as volume and speech speed.

[0068] Dimensionality compression is to reduce the dimension of the feature vector, thereby reducing computational complexity and increasing the speed of comparison. The purpose of dimensionality compression is to remove redundant data while retaining the core information of speech features, making the feature vector more compact and easier to process.

[0069] In the implementation process, dimensionality reduction algorithms such as principal component analysis (PCA) and linear discriminant analysis (LDA) can be used to compress high-dimensional feature vectors into low-dimensional representations. At the same time, different compression strategies can be selected according to the requirements of the comparison model to ensure that the compressed feature vector still contains sufficient speech feature information.

[0070] Speech feature data is the final data format used for comparison processing. By extracting spectral features, duration features and energy features from the processed speech information, and undergoing feature vector conversion and dimensional compression processing, the generated speech feature data is highly comparable and recognizable.

[0071] During the implementation process, the speech feature data can be stored as a structured data file (such as JSON, CSV format) to facilitate the subsequent loading and calling of the comparison model. At the same time, according to different application scenarios, the speech feature data can be divided into multiple types such as phrase features and long language features to meet different comparison requirements.

[0072] Example description: In the patient record management in the medical and health field, the patient's voice information may contain recordings from different scenarios, such as online consultations, telephone consultations, and office recordings. By performing spectral analysis and feature extraction on these recording files, vector data containing the patient's voice features can be generated. When matching multiple medical records, these feature data can be used to effectively distinguish the voice information of different patients, avoid file confusion, and improve the management accuracy and efficiency of patient records.

[0073] In identity verification and complaint verification in the financial sector, the voice information of customers usually comes from telephone recordings, online recording uploads and other channels. By performing spectrum analysis, duration feature extraction and energy feature analysis on the customer's voice data, the customer's voice feature data can be generated. During the customer identity verification process, these feature data can be compared with the existing voiceprint template to confirm the customer's identity, improve the accuracy of identity verification, and reduce the time and risk of misjudgment in complaint handling.

[0074] By extracting spectral features, duration features, and energy features from the processed speech information, converting these features into feature vectors, and then generating speech feature data through dimensional compression, efficient structured processing of speech data is achieved. By extracting highly recognizable speech feature data, the error caused by audio quality differences during the comparison process is reduced, and the accuracy of speech comparison and system processing efficiency are improved.

[0075] S40, based on the comparison strategy of the voice comparison task, calling the voiceprint comparison model to perform comparison processing on the voice feature data to generate a comparison result;

[0076] In this embodiment, one of the core contents of the matching strategy is the configuration of model calling parameters, which include the selection of matching models, format requirements of input data, setting of matching thresholds, etc. Different scenarios may require different model calling parameters. For example, in the identity authentication scenario in the financial field, a model with higher accuracy can be selected and a stricter matching threshold can be set; in the voice data matching scenario in the medical and health field, looser matching parameters can be selected to ensure a wider coverage of data matching.

[0077] During the implementation process, multiple comparison strategy templates can be preset through the strategy management module, and when a speech comparison task is generated, the appropriate model call parameters can be automatically matched according to the task type and scenario. For specific comparison tasks, parameters can also be dynamically adjusted to ensure the reliability and accuracy of the comparison results.

[0078] The call of the voiceprint comparison model is the core of the comparison process. The comparison model needs to receive the voice feature data generated from the voice feature extraction stage and match and analyze these feature data with the voiceprint template stored in the database. The comparison model is usually trained using a deep learning algorithm and has the ability to calculate the similarity of high-dimensional feature vectors.

[0079] During the implementation process, the voiceprint comparison model can be loaded through the API interface or internal module call. The voice feature data needs to be in a consistent format when loaded and is standardized to reduce the impact of input data differences on the comparison results. In actual applications, a distributed comparison architecture can be used to improve the concurrent processing capabilities of model calls.

[0080] Feature vector matching is the core calculation process performed by the comparison model. Both the voice feature data and the voiceprint template are represented in the form of feature vectors. The comparison model determines whether the voice data belongs to the same person by calculating the similarity between the two feature vectors. The result of feature vector matching is output in the form of a similarity score.

[0081] In the implementation process, algorithms such as Euclidean distance and cosine similarity can be used to match feature vectors, and the validity of the matching result can be determined based on the threshold set in the comparison strategy. In order to improve the accuracy of the match, the feature vector can be subjected to noise reduction or feature enhancement processing to reduce unnecessary interference information in the speech data.

[0082] After the comparison model completes the feature vector matching, it will generate a similarity score, which indicates the matching degree between the input voice feature data and the voiceprint template. The similarity score is usually a percentage value, and the higher the score, the higher the similarity between the two voice data.

[0083] In the implementation process, the similarity scores can be classified according to the similarity threshold set in the comparison strategy. If the similarity score is higher than the threshold, it is determined to be the same person; if it is lower than the threshold, it is determined to be different people. At the same time, the similarity score can be recorded in the log of the comparison task for subsequent analysis and auditing.

[0084] The comparison result is the final output of the comparison process, including the status of the comparison task, similarity score, comparison conclusion, etc. The comparison result can be a simple "match" or "mismatch" binary result, or it can include detailed comparison analysis information, such as the version number of the comparison model, the format of the input data, etc.

[0085] During the implementation process, the comparison results can be stored in the task management system in the form of structured data, and the relevant systems or personnel can be notified through message push or interface call. The generation process of the comparison results needs to ensure the integrity and accuracy of the data to avoid comparison errors caused by system failure or data loss.

[0086] Example description: In the field of medical health, the patient's voice information may be recorded multiple times in different medical treatment scenarios. The system calls the voiceprint comparison model to compare the voice feature data of the current recording with the voiceprint template in the historical record to determine whether it is the same patient. This comparison processing method can effectively solve the problem of patient identity confusion, ensure the accuracy of patient file management, and thus improve the quality and efficiency of medical services.

[0087] In the identity authentication scenario in the financial field, the customer's voice information may come from a variety of channels such as telephone recordings and online recording uploads. The system calls the voiceprint comparison model to match the voice feature data currently provided by the customer with the voiceprint template in the database, calculates the similarity score, and generates the identity authentication result based on the comparison strategy. In the process of verifying customer complaints, this comparison method can quickly determine whether the complaint voice and the original business voice are from the same person, effectively reducing misjudgments and improving the efficiency and accuracy of complaint handling.

[0088] By configuring the comparison strategy and calling the voiceprint comparison model, efficient comparison processing of voice feature data is achieved. Through automated feature vector matching and similarity score analysis, manual intervention is reduced and the accuracy of the comparison results is improved. At the same time, based on the dynamic adjustment capability of the comparison strategy, it can adapt to the comparison requirements of different scenarios and improve the flexibility and intelligence of the system.

[0089] S50, generating an identification report based on the comparison result, and sending the identification report to a target processing end.

[0090] In this embodiment, the comparison result usually includes similarity score, comparison conclusion, comparison time, comparison task identifier, etc. In order to generate an identification report, it is necessary to extract these core information from the comparison result so as to display key comparison data in the report for reference by the target processing end.

[0091] In the implementation process, the data structure of the comparison results can be parsed to automatically extract core information such as similarity scores and comparison conclusions. The extracted core information can be stored in a structured data table and associated with the comparison task identifier to ensure information traceability and consistency.

[0092] The content of the appraisal report needs to include comparison description, similarity analysis and comparison conclusion. The comparison description includes the comparison task overview and the comparison strategy used; the similarity analysis lists the similarity score and analysis process of the comparison in detail; the comparison conclusion is the final judgment result given according to the comparison threshold, such as "match" or "mismatch".

[0093] During the implementation process, the template generation tool can be used to automatically generate a report content template based on the extracted core information and fill in the relevant data. The template generation tool can support multiple report formats (such as PDF, HTML, DOC, etc.) to adapt to different business needs and display scenarios.

[0094] Formatting refers to the layout and structure optimization of the generated report content for easy reading and archiving. Formatting includes setting the styles of elements such as titles, paragraphs, tables, and charts to ensure that the report content is clearly structured and logically structured.

[0095] During the implementation process, you can use document generation tools (such as JasperReports, iText, etc.) to format the report content. The formatted report file can be directly used for archiving or sent to the target processing end. To improve readability, you can add visual comparison result charts during the formatting process, such as a bar chart or pie chart of similarity scores.

[0096] In order to ensure the integrity and security of the identification report file, the report file needs to be encapsulated and packaged into a data packet for transmission. The data packet usually contains the report file, metadata information of the report (such as task ID, generation time, etc.), and a checksum used to verify the integrity of the report.

[0097] During the implementation process, a data encapsulation tool (such as a ZIP compression tool or an encryption encapsulation tool) can be used to encapsulate the report file. At the same time, in order to improve the security of data transmission, the data packet can be encrypted to prevent illegal tampering or leakage during transmission.

[0098] The encapsulated data packets need to be sent to the target processing end through the communication module. The target processing end can be an internal business system, a regulatory agency's platform, or a customer-specified receiving end. The communication module needs to support multiple data transmission protocols (such as HTTP, FTP, MQ, etc.) to adapt to different network environments and business needs.

[0099] In the implementation process, the communication interface can be constructed to realize automatic data transmission with the target processing end. At the same time, in order to ensure the stability and reliability of transmission, functions such as breakpoint resume and transmission retry can be implemented in the communication module to prevent data loss or transmission failure caused by network abnormalities.

[0100] Example description: In the patient record management in the medical and health field, the hospital needs to perform voice comparison on the patient's multiple medical records to ensure the accuracy and consistency of the records. After completing the voice comparison, the system automatically generates a patient identity verification report containing the comparison description, similarity analysis and comparison conclusion. After the hospital's archive management system receives the report, it triggers the automatic seal process according to the internal seal policy. After the seal process is completed, the system will package the report file with the seal mark and send the packaged report file to the hospital's archive management department or supervision platform as the basis for verifying the patient's file. This automated seal and report transmission mechanism can effectively reduce manual intervention and improve the security and compliance of archive management.

[0101] In the customer complaint handling scenario in the financial field, banks or financial institutions need to compare the complaint voice submitted by the customer with the original business recording to determine whether it is the same customer. After completing the voice comparison, the system will generate a customer identity verification report, which contains the core information of the comparison task, the similarity score and the comparison conclusion. After the customer service system of the financial institution receives the report, it automatically triggers the seal operation of the report according to the preset seal strategy. After the seal process is completed, the system will package the report file with the seal mark and send it to the customer complaint management platform or regulatory agency through a secure communication channel. Through this automated seal and report transmission method, financial institutions can quickly submit seal reports, improve complaint handling efficiency, reduce the risk of misjudgment, and enhance communication compliance with regulatory agencies.

[0102] By extracting core information from the comparison results, generating an identification report, performing formatting, encapsulating the report file into a data package and sending the data package to the target processing end, the full process management of automatic processing of the comparison results and report generation is realized. By reducing manual operations, the efficiency of report generation and transmission is improved, ensuring that the comparison results can be quickly and accurately delivered to the target processing end, and improving the automation of business processing and data security.

[0103] The present invention relates to the fields of artificial intelligence technology, medical health and financial technology, and discloses an intelligent voice comparison and identification method, including: constructing a voice comparison task, processing voice information to generate processed voice data, extracting voice feature data from the processed voice data, calling a voiceprint comparison model to perform comparison processing to generate a comparison result, and generating an identification report based on the comparison result and sending it to a target processing end. The present invention realizes efficient processing of voice comparison-assisted identification by automating the processes of voice task construction, voice data processing, feature extraction, comparison processing and report generation, reduces manual intervention, improves the accuracy of voice comparison and the efficiency of report generation, and is applicable to a variety of voice data source scenarios, thereby improving the intelligence level of business processes and the reliability of data processing.

[0104] In one embodiment, the above S10 includes:

[0105] S101, receiving and parsing voice information, and determining the source channel and format type of the voice information;

[0106] S102, determining a voice processing strategy for processing the voice information based on a source channel and a format type of the voice information;

[0107] S103, configuring model calling parameters and comparison thresholds in the comparison strategy;

[0108] S104, generating a speech comparison task identifier based on the speech information, the speech processing strategy and the comparison strategy, and storing the speech comparison task identifier in association with the speech information to complete the generation of the speech comparison task.

[0109] In this embodiment, receiving voice information is the first step in building a voice comparison task, which requires the system to be compatible with voice data from multiple sources, such as telephone recordings, online recordings, uploaded audio files, etc. These voice data may be provided in different formats (such as wav, mp3, m4a), and the purpose of parsing this information is to determine the source channel and format type of the data, so as to select the appropriate strategy for subsequent voice processing.

[0110] In the implementation process, the automatic uploading and parsing of voice files can be realized through the data receiving interface. The system can use audio parsing tools (such as FFmpeg) to identify the audio format and record the source channel of the audio file, such as the telephone system, online customer service system or user upload port. The parsing results can be stored as metadata for subsequent task processing.

[0111] The determination of speech processing strategy is to select the most appropriate processing method according to the source channel and format type of speech information. Speech files of different sources and formats may require different pre-processing operations, such as format conversion, noise reduction or speech separation.

[0112] During the implementation process, the system can preset multiple voice processing strategy templates. For example, for telephone recordings, the system can choose to prioritize format conversion and noise reduction; for high-quality online recordings, the system can skip format conversion and only perform voice segment separation. The strategy selection can be automatically completed through the rule engine or policy management module to ensure that voice data is effectively processed.

[0113] The matching strategy includes the configuration of the matching model calling parameters and the similarity threshold. These configurations determine the voiceprint model, matching accuracy and similarity judgment criteria selected by the system when performing voice matching.

[0114] During the implementation process, different matching policy templates can be preset through the policy management module. The matching model call parameters can include model version number, input data format, feature vector dimension, etc.; the similarity threshold can be dynamically adjusted according to business needs. For example, a higher threshold can be set in the identity authentication scenario to ensure accuracy, while a lower threshold can be set in the patient profile matching scenario to expand the matching range.

[0115] The voice comparison task ID is used to uniquely identify each voice comparison task so that the system can track the execution status and processing results of the task. When the task ID is generated, it needs to be associated and stored with the voice information, voice processing strategy, and comparison strategy to ensure the integrity and traceability of the task.

[0116] During the implementation process, the system can automatically generate a unique task ID through the task management module and bind the task ID to the relevant voice data and policy configuration. The implementation of task identification and associated storage can use a database management system to store task data as structured records for easy query and traceability. The system can also include information such as timestamps and task types in the task identification to further improve the accuracy and flexibility of task management.

[0117] Example description: In the field of healthcare, hospitals need to perform voice comparisons on different medical records of patients to ensure the accuracy and consistency of patient files. The system receives and parses the patient's voice recording file to determine whether the recording source is a clinic recording, a telephone consultation recording, or a self-service voice input. Based on the source and format type of the voice file, the system automatically selects the corresponding voice processing strategy, such as format conversion for telephone recordings and noise reduction for clinic recordings. When generating a voice comparison task, the system configures an appropriate comparison strategy based on the business scenario, such as setting a looser comparison threshold in the patient file matching scenario to ensure a higher matching rate. Finally, the system generates a unique voice comparison task identifier, associates the task identifier with the patient's voice data and stores it for subsequent file management and query.

[0118] In the financial field, when banks handle customer complaints, they need to compare the complaint recordings submitted by customers with the original business recordings to determine whether they are from the same customer. The system receives the voice files submitted by customers, analyzes the source channel and format type of the files, and determines whether they are telephone customer service recordings or online recordings uploaded. According to different voice file types, the system selects corresponding voice processing strategies, such as format conversion and denoising for telephone recordings, and direct voice separation for high-quality files uploaded online. The system configures comparison strategies based on the business needs of complaint verification, such as setting higher similarity thresholds for high-risk complaints. Finally, the system generates a unique comparison task identifier, associates the comparison task identifier with the customer's recording data for storage, and submits the generated report to the customer complaint management system for subsequent auditing and processing.

[0119] This embodiment realizes the full process automated management of voice comparison tasks by receiving and parsing voice information, selecting appropriate voice processing strategies, configuring comparison strategies, and generating voice comparison task identifiers. The associated storage of task identifiers improves the traceability of task management and the execution efficiency of the system, effectively reduces the need for manual intervention, and improves the accuracy and processing speed of voice comparison tasks.

[0120] In one embodiment, the above S20 includes:

[0121] S201, based on the format conversion identifier in the voice processing strategy, detecting the format type of the voice information to determine whether it meets the format requirements;

[0122] S202, if the voice information does not meet the format requirement, performing a format conversion operation on the voice information based on a format conversion identifier in the voice processing strategy;

[0123] S203, analyzing the background noise intensity in the voice information based on the denoising flag in the voice processing strategy to determine whether denoising processing is required;

[0124] S204: If the voice information needs to be denoised, a denoising operation is performed on the voice information based on the denoising flag in the voice processing strategy to generate processed voice information.

[0125] In this embodiment, the detection of the format type is the first step in speech processing. By analyzing the format type of the voice file, it is determined whether it meets the input requirements of the comparison model. Common voice file formats include wav, mp3, m4a, etc. Audio files of different formats may affect the recognition accuracy during the comparison process, so it is necessary to detect the format of the voice file.

[0126] In the implementation process, an audio format parsing tool (such as FFmpeg or SoX) can be used to automatically detect the format type of the voice file. After receiving the audio file, the system parses its file header information and identifies key attributes such as format type, sampling rate, bit rate, etc. If the file format is inconsistent with the preset format requirements, the system will record the format non-compliance mark and trigger subsequent format conversion operations.

[0127] The format conversion operation is to convert the voice files that do not meet the system requirements into the standard format required by the comparison model. For example, convert files in mp3, m4a and other formats into wav format with a sampling rate of 8kHz to ensure the accuracy of voice feature extraction and comparison.

[0128] During the implementation process, the system can call audio processing tools (such as FFmpeg) to perform format conversion operations. The format conversion process includes sampling rate adjustment, bit rate adjustment, and encoding method conversion. In actual applications, the system can preset multiple format conversion templates according to the requirements of the comparison model, and automatically select the appropriate conversion strategy according to the format type.

[0129] The presence of background noise may affect the accuracy of speech comparison, so it is necessary to analyze the background noise intensity in the speech file to determine whether denoising is required. The goal of background noise analysis is to identify the noise components in the audio signal and evaluate their impact on speech feature extraction.

[0130] In the implementation process, a noise detection algorithm based on frequency domain analysis, such as short-time Fourier transform (STFT) or Mel-frequency cepstral coefficients (MFCC), can be used to analyze the frequency components of the audio file. The system can automatically determine whether denoising is required by setting a noise intensity threshold. If the background noise intensity exceeds the preset threshold, the system will mark the need for denoising.

[0131] Denoising refers to reducing the background noise components in the audio file to an acceptable range through audio filtering or noise reduction algorithms, thereby improving the extraction effect of speech features. Denoising operations usually include static noise removal, dynamic noise suppression and filtering processing.

[0132] In the implementation process, you can use traditional frequency domain filtering methods such as low-pass filters and band-pass filters, or use deep learning-based denoising models (such as RNN or CNN) for denoising. For phone recordings, you can focus on removing line noise; for conference recordings, you can remove environmental noise. The denoised audio file needs to be integrity checked to ensure that the denoising operation does not affect the recognition effect of the voice content.

[0133] This embodiment achieves standardized processing of voice information through steps such as format type detection, format conversion, background noise analysis and denoising. Through automated format conversion and denoising operations, the clarity and consistency of voice data are improved, providing more accurate input data for subsequent voice feature extraction and comparison, effectively reducing comparison errors caused by audio quality differences, and improving the system's processing efficiency and the accuracy of comparison results.

[0134] In one embodiment, the above S30 includes:

[0135] S301, performing spectrum analysis on the processed voice information to extract frequency features of the voice information;

[0136] S302, determining a duration feature and an energy feature of the voice information based on the frequency feature, and converting the duration feature and the energy feature into a feature vector;

[0137] S303, performing dimension compression processing on the feature vector, and generating speech feature data based on the feature vector after the dimension compression processing.

[0138] In this embodiment, spectrum analysis is a key step in speech feature extraction. By converting speech signals from the time domain to the frequency domain, the frequency characteristics of speech data can be captured. These frequency characteristics can reflect the changes in pitch, timbre and intensity of speech, and are an important basis for distinguishing different speech individuals.

[0139] In the implementation process, algorithms such as short-time Fourier transform (STFT) or Mel-frequency cepstral coefficient (MFCC) can be used to perform spectral analysis on speech signals. STFT divides speech signals into multiple short time periods and analyzes the frequency components in each time period, while MFCC compresses the spectral features of speech by simulating the auditory perception of the human ear, thereby extracting more stable speech features.

[0140] Duration features and energy features are key indicators further extracted from frequency features.

[0141] Duration feature: indicates the duration of a speech segment and can reflect the rhythm and speed changes of speech.

[0142] Energy feature: It indicates the volume change of the speech segment and can reflect the loudness and strength changes of the speech.

[0143] In the implementation process, the duration feature can be determined by calculating the duration of each frequency segment, and the energy feature can be determined by calculating the energy value of each frequency segment. At the same time, these features can be normalized to enable consistent comparison between different speech segments.

[0144] Feature vector is the core representation form of speech feature data. By converting duration features and energy features into high-dimensional vectors, speech data can be structured as data points for use by subsequent comparison models. Feature vector is the key data structure for feature matching during the comparison process.

[0145] In the implementation process, a fixed-length vector can be used to represent the feature data of each speech segment. The dimension of the feature vector can be adjusted according to the requirements of the comparison model, and the data in the feature vector can be standardized to reduce the errors caused by external factors such as speech speed and volume.

[0146] The purpose of dimensionality compression is to reduce the dimension of the feature vector, thereby reducing the computational complexity and increasing the processing speed of the comparison model. Dimensionality compression can remove redundant data and retain the most distinguishing feature information in the speech data.

[0147] In the implementation process, algorithms such as principal component analysis (PCA) or linear discriminant analysis (LDA) can be used to reduce the dimension of the feature vector. At the same time, different compression strategies can be selected according to the requirements of the comparison model to ensure that the compressed feature vector still contains sufficient speech recognition information. The generated speech feature data can be used for subsequent comparison processing to achieve efficient and accurate voice identity authentication.

[0148] In one embodiment, the system can select different spectrum analysis methods according to the type of voice data. For example, for telephone recording data, short-time Fourier transform (STFT) can be used to extract frequency changes in a short period of time; for long voice recordings, Mel-frequency cepstral coefficients (MFCC) can be used to extract more stable features.

[0149] In another embodiment, the feature vector conversion process can combine multiple features, including pitch features, speech speed features, and tone features, to improve the recognition accuracy of the comparison model. For example, in identity verification in the financial field, more attention can be paid to the frequency and energy features of speech; in patient file matching in the medical and health field, more attention can be paid to the duration features and rhythm changes of speech.

[0150] In another embodiment, the system can dynamically adjust the dimension of the feature vector according to the requirements of the comparison model. For example, in a high-precision comparison scenario, more feature dimensions can be retained to improve the accuracy of the comparison; in a fast comparison scenario, the comparison speed can be increased by a greater degree of dimensional compression.

[0151] This embodiment generates efficient and comparable voice feature data by performing spectrum analysis on the processed voice information, extracting frequency features, determining duration features and energy features, converting these features into feature vectors, and performing dimensional compression processing on the feature vectors. Through this automated feature extraction process, the voice data is structured and standardized, the comparison error caused by voice quality differences is reduced, and the accuracy of the comparison results and the processing efficiency of the system are improved.

[0152] In one embodiment, the above S40 includes:

[0153] S401, determining a voiceprint comparison model for performing comparison processing based on the comparison strategy of the voice comparison task;

[0154] S402, calling the voiceprint comparison model, performing feature vector matching between the voice feature data and the voiceprint template in the database, and obtaining a feature vector matching result;

[0155] S403: Analyze the similarity score between the speech feature data and the voiceprint template based on the feature vector matching result, and generate a comparison result according to the similarity score.

[0156] In this embodiment, in the voice matching task, selecting an appropriate voiceprint matching model is the core step of the matching process. Different matching strategies correspond to different model calling requirements, such as short voice model, long voice model or multilingual model. The matching strategy determines the matching model to be used according to the type and scenario of the matching task (such as identity authentication, file matching, complaint verification, etc.).

[0157] In the implementation process, multiple matching strategy templates can be preset through the strategy management module, and when matching tasks are generated, the appropriate voiceprint matching model can be automatically matched according to the type of task and business requirements. The selection of the model can be based on parameters of multiple dimensions, such as voice segment length, voice quality, input data format, etc., so as to improve the accuracy and applicability of the matching.

[0158] The core function of the voiceprint comparison model is to match the input voice feature data with the voiceprint template stored in the database. The voiceprint template is a high-dimensional feature vector generated from historical voice data, which is used to represent the voice characteristics of an individual. By matching the feature vector, it can be determined whether the two voices are from the same person.

[0159] In the implementation process, the comparison model loads voice feature data and voiceprint templates to perform high-dimensional feature vector comparison calculations. Common algorithms for feature vector matching include Euclidean distance, cosine similarity, etc. The matching process will perform filtering and sorting according to the preset comparison strategy parameters, output the matching score of each template, and obtain the feature vector matching result.

[0160] The similarity score is an indicator calculated by the comparison model based on the feature vector matching results, indicating the degree of match between the input voice feature data and the voiceprint template. The similarity score is usually expressed as a percentage or a score. The higher the score, the higher the similarity between the voice data.

[0161] In the implementation process, the calculation of the similarity score can be based on the output of the matching algorithm, and the matching values ​​of different features can be weighted to obtain a comprehensive similarity score. At the same time, different thresholds can be set according to different matching strategies to classify the similarity score as "match" or "mismatch". The analysis of the similarity score can also be combined with the business needs of the matching task. For example, in the identity authentication scenario, a higher similarity threshold can be set to reduce the risk of misjudgment.

[0162] The comparison result is the final output of the comparison process, including the status of the comparison task, similarity score, comparison conclusion, etc. The comparison result can be a simple "match" or "mismatch" binary result, or it can include detailed comparison analysis information, such as the matching template ID, comparison time, comparison model version, etc.

[0163] During the implementation process, the comparison results can be stored in the database in the form of structured data, and sent to the target processing end through message push or interface call. In order to improve the traceability of the comparison results, detailed log information of the comparison task can be included in the results for subsequent audit and analysis.

[0164] Example description: In the process of matching patient records in the medical and health field, the voice data of different medical records needs to be compared with the patient's historical voiceprint template to confirm whether it is the same patient. The system selects a suitable voiceprint comparison model according to the scenario of the comparison task, such as selecting a short voice model for online consultation data with short voice records, and selecting a long voice model for long recordings of clinic records. The system matches the feature vectors of the patient's current voice feature data with the historical voiceprint template in the database and calculates the similarity score. If the similarity score exceeds the preset threshold, the system will generate a matching comparison result and send the result to the hospital's archive management system to update the patient's archive information.

[0165] In the identity authentication and complaint verification scenarios in the financial field, the complaint recording submitted by the customer needs to be compared with the original business recording to determine whether it is the same customer. The system selects a suitable voiceprint comparison model based on the comparison strategy of the complaint task, and calls the model to load the customer's voice feature data. The system matches the customer's voice feature data with the voiceprint template in the database by feature vector and calculates the similarity score. If the similarity score exceeds the threshold, the system will generate a "matched" comparison result and send the result to the customer complaint management platform for customer service personnel to quickly verify the customer's identity. Through this automated comparison process, financial institutions can effectively improve the accuracy of identity authentication and the efficiency of complaint handling, and reduce the risk of misjudgment of customer complaints.

[0166] This embodiment realizes the full process automation of voice matching tasks through configuration of matching strategies, calling of voiceprint matching models, feature vector matching, similarity score analysis and matching result generation. By selecting the optimal matching model, the accuracy of voice matching is improved; by matching feature vectors and calculating similarity scores, manual intervention is reduced and the processing efficiency of matching tasks is improved; by automatically generating matching results, the traceability of the matching process and the compliance of data management are ensured.

[0167] In one embodiment, the above S50 includes:

[0168] S501, extracting core information including similarity scores from the comparison results;

[0169] S502, generating an identification report including a comparison description, a similarity analysis and / or a comparison conclusion according to the core information;

[0170] S503, performing formatting processing on the content of the appraisal report to generate an appraisal report file that meets the output requirements;

[0171] S504, encapsulating the identification report file into a data packet, and sending the data packet to a target processing end through a communication module;

[0172] S505, receiving reception confirmation information from the target processing end, and storing the reception confirmation information to complete the report sending process.

[0173] In this embodiment, the comparison result generally includes key information such as similarity score, comparison time, matching template ID, comparison task ID, etc. When generating an identification report, it is necessary to extract these core information from the comparison result so as to display the key content of the comparison data in the report for reference by the target processing end.

[0174] During the implementation process, the data analysis module can be used to analyze the structured data of the comparison results and extract key information including similarity scores, comparison conclusions, comparison model versions, etc. This information will serve as the basic data for report generation to ensure the accuracy and completeness of the report content.

[0175] The content of the identification report presents the results of the comparison task in a structured manner, including a comparison description (such as the purpose of the comparison task and the model calling situation), a similarity analysis (such as the similarity score and score explanation), and a comparison conclusion (such as whether it is the same voice individual).

[0176] During the implementation process, the content of the identification report can be automatically generated through the template generation tool. The system fills the preset report template based on the extracted core information, and selects different report formats and contents according to different comparison task types. For example, in the identity authentication scenario, the similarity analysis results and comparison conclusions can be highlighted; in the archive matching scenario, a more detailed comparison description and analysis process can be displayed.

[0177] In order to ensure the readability and standardization of the appraisal report content, the system needs to format the generated report content. Formatting includes setting styles for elements such as titles, paragraphs, tables, and charts to ensure that the report structure is clear and well-organized.

[0178] During the implementation process, you can use document generation tools (such as JasperReports, iText, etc.) to format the report content. The system can generate report files in different formats such as PDF, HTML, DOC, etc. according to the needs of the target processing end. During the formatting process, you can also add data charts (such as similarity score charts, comparison result pie charts, etc.) to improve the intuitiveness of the report.

[0179] In order to ensure the integrity and security of the identification report file, the system needs to encapsulate the generated report file into a data packet and send it to the target processing end through the communication module. The data packet usually contains the report file, metadata information of the report (such as task ID, generation time, etc.), and a checksum used to verify the integrity of the report.

[0180] During the implementation process, data encapsulation tools (such as ZIP compression tools or encryption encapsulation tools) can be used to encapsulate the report file. At the same time, in order to improve the security of data transmission, the data packet can be encrypted to prevent illegal tampering or leakage during transmission. The communication module can support multiple data transmission protocols such as HTTP, FTP, MQ, etc. to adapt to different transmission environments.

[0181] After the data packet transmission is completed, the target processing end needs to feedback the receiving confirmation information to the system to ensure that the transmission process of the report file is successfully completed. The receiving confirmation information can include the receiving status, feedback time, recipient signature, etc.

[0182] During the implementation process, after completing the data packet transmission, the communication module will automatically monitor the feedback information from the target processing end and store the received confirmation information in the task management system. The system can update the status of the comparison task based on the received confirmation information, for example, updating the task status to "completed" or "pending feedback", ensuring the full process traceability of the task.

[0183] This embodiment realizes the automatic report generation and transmission process of the comparison results by extracting core information from the comparison results, generating structured identification report content, formatting the report content, encapsulating the report file, and sending it to the target processing end through the communication module. By receiving feedback confirmation information from the target processing end, the transmission process of the report file is ensured to be complete and traceable, and the system's automatic processing capability and data management compliance are improved.

[0184] In one embodiment, after the above S50, the method further includes:

[0185] S601, triggering the seal usage process of the appraisal report based on the seal usage policy of the target processing end;

[0186] S602, performing a stamping process on the appraisal report and generating a report file with a stamping mark;

[0187] S603, encapsulating the report file with the seal mark, and sending the encapsulated report file to the report receiver;

[0188] S604, receiving feedback information including report receiving status and / or seal use confirmation information from the report receiver.

[0189] In this embodiment, the seal policy is a rule preset by the target processing end according to the business scenario, which determines under what circumstances the identification report is to be processed with a seal. The seal policy may include parameters such as report type, report purpose, seal permission, and seal level. For example, for reports involving identity authentication, the seal process can be automatically triggered; for general archive record reports, it can be set not to require a seal.

[0190] During the implementation process, after receiving the appraisal report, the target processing end will automatically judge according to the internal seal policy. If the seal policy requires the report to be sealed, the system will trigger the seal process and mark the task status as "in process". The management of the seal policy can be implemented through the policy management module, which supports multi-dimensional configuration by task type, report type, time, etc.

[0191] Seal processing refers to the digital or physical signature of the appraisal report to confirm the validity and authority of the report. The seal mark can be in the form of digital signature, electronic watermark, seal image, etc. The mark content includes information such as the unit using the seal, the time of use, and the seal number.

[0192] During the implementation process, the system can call the digital signature service (such as the signature interface of the CA certification agency) or the internal seal module to perform seal processing on the identification report file. After the seal processing, a new version of the report file will be generated with a seal mark. The style and content of the seal mark can be dynamically adjusted according to the seal strategy, such as selecting different seal styles or adding anti-counterfeiting watermarks.

[0193] The purpose of encapsulation is to ensure the integrity and security of the report file with the seal mark during transmission. The encapsulated report file can contain the metadata information of the report (such as task ID, seal time, seal unit, etc.), as well as the check code used to verify the integrity of the report.

[0194] During the implementation process, data encapsulation tools can be used to compress and encrypt the seal report file. After encapsulation, the system sends the data packet to the report recipient through the communication module. The communication module can support multiple transmission protocols (such as HTTP, MQ, FTP, etc.) to ensure the stability and security of data transmission.

[0195] After the report file is transmitted, the system needs to receive feedback information from the report recipient to confirm the receipt status of the report file and the validity of the seal processing. The feedback information may include the receipt status (such as "received", "not received"), the seal confirmation status (such as "confirmed", "not confirmed"), the feedback time, etc.

[0196] During the implementation process, after completing the data packet transmission, the communication module will automatically monitor the feedback information from the receiver and store the received confirmation information in the task management system. The system can update the task status based on the feedback information and generate a task log for subsequent auditing. If no feedback information is received, the system can set a retransmission mechanism or notify the administrator to intervene.

[0197] Example description: In the process of patient file verification in the medical and health field, after the hospital generates a patient file matching report, it will automatically trigger the report's seal process according to the seal policy of the file management system. The system calls the digital signature service, seals the report file, and generates a report file with the hospital seal. The sealed report file is packaged and sent to the patient management system or external supervision platform through the hospital's internal network. After confirming the report file, the recipient will feedback the "received" and "seal confirmation" status. The system will store the feedback information in the file management system to complete the closed-loop management of report transmission and seal process. Through the automated seal and report transmission process, hospitals can effectively reduce manual operations in the file verification process and improve the accuracy and compliance of file management.

[0198] In the scenarios of customer identity verification and complaint verification in the financial sector, after completing the voice matching task, the bank generates an identity verification report and automatically triggers the seal-using process according to the seal-using policy of the customer service system. The system calls the bank's internal seal-using module to electronically sign the report file and generate a PDF report file with the bank's seal logo. The sealed report file is encrypted and packaged, and then sent to the customer complaint management platform or external regulatory agency through the secure communication module. After confirming the report file, the report recipient will feedback the "received" and "seal-using confirmation" status, and the system will record the feedback information in the task log of the complaint management system. Through the automated seal-using and report transmission process, banks can effectively reduce the risk of manual operations in the customer complaint verification process, improve the efficiency and compliance of complaint handling, and improve customer service quality.

[0199] This embodiment triggers the seal-using process based on the seal-using strategy of the target processing end, performs seal-using processing on the appraisal report, encapsulates the report file with the seal-using mark and sends it to the report recipient, and receives feedback from the report recipient, thus realizing the automated management of the report seal-using process and the closed-loop control of the entire process of report transmission. Through automated seal-using processing and feedback confirmation, manual intervention is reduced, and the security, compliance and reliability of report management are improved.

[0200] In one embodiment, an intelligent voice comparison and identification device is provided, and the intelligent voice comparison and identification device corresponds one-to-one to the intelligent voice comparison and identification method in the above embodiment. Figure 3 , Figure 3This is a functional module diagram of a preferred embodiment of the intelligent speech comparison and identification device of the present invention. Task construction module 10, speech preprocessing module 20, feature extraction module 30, comparison processing module 40 and report generation and sending module 50. The functional modules are described in detail as follows:

[0201] A task construction module 10, used to construct a speech comparison task including speech information, speech processing strategy and comparison strategy;

[0202] A speech preprocessing module 20, configured to process the speech information based on the speech processing strategy to generate processed speech information;

[0203] A feature extraction module 30, used to extract voice feature data from the processed voice information;

[0204] A comparison processing module 40 is used to call a voiceprint comparison model to perform comparison processing on the voice feature data based on the comparison strategy of the voice comparison task to generate a comparison result;

[0205] The report generating and sending module 50 is used to generate an identification report based on the comparison result and send the identification report to the target processing end.

[0206] In one embodiment, the task construction module 10 is specifically used to:

[0207] Receive and analyze voice information to determine the source channel and format type of the voice information;

[0208] Determining a voice processing strategy for processing the voice information based on a source channel and a format type of the voice information;

[0209] Configure the model calling parameters and comparison thresholds in the comparison strategy;

[0210] Based on the voice information, the voice processing strategy and the comparison strategy, a voice comparison task identifier is generated, and the voice comparison task identifier is associated with the voice information and stored to complete the generation of the voice comparison task.

[0211] In one embodiment, the speech preprocessing module 20 is specifically used for:

[0212] Based on the format conversion identifier in the voice processing strategy, the format type of the voice information is detected to determine whether it meets the format requirements;

[0213] If the voice information does not meet the format requirement, performing a format conversion operation on the voice information based on the format conversion identifier in the voice processing strategy;

[0214] Based on the denoising flag in the speech processing strategy, analyzing the background noise intensity in the speech information to determine whether denoising processing is required;

[0215] If the voice information needs to be subjected to denoising processing, a denoising operation is performed on the voice information based on the denoising flag in the voice processing strategy to generate processed voice information.

[0216] In one embodiment, the feature extraction module 30 is specifically used for:

[0217] Performing spectrum analysis on the processed voice information to extract frequency features of the voice information;

[0218] Based on the frequency feature, determining the duration feature and energy feature of the voice information, and converting the duration feature and energy feature into a feature vector;

[0219] A dimension compression process is performed on the feature vector, and speech feature data is generated based on the feature vector after the dimension compression process.

[0220] In one embodiment, the comparison processing module 40 is specifically used for:

[0221] Based on the comparison strategy of the voice comparison task, determining a voiceprint comparison model for performing the comparison process;

[0222] Calling the voiceprint comparison model, performing feature vector matching between the voice feature data and the voiceprint template in the database, and obtaining a feature vector matching result;

[0223] Based on the feature vector matching result, the similarity score between the speech feature data and the voiceprint template is analyzed, and a comparison result is generated according to the similarity score.

[0224] In one embodiment, the report generating and sending module 50 is specifically configured to:

[0225] Extracting core information including similarity scores from the comparison results;

[0226] Generate an identification report including a comparison description, a similarity analysis and / or a comparison conclusion based on the core information;

[0227] Performing formatting processing on the content of the appraisal report to generate an appraisal report file that meets the output requirements;

[0228] Encapsulating the identification report file into a data packet, and sending the data packet to a target processing end through a communication module;

[0229] Receiving confirmation information from the target processing end, and storing the receiving confirmation information to complete the report sending process.

[0230] In one embodiment, the report generating and sending module 50 is specifically configured to:

[0231] Based on the seal usage policy of the target processing end, trigger the seal usage process of the appraisal report;

[0232] Performing stamp processing on the appraisal report and generating a report file with a stamp mark;

[0233] Packaging the report file with the seal mark, and sending the packaged report file to the report recipient;

[0234] Receive feedback information including report receiving status and / or seal confirmation information from the report receiver.

[0235] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external user terminal through a network connection. When the computer program is executed by the processor, it realizes the functions or steps of the service side of an intelligent voice comparison and identification method.

[0236] In one embodiment, a computer device is provided. The computer device may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps of a user side of an intelligent voice comparison and identification method

[0237] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the following steps are implemented:

[0238] Construct a speech comparison task including speech information, speech processing strategy and comparison strategy;

[0239] Processing the voice information based on the voice processing strategy to generate processed voice information;

[0240] Extracting speech feature data from the processed speech information;

[0241] Based on the comparison strategy of the voice comparison task, calling the voiceprint comparison model to perform comparison processing on the voice feature data to generate a comparison result;

[0242] An identification report is generated based on the comparison result, and the identification report is sent to a target processing end.

[0243] In one embodiment, a computer readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0244] Construct a speech comparison task including speech information, speech processing strategy and comparison strategy;

[0245] Processing the voice information based on the voice processing strategy to generate processed voice information;

[0246] Extracting speech feature data from the processed speech information;

[0247] Based on the comparison strategy of the voice comparison task, calling the voiceprint comparison model to perform comparison processing on the voice feature data to generate a comparison result;

[0248] An identification report is generated based on the comparison result, and the identification report is sent to a target processing end.

[0249] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can refer to the relevant descriptions on the server side and the user side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0250] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0251] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0252] It should be noted that if software tools or components other than those of the Company appear in the embodiments of the present application, they are only used for illustration and do not represent actual use. The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the above-mentioned embodiments, a person of ordinary skill in the art should understand that the technical solutions described in the above-mentioned embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents; and these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. An intelligent voice comparison and identification method, characterized in that: The following steps are involved: Construct a speech comparison task including speech information, speech processing strategy and comparison strategy; Processing the voice information based on the voice processing strategy to generate processed voice information; Extracting speech feature data from the processed speech information; Based on the comparison strategy of the voice comparison task, calling the voiceprint comparison model to perform comparison processing on the voice feature data to generate a comparison result; An identification report is generated based on the comparison result, and the identification report is sent to a target processing end.

2. The intelligent voice comparison and identification method according to claim 1, characterized in that: Construct a speech comparison task that includes speech information, speech processing strategy, and comparison strategy, including: Receive and analyze voice information to determine the source channel and format type of the voice information; Determining a voice processing strategy for processing the voice information based on a source channel and a format type of the voice information; Configure the model calling parameters and comparison thresholds in the comparison strategy; Based on the voice information, the voice processing strategy and the comparison strategy, a voice comparison task identifier is generated, and the voice comparison task identifier is associated with the voice information and stored to complete the generation of the voice comparison task.

3. The intelligent voice comparison and identification method according to claim 1, characterized in that: Processing the voice information based on the voice processing strategy to generate processed voice information includes: Based on the format conversion identifier in the voice processing strategy, the format type of the voice information is detected to determine whether it meets the format requirements; If the voice information does not meet the format requirement, performing a format conversion operation on the voice information based on the format conversion identifier in the voice processing strategy; Based on the denoising flag in the speech processing strategy, analyzing the background noise intensity in the speech information to determine whether denoising processing is required; If the voice information needs to be subjected to denoising processing, a denoising operation is performed on the voice information based on the denoising flag in the voice processing strategy to generate processed voice information.

4. The intelligent voice comparison and identification method according to claim 1, characterized in that: Extracting voice feature data from the processed voice information includes: Performing spectrum analysis on the processed voice information to extract frequency features of the voice information; Based on the frequency feature, determining the duration feature and energy feature of the voice information, and converting the duration feature and energy feature into a feature vector; A dimension compression process is performed on the feature vector, and speech feature data is generated based on the feature vector after the dimension compression process.

5. The intelligent voice comparison and identification method according to claim 1, characterized in that: Based on the comparison strategy of the voice comparison task, the voiceprint comparison model is called to perform comparison processing on the voice feature data to generate a comparison result, including: Based on the comparison strategy of the voice comparison task, determining a voiceprint comparison model for performing the comparison process; Calling the voiceprint comparison model, performing feature vector matching between the voice feature data and the voiceprint template in the database, and obtaining a feature vector matching result; Based on the feature vector matching result, the similarity score between the speech feature data and the voiceprint template is analyzed, and a comparison result is generated according to the similarity score.

6. The intelligent voice comparison and identification method according to claim 1, characterized in that: Generating an identification report based on the comparison result and sending the identification report to the target processing end includes: Extracting core information including similarity scores from the comparison results; Generate an identification report including a comparison description, a similarity analysis and / or a comparison conclusion based on the core information; Performing formatting processing on the content of the appraisal report to generate an appraisal report file that meets the output requirements; Encapsulating the identification report file into a data packet, and sending the data packet to a target processing end through a communication module; Receiving confirmation information from the target processing end, and storing the reception confirmation information to complete the report sending process.

7. The intelligent voice comparison and identification method as claimed in claim 1, characterized in that: After generating an identification report based on the comparison result and sending the identification report to the target processing end, the method further includes: Based on the seal usage policy of the target processing end, trigger the seal usage process of the appraisal report; Performing stamp processing on the appraisal report and generating a report file with a stamp mark; Packaging the report file with the seal mark, and sending the packaged report file to the report recipient; Receive feedback information including report receiving status and / or seal confirmation information from the report receiver.

8. An intelligent voice comparison and identification device, characterized in that: The intelligent voice comparison and identification device comprises: A task construction module, used to construct a speech comparison task including speech information, speech processing strategy and comparison strategy; A speech preprocessing module, used for processing the speech information based on the speech processing strategy to generate processed speech information; A feature extraction module, used to extract voice feature data from the processed voice information; A comparison processing module, used to call a voiceprint comparison model to perform comparison processing on the voice feature data based on the comparison strategy of the voice comparison task, and generate a comparison result; The report generation and sending module is used to generate an identification report based on the comparison result and send the identification report to the target processing end.

9. A computer device, characterized in that: The computer device includes a memory, a processor, and an intelligent voice comparison and identification program stored in the memory and executable on the processor. When the intelligent voice comparison and identification program is executed by the processor, the steps of the intelligent voice comparison and identification method as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: The storage medium stores an intelligent voice comparison and identification program, which, when executed by the processor, implements the steps of the intelligent voice comparison and identification method as described in any one of claims 1 to 7.