Speech analysis method and device, electronic equipment and computer readable storage medium

By preprocessing, vocalprint segmentation and identification classification of recorded data, and using preset prompt word extraction and execution strategies, the problem of inaccurate speech analysis in the existing technology is solved, and more efficient strategy push is achieved.

CN120544566APending Publication Date: 2025-08-26CHINA PING AN LIFE INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510728677.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

When existing speech analysis methods handle interaction records between the insurance industry and the smart medical industry, they cannot accurately analyze semantic information, which affects the accuracy of strategy push.

Method used

By obtaining recording data, preprocessing, voiceprint segmentation, conversion processing and identification classification are performed, and target execution strategies are used to extract preset prompt words to achieve semantic analysis.

Benefits of technology

Improve the accuracy of semantic analysis, thereby improving the accuracy of strategy push.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544566A_ABST
    Figure CN120544566A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice analysis and insurance services or smart medical treatment, and provides a voice analysis method and device, electronic equipment and a computer readable storage medium, and the method comprises the steps: obtaining recording data; preprocessing the recording data to obtain preprocessed recording information; performing voiceprint segmentation processing on the preprocessed recording information to obtain user voice information; performing conversion processing on the user voice information to obtain user voice text data; performing identification and classification processing on the user voice text data based on a preset prompt word to obtain user classification information; and extracting a target execution strategy from a preset execution strategy set based on the user classification information. According to the technical scheme, semantic analysis processing can be accurately carried out from the interaction records of the insurance industry or the medical industry, and then the accuracy of policy pushing of the insurance business or the medical industry can be well improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to, but are not limited to, the field of speech analysis, and in particular to a speech analysis method, apparatus, electronic device, and computer-readable storage medium. Background Art

[0002] With the continuous development of the social economy and the continuous improvement of people's living standards, the insurance and smart healthcare industries have also been continuously promoted and applied. In the insurance industry, the interaction records between customers and insurance customer service are important data resources. By analyzing these interaction records, insurance companies can better understand customer needs and thus optimize service strategies. In the smart healthcare industry, the interaction records between patients and intelligent question-and-answer robots are also important data resources. By analyzing these interaction records, medical staff can also better understand patients' needs. However, traditional speech analysis methods have limitations in processing large and complex interaction records. They cannot accurately analyze semantic information from interaction records, which affects the accuracy of policy push. Summary of the Invention

[0003] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.

[0004] In order to solve the problems mentioned in the above background technology, the embodiments of the present application provide a speech analysis method, device, electronic device and computer-readable storage medium, which can accurately perform semantic analysis and processing from interaction records, thereby greatly improving the accuracy of policy push.

[0005] In a first aspect, an embodiment of the present application provides a speech analysis method, comprising:

[0006] Get recording data;

[0007] Preprocessing the recorded data to obtain preprocessed recorded information;

[0008] Performing voiceprint segmentation processing on the pre-processed recording information to obtain user voice information;

[0009] Converting the user voice information to obtain user voice text data;

[0010] Recognizing and classifying the user's voice and text data based on preset prompt words to obtain user classification information;

[0011] A target execution policy is extracted from a preset execution policy set based on the user classification information.

[0012] In a second aspect, an embodiment of the present application further provides a speech analysis device, comprising:

[0013] An acquisition unit, used for acquiring recording data;

[0014] A preprocessing unit, configured to preprocess the recording data to obtain preprocessed recording information;

[0015] a segmentation unit, configured to perform voiceprint segmentation processing on the pre-processed recording information to obtain user voice information;

[0016] A conversion unit, configured to convert the user voice information to obtain user voice text data;

[0017] a recognition unit, configured to perform recognition and classification processing on the user voice and text data based on a preset prompt word to obtain user classification information;

[0018] The extraction unit is configured to extract a target execution strategy from a preset execution strategy set based on the user classification information.

[0019] In a third aspect, an embodiment of the present application further provides an electronic device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the speech analysis method as described in the first aspect above is implemented.

[0020] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are used to execute the speech analysis method described in the first aspect above.

[0021] According to the speech analysis method of the embodiment provided by the present application, there are at least the following beneficial effects: in the process of speech analysis, first, the recording data is obtained; then, the recording data is preprocessed to obtain preprocessed recording information; then, the preprocessed recording information is subjected to voiceprint segmentation processing to obtain user voice information; then, the user voice information is converted to obtain user voice text data; then, based on the preset prompt words, the user voice text data is identified and classified to obtain user classification information; finally, based on the user classification information, the target execution strategy is extracted from the preset execution strategy set. Through the above technical solution, the preprocessed recording information is subjected to voiceprint segmentation processing to obtain user voice information, then the user voice information is converted to obtain user voice text data, then, based on the preset prompt words, the user voice text data is identified and classified to obtain user classification information, and finally, based on the user classification information, the target execution strategy is extracted from the preset execution strategy set. This can accurately perform semantic analysis processing on the interaction record, improve the accuracy of semantic analysis, and thus greatly improve the accuracy of policy push. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The accompanying drawings are used to provide a further understanding of the technical solution of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the technical solution of the present application and do not constitute a limitation on the technical solution of the present application.

[0023] Figure 1 This is a flow chart of a speech analysis method provided by one embodiment of the present application;

[0024] Figure 2 yes Figure 1 A schematic flow chart of a specific implementation of step S200;

[0025] Figure 3 yes Figure 1 A schematic flow chart of a specific implementation of step S300;

[0026] Figure 4 yes Figure 1 A schematic flow chart of a specific implementation of step S400;

[0027] Figure 5 yes Figure 1 A schematic flow chart of a specific implementation of step S500;

[0028] Figure 6 yes Figure 1 A flowchart of a specific implementation of step S600;

[0029] Figure 7 It is executed Figure 1 A schematic flow chart of a specific implementation method after step S600;

[0030] Figure 8 is a schematic diagram of a speech analysis device provided by one embodiment of the present application;

[0031] Figure 9 This is a schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0032] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0033] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, used in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.

[0034] It should be noted that, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0035] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.

[0036] AI is a new technical discipline that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. Artificial intelligence is a branch of computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. Artificial intelligence can simulate the information processes of human consciousness and thinking. It also refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results.

[0037] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0038] Artificial intelligence, or AI, is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.

[0039] The servers involved in artificial intelligence technology can be independent servers or cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), as well as big data and artificial intelligence platforms.

[0040] The present application provides a speech analysis method, device, electronic device and computer-readable storage medium. In the process of speech analysis, first, recording data is obtained; then, the recording data is pre-processed to obtain pre-processed recording information; then, the pre-processed recording information is subjected to voiceprint segmentation processing to obtain user voice information; then, the user voice information is converted to obtain user voice text data; then, based on a preset prompt word, the user voice text data is identified and classified to obtain user classification information; finally, based on the user classification information, a target execution strategy is extracted from a preset execution strategy set. Through the above technical solution, the pre-processed recording information is subjected to voiceprint segmentation processing to obtain user voice information, then the user voice information is converted to obtain user voice text data, then, based on a preset prompt word, the user voice text data is identified and classified to obtain user classification information, and finally, based on the user classification information, a target execution strategy is extracted from a preset execution strategy set. This can accurately perform semantic analysis processing on the interaction record, improve the accuracy of semantic analysis, and thus greatly improve the accuracy of policy push.

[0041] The speech analysis method provided in the embodiment of the present application relates to the field of speech analysis technology. The speech analysis method provided in the embodiment of the present application can be applied in a terminal, can also be applied in a server side, and can also be software running in a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers, or as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0042] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.

[0043] It should be noted that in each specific embodiment of the present application, when it comes to the need to perform relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of such data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.

[0044] The embodiments of the present application are further described below with reference to the accompanying drawings.

[0045] like Figure 1 As shown, Figure 1 This is a flow chart of a speech analysis method provided by an embodiment of the present application, which includes the following steps:

[0046] Step S100: Acquire recording data.

[0047] The speech analysis method provided in the embodiment of the present application first obtains the recording data during the speech analysis process; then, preprocessing can be performed based on the recording data to obtain preprocessed recording information, in preparation for the subsequent speech analysis process.

[0048] For example, in the insurance business, the recorded data can be the content of a conversation between an insurance customer and an insurance salesperson, a telephone conversation between the insurance salesperson and the insurance customer, or a face-to-face conversation between the insurance salesperson and the insurance customer recorded using a recording device, where the recording device can be a mobile phone or a voice recorder; or it can be recorded during the communication between the insurance customer and the intelligent question-and-answer system. Alternatively, in the field of smart healthcare, the recorded data can be the content of a conversation between a patient and a medical staff member, a face-to-face conversation between the medical staff and the patient recorded using a recording device, or a conversation between the intelligent question-and-answer system and the patient.

[0049] It is worth noting that in each specific embodiment of the present application, when the recording data involves the need to perform relevant processing based on data related to the user's identity or characteristics such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of such data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.

[0050] Step S200: pre-processing the recording data to obtain pre-processed recording information.

[0051] The speech analysis method provided in the embodiment of the present application can pre-process the recording data after obtaining the recording data during the speech analysis process, thereby obtaining pre-processed recording information to prepare for subsequent speech analysis processing.

[0052] It is worth noting that during the preprocessing of the recorded data, the recorded data can be denoised, and then the denoised recorded information can be extracted based on the time threshold to obtain the preprocessed recorded information. Denoising the recorded data can effectively remove interference information from the recorded data, making subsequent speech analysis more accurate. Extracting the denoised recorded information based on the time threshold can obtain the preprocessed recorded information. Extracting the denoised recorded information based on the time threshold removes short recorded information, thereby significantly improving the accuracy of semantic analysis.

[0053] like Figure 2 As shown, preprocessing the recording data to obtain preprocessed recording information may include the following steps:

[0054] Step S210, performing denoising processing on the recorded data to obtain denoised recorded information;

[0055] Step S220: extracting and processing the denoised recording information according to a preset time threshold to obtain pre-processed recording information.

[0056] In the process of preprocessing the recording data to obtain preprocessed recording information in steps S210 to S220, the recording data is first denoised to obtain denoised recording information; then, according to a pre-set time threshold, the denoised recording information is extracted to obtain preprocessed recording information; based on the above technical solution, the accuracy of subsequent speech analysis can be greatly improved.

[0057] It is worth noting that denoising the recorded data to obtain denoised recorded information can effectively remove interference information in the recorded data, so that subsequent voice analysis can be more accurate. Then, according to a pre-set time threshold, the denoised recorded information is extracted and processed to eliminate shorter recorded information to further improve the accuracy of subsequent voice analysis. The time threshold can be set according to actual needs. For example, it can be set to 1 minute, and then recorded information with a duration of less than 1 minute can be eliminated to simplify the recorded information and make the recognition of the recorded information faster.

[0058] For example, in the insurance business, after obtaining audio recordings between customers and insurance salespeople, the acquired audio recordings can be denoised to obtain denoised audio recording information; then, audio recordings with a duration of less than 1 minute in the denoised audio recordings can be removed to obtain pre-processed audio recording information. The above technical solution can not only speed up the efficiency of subsequent voice analysis, but also improve the accuracy of subsequent voice analysis. Alternatively, in the smart healthcare field, after obtaining audio recordings between patients and intelligent question-and-answer robots, the acquired audio recordings can be denoised to obtain denoised audio recording information; then, audio recordings with a duration of less than 30 seconds in the denoised audio recordings can be removed to obtain pre-processed audio recording information. The above technical solution can also speed up the efficiency of subsequent voice analysis and improve the accuracy of subsequent voice analysis.

[0059] Step S300: Perform voiceprint segmentation processing on the pre-processed recording information to obtain user voice information.

[0060] The speech analysis method provided in the embodiment of the present application pre-processes the recording data to obtain pre-processed recording information; then performs voiceprint segmentation processing on the pre-processed recording information to obtain the corresponding user voice information; and extracts the user voice information for the user from the pre-processed recording information to prepare for the user's voice information analysis and processing.

[0061] It is worth noting that the pre-processed recording information is subjected to voiceprint segmentation processing, so that the voice information of different objects can be segmented from the pre-processed recording information. The embodiment of the present application is mainly for recommending the execution strategy for the user, so only the separated user voice information is analyzed and identified.

[0062] For example, in the insurance business, voiceprint segmentation can be performed on pre-processed audio recordings of conversations between users and insurance salespeople to obtain user voice information specific to the user, which can then be recognized. Alternatively, in the smart healthcare field, voiceprint segmentation can be performed on pre-processed audio recordings of conversations between patients and intelligent question-and-answer robots to obtain user voice information specific to the patient, which can then be recognized.

[0063] like Figure 3 As shown, performing voiceprint segmentation on the pre-processed recording information to obtain user voice information may include the following steps:

[0064] Step S310, extracting voiceprint features from the pre-processed recording information to obtain recording voiceprint feature information;

[0065] Step S320: extracting a feature vector from the recorded voiceprint feature information based on the pre-trained voiceprint recognition model to obtain a recorded voiceprint feature vector;

[0066] Step S330: segmenting the pre-processed recording information based on the pre-trained voice conversion detection model to obtain a plurality of pre-processed recording segments;

[0067] Step S340: clustering the plurality of pre-processed recording segments according to the recording voiceprint feature vector to obtain user voice information.

[0068] For steps S310 to S340, in the process of performing voiceprint segmentation processing on the pre-processed recording information to obtain user voice information, first, voiceprint feature extraction processing is performed on the pre-processed recording information to obtain recording voiceprint feature information; then, feature vector extraction processing is performed on the recording voiceprint feature information based on the pre-trained voiceprint recognition model to obtain recording voiceprint feature vector; then, segmentation processing is performed on the pre-processed recording information based on the pre-trained voice conversion detection model to obtain multiple pre-processed recording segments; finally, clustering processing is performed on the multiple pre-processed recording segments according to the recording voiceprint feature vector to obtain the corresponding user voice information, in preparation for subsequent user voice information analysis.

[0069] It is worth noting that by performing voiceprint feature extraction on the pre-processed recording information, the recording voiceprint feature information can be obtained, in preparation for the subsequent feature vector extraction; then, based on the pre-trained voiceprint recognition model, the recording voiceprint feature information can be subjected to feature vector extraction processing to obtain the corresponding recording voiceprint feature vector, and subsequently, clustering processing can be performed based on the recording voiceprint feature vector; the pre-processed recording information can also be segmented based on the pre-trained voice conversion detection model to obtain multiple pre-processed recording segments; subsequently, the multiple pre-processed recording segments can be clustered based on the recording voiceprint feature vector to obtain the corresponding user voice information.

[0070] It's worth noting that a pre-trained voiceprint recognition model is one that has already completed model training. This model can extract feature vectors from recorded voiceprint information to generate recording voiceprint feature vectors. A pre-trained speech transition detection model can identify transition pauses in recorded information to segment multiple pre-processed recording segments. The recording voiceprint feature vectors cluster multiple pre-processed recording segments to cluster and combine pre-processed recording segments belonging to the same subject.

[0071] For example, in the insurance business, voiceprint feature extraction can be performed on pre-processed audio recordings between customers and insurance salespeople to obtain audio recording voiceprint feature information; then, feature vector extraction can be performed on the audio recording voiceprint feature information based on a pre-trained voiceprint recognition model to obtain audio recording voiceprint feature vectors; then, the pre-processed audio recording information can be segmented based on a pre-trained voice conversion detection model to obtain multiple pre-processed audio segments; then, the multiple pre-processed audio segments can be clustered based on the audio recording voiceprint feature vectors to obtain the customer's user voice information. Alternatively, in a smart medical system, voiceprint feature extraction can be performed on pre-processed audio recordings between patients and intelligent question-and-answer robots to obtain audio recording voiceprint feature information; then, feature vector extraction can be performed on the audio recording voiceprint feature information based on a pre-trained voiceprint recognition model to obtain audio recording voiceprint feature vectors; then, the pre-processed audio recording information can be segmented based on a pre-trained voice conversion detection model to obtain multiple pre-processed audio segments; then, the multiple pre-processed audio segments can be clustered based on the audio recording voiceprint feature vectors to obtain the patient's user voice information.

[0072] Step S400: converting the user voice information into user voice text data.

[0073] The speech analysis method provided in the embodiment of the present application can convert the user voice information into user voice text data after performing voiceprint segmentation processing on the pre-processed recording information to prepare for subsequent speech analysis processing.

[0074] It is worth noting that in the process of converting the user voice information to obtain the user voice text data, the user voice information is first subjected to energy detection processing to obtain energy detection information; then, according to the pre-set energy threshold and energy detection information, the user voice information is subjected to silence elimination processing to obtain the adjusted voice information; then, the adjusted voice information is formatted and adjusted to obtain the standardized voice information; finally, the standardized voice information is subjected to text recognition processing based on the pre-trained speech recognition model to obtain the corresponding user voice text data.

[0075] like Figure 4 As shown, converting the user voice information to obtain user voice text data may include the following steps:

[0076] Step S410, performing energy detection processing on the user voice information to obtain energy detection information;

[0077] Step S420, performing silence removal processing on the user voice information according to the preset energy threshold and energy detection information to obtain adjusted voice information;

[0078] Step S430, formatting and adjusting the adjusted voice information to obtain standardized voice information;

[0079] Step S440: Perform text recognition processing on the standardized voice information based on the pre-trained voice recognition model to obtain user voice text data.

[0080] In steps S410 to S440, during the conversion of user voice information to obtain user voice text data, energy detection is first performed on the user voice information to obtain energy detection information; then, silence removal is performed on the user voice information based on a preset energy threshold and the energy detection information to obtain adjusted voice information; then, the adjusted voice information is formatted and adjusted to obtain standardized voice information; and finally, text recognition is performed on the standardized voice information based on a pre-trained voice recognition model to obtain the corresponding user voice text data. Through the above technical solution, user voice information can be converted into user voice text data, preparing for subsequent voice analysis.

[0081] It is worth noting that energy detection processing is performed on the user's voice information to obtain energy detection information. Then, silence removal processing is performed on the user's voice information based on a preset energy threshold and the energy detection information to obtain adjusted voice information. Through this technical means, the silence portion of the user's voice information can be removed, preparing for subsequent voice analysis processing. In addition, formatting and adjusting the adjusted voice information can obtain corresponding standardized voice information, preparing for subsequent text recognition processing.

[0082] For example, in the insurance business, after obtaining the customer's user voice information, the user voice information can be subjected to energy detection processing to obtain energy detection information; then, according to the preset energy threshold and energy detection information, the user voice information can be subjected to mute elimination processing to obtain adjusted voice information; then, the adjusted voice information can be subjected to formatting and adjustment processing to obtain standardized voice information; finally, based on the pre-trained speech recognition model, the standardized voice information can be subjected to text recognition processing to obtain user voice text data corresponding to the customer, and the customer's voice information can be subsequently analyzed and processed. Alternatively, in the smart medical field, after obtaining the patient's user voice information, the user voice information can be subjected to energy detection processing to obtain energy detection information; then, according to the preset energy threshold and energy detection information, the user voice information can be subjected to mute elimination processing to obtain adjusted voice information; then, the adjusted voice information can be subjected to formatting and adjustment processing to obtain standardized voice information; finally, based on the pre-trained speech recognition model, the standardized voice information can be subjected to text recognition processing to obtain user voice text data corresponding to the patient, and the patient's voice information can be subsequently analyzed and processed.

[0083] Step S500: Recognition and classification processing is performed on user voice and text data based on preset prompt words to obtain user classification information.

[0084] The speech analysis method provided in the embodiment of the present application can, after converting the user voice information to obtain the user voice text data, identify and classify the user voice text data based on pre-set prompt words to obtain user classification information; based on the user classification information, preparations can be made for subsequent execution strategy extraction.

[0085] It is worth noting that after obtaining the user voice and text data, the obtained user voice and text data can be matched and identified based on the pre-set prompt words to determine the user classification information corresponding to the user voice and text data. Subsequently, the corresponding target execution strategy can be determined from the execution strategy set based on the user classification information.

[0086] For example, in the insurance business field, when it is necessary to determine the reason why the user does not pay the insurance premium from the user's voice text data, the prompt words may include "financial difficulties and temporary inability to pay", "dissatisfied with the insurance company's service", "dissatisfied with the insurance product" and "other insurance companies are cheaper", etc., the user's voice text data can be matched with the above prompt words to determine the reason why the user does not pay the insurance premium, and then determine the corresponding user classification information.

[0087] like Figure 5 As shown, the user voice and text data is identified and classified based on the preset prompt words to obtain user classification information, which may include the following steps:

[0088] Step S510, segmenting the user's voice text data to obtain multiple voice text segments;

[0089] Step S520, matching multiple voice text segments based on preset prompt words to obtain target prompt words;

[0090] Step S530: Analyze and process the user attribute information corresponding to the target prompt word to obtain user classification information.

[0091] For steps S510 to S530, in the process of identifying and classifying the user voice text data based on the pre-set prompt words to obtain the user classification information, the user voice text data is first segmented to obtain multiple voice text segments, and then the multiple voice text segments are matched based on the pre-set prompt words to obtain the target prompt words; finally, the user attribute information corresponding to the target prompt words can be analyzed and processed to obtain the corresponding user classification information, in preparation for the subsequent execution strategy recommendation.

[0092] It is worth noting that the user's voice and text data is segmented to obtain multiple voice and text segments, which prepares for the subsequent matching process. Based on the pre-set prompt word, the multiple voice and text segments are matched to obtain the target prompt word. Subsequently, the user attribute information corresponding to the target prompt word can be analyzed and processed to obtain the corresponding user classification information.

[0093] For example, in the insurance business, the user voice text data corresponding to the customer and the insurance salesperson is first segmented and processed to obtain multiple voice text segments; then, based on pre-set prompt words, the multiple voice text segments are matched to obtain the target prompt words; finally, the user attribute information corresponding to the target prompt words is analyzed and processed to obtain the user classification information corresponding to the customer. Alternatively, in the smart medical field, the user voice text data corresponding to the patient and the intelligent dialogue robot is first segmented and processed to obtain multiple voice text segments; then, based on pre-set prompt words, the multiple voice text segments are matched to obtain the target prompt words; finally, the user attribute information corresponding to the target prompt words is analyzed and processed to obtain the user classification information corresponding to the patient.

[0094] Step S600: extracting a target execution policy from a preset execution policy set based on user classification information.

[0095] The speech analysis method provided in the embodiment of the present application can, after identifying and classifying the user's speech text data based on pre-set prompt words to obtain user classification information, extract the corresponding target execution strategy from the pre-set execution strategy set based on the user classification information, and then perform execution processing according to the target execution strategy.

[0096] It is worth noting that the execution policy set includes multiple execution policies. Based on user classification information, the corresponding target execution policy can be extracted from the multiple execution policies to prepare for subsequent policy execution.

[0097] like Figure 6As shown, the execution policy set includes multiple execution policies, each of which carries user category information. Extracting a target execution policy from the preset execution policy set based on the user category information may include the following steps:

[0098] Step S610, matching the user classification information with the user category information of the execution policy to obtain target user category information;

[0099] Step S620: Determine the execution policy corresponding to the target user category information as the target execution policy.

[0100] For steps S610 to S620, in the process of extracting the target execution policy from a pre-set execution policy set based on user classification information, the user classification information is first matched with the user category information of the execution policy to obtain the target user category information; then the execution policy corresponding to the target user category information can be determined as the target execution policy; through the above technical solution, the accuracy of execution strategy selection can be improved.

[0101] It is worth noting that in the process of determining the target execution strategy, the user classification information is first matched with the user category information of the execution strategy to obtain the target user category information; subsequently, the execution strategy corresponding to the target user category information can be determined as the target execution strategy, and then execution processing can be performed based on the target execution strategy.

[0102] like Figure 7 As shown, after extracting the target execution policy from the preset execution policy set based on the user classification information, the following steps may also be included:

[0103] Step S710, verifying the target execution policy to obtain a verification value;

[0104] Step S720: When the verification value is greater than the preset verification threshold, the target execution strategy remains unchanged.

[0105] In steps S710 to S720, after extracting the target execution policy from the preset execution policy set based on the user classification information, the target execution policy can be verified to obtain a verification value. If the verification value is greater than a preset verification threshold, the target execution policy can be maintained unchanged. If the verification value is not greater than the preset verification threshold, a new target execution policy needs to be selected. Through the above technical solution, the recommended execution of the policy can be made more reasonable and flexible.

[0106] It is worth noting that the target execution strategy is verified to obtain a verification value corresponding to the target execution strategy. For example, in the insurance business, the verification value can be measured based on the positive feedback of the target execution strategy. For example, in the process of collecting insurance premiums, if the target execution strategy can urge customers to pay, the verification value will be assigned a higher score; otherwise, the verification value will be assigned a lower score. The verification threshold can be set according to actual needs and is not limited here.

[0107] In addition, if Figure 8 As shown, an embodiment of the present application further provides a speech analysis device 10, comprising:

[0108] An acquisition unit 100 is used to acquire recording data;

[0109] The pre-processing unit 200 is used to pre-process the recording data to obtain pre-processed recording information;

[0110] The segmentation unit 300 is used to perform voiceprint segmentation processing on the pre-processed recording information to obtain user voice information;

[0111] The conversion unit 400 is used to convert the user voice information to obtain user voice text data;

[0112] The recognition unit 500 is used to recognize and classify the user's voice and text data based on the preset prompt words to obtain user classification information;

[0113] The extraction unit 600 is configured to extract a target execution policy from a preset execution policy set based on user classification information.

[0114] It should be noted that, in the process of voice analysis, the recording data is first obtained; then the recording data is pre-processed to obtain the pre-processed recording information; then the pre-processed recording information is subjected to voiceprint segmentation processing to obtain the user voice information; then the user voice information is converted to obtain the user voice text data; then the user voice text data is identified and classified based on the preset prompt words to obtain the user classification information; finally, the target execution strategy is extracted from the preset execution strategy set based on the user classification information. Through the above technical solution, the pre-processed recording information is subjected to voiceprint segmentation processing to obtain the user voice information, then the user voice information is converted to obtain the user voice text data, then the user voice text data is identified and classified based on the preset prompt words to obtain the user classification information, and finally, the target execution strategy is extracted from the preset execution strategy set based on the user classification information. This can accurately perform semantic analysis processing on the interaction records, improve the accuracy of semantic analysis, and thus greatly improve the accuracy of policy push.

[0115] The specific implementation of the speech analysis device 10 is substantially the same as the specific embodiment of the speech analysis method described above, and will not be described in detail herein.

[0116] In addition, if Figure 9 As shown, an embodiment of the present application further provides an electronic device 700 , which includes: a memory 720 , a processor 710 , and a computer program stored in the memory 720 and executable on the processor 710 .

[0117] The processor 710 and the memory 720 may be connected via a bus or other means.

[0118] The non-transient software programs and instructions required to implement the speech analysis methods of the above embodiments are stored in the memory 720 , and when executed by the processor 710 , the speech analysis methods of the above embodiments are performed.

[0119] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0120] In addition, an embodiment of the present application also provides a computer-readable storage medium, which stores computer-executable instructions, and the computer-executable instructions are executed by a processor 710 or a controller, for example, by a processor 710 in the above-mentioned device embodiment, so that the above-mentioned processor 710 can execute the speech analysis method in the above-mentioned embodiment.

[0121] The above embodiments may be used in combination, and modules with the same name in different embodiments may be the same or different.

[0122] The foregoing description describes specific embodiments of the present application, and other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0123] The various embodiments in this application are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from the other embodiments. In particular, the device, equipment, and computer-readable storage medium embodiments are generally similar to the method embodiments, so their descriptions are relatively simple. For relevant portions, refer to the descriptions of the method embodiments.

[0124] The apparatus, device, computer-readable storage medium and method provided in the embodiments of the present application correspond to each other. Therefore, the apparatus, device and non-volatile computer storage medium also have similar beneficial technical effects as the corresponding method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the corresponding apparatus, device and computer storage medium will not be repeated here.

[0125] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD through their own programming, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly performed using "logic compiler" software. This is similar to the software compilers used during program development. Before compilation, the original code must be written in a specific programming language, called a hardware description language (HDL). There are many HDLs, including ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also appreciate that simply programming a method flow in one of these hardware description languages ​​and programming it into an integrated circuit can easily create a hardware circuit that implements the logic method flow.

[0126] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC625D, Atmel AT91 SAM, Microchip PIC18F26K20, and Silicone Labs C8051 F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code format, the controller can be implemented in the form of logic gates, switches, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the means for implementing various functions included therein can also be considered as structures within the hardware component. Or even, the means for implementing various functions can be considered as both a software module implementing the method and a structure within the hardware component.

[0127] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0128] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing the embodiments of the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0129] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the embodiments of the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0130] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0131] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0132] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0133] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0134] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0135] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0136] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0137] In the embodiments of the present application, "at least one" refers to one or more, and "more" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent the existence of A alone, the existence of A and B at the same time, and the existence of B alone. Among them, A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b and c can represent: a, b, c, a and b, a and c, b and c or a and b and c, where a, b, c can be single or multiple.

[0138] Embodiments of the present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. Embodiments of the present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0139] The various embodiments in this application are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiment is generally similar to the method embodiment, so the description is relatively simple. For relevant parts, refer to the partial description of the method embodiment.

[0140] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included within the scope of the claims of the present application.

Claims

1. A speech analysis method, characterized in that: include: Get recording data; Preprocessing the recorded data to obtain preprocessed recorded information; Performing voiceprint segmentation processing on the pre-processed recording information to obtain user voice information; Converting the user voice information to obtain user voice text data; Recognizing and classifying the user's voice and text data based on preset prompt words to obtain user classification information; A target execution policy is extracted from a preset execution policy set based on the user classification information.

2. The speech analysis method according to claim 1, wherein: The preprocessing of the recording data to obtain preprocessed recording information includes: Performing denoising on the recorded data to obtain denoised recorded information; The denoised recording information is extracted and processed according to a preset time threshold to obtain the pre-processed recording information.

3. The speech analysis method according to claim 1, wherein: The performing voiceprint segmentation processing on the pre-processed recording information to obtain user voice information includes: Performing voiceprint feature extraction on the pre-processed recording information to obtain recording voiceprint feature information; Extracting feature vectors from the recorded voiceprint feature information based on a pre-trained voiceprint recognition model to obtain a recorded voiceprint feature vector; Segmenting the pre-processed recording information based on a pre-trained voice conversion detection model to obtain a plurality of pre-processed recording segments; Clustering is performed on the plurality of pre-processed recording segments according to the recording voiceprint feature vector to obtain the user voice information.

4. The speech analysis method according to claim 1, wherein: The converting the user voice information to obtain user voice text data includes: Performing energy detection processing on the user voice information to obtain energy detection information; performing silence removal processing on the user voice information according to a preset energy threshold and the energy detection information to obtain adjusted voice information; Performing formatting and adjustment processing on the adjusted voice information to obtain standardized voice information; The standardized voice information is subjected to text recognition processing based on a pre-trained voice recognition model to obtain the user voice text data.

5. The speech analysis method according to claim 1, wherein: The user voice and text data is identified and classified based on the preset prompt words to obtain user classification information, including: Segmenting the user's voice and text data to obtain multiple voice and text segments; Matching the plurality of voice text segments based on the preset prompt word to obtain a target prompt word; The user attribute information corresponding to the target prompt word is analyzed and processed to obtain the user classification information.

6. The speech analysis method according to claim 1, wherein: The execution policy set includes a plurality of execution policies, each of which carries user category information; and extracting a target execution policy from the preset execution policy set based on the user category information includes: Matching the user classification information with the user category information of the execution policy to obtain target user category information; The execution policy corresponding to the target user category information is determined as the target execution policy.

7. The speech analysis method according to claim 1, wherein: After extracting the target execution policy from the preset execution policy set based on the user classification information, the method further includes: Performing verification processing on the target execution strategy to obtain a verification value; When the verification value is greater than the preset verification threshold, the target execution strategy remains unchanged.

8. A speech analysis device, characterized in that: include: An acquisition unit, used for acquiring recording data; A preprocessing unit, configured to preprocess the recording data to obtain preprocessed recording information; a segmentation unit, configured to perform voiceprint segmentation processing on the pre-processed recording information to obtain user voice information; A conversion unit, configured to convert the user voice information to obtain user voice text data; a recognition unit, configured to perform recognition and classification processing on the user voice and text data based on a preset prompt word to obtain user classification information; The extraction unit is configured to extract a target execution strategy from a preset execution strategy set based on the user classification information.

9. An electronic device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the speech analysis method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing computer-executable instructions, characterized in that: The computer-executable instructions are used to execute the speech analysis method according to any one of claims 1 to 7.