Voice detection processing method and device

By combining the text and voice data of voice calls, abnormal detection and emotion recognition are performed, the user's emotions detection problems in online services are solved, and the accuracy and efficiency of resource recommendations are improved.

CN120126510APending Publication Date: 2025-06-10ANT GALAXY (CHONGQING) INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510279959.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

During the voice call of online services, the prior art is difficult to effectively detect user emotional changes and abnormal situations, resulting in the limitation of customer service's efficiency and accuracy of resource recommendations to users.

Method used

By obtaining the call text of the voice call and the storage address of the user's voice clip, combining the text detection strategy to perform abnormal detection, and inputting the user's voice clips into the emotion recognition model for emotional recognition. If the detection results and emotion recognition results meet the resource recommendation conditions, the user's mark will be cancelled.

Benefits of technology

It realizes the detection and analysis of user emotions from the two dimensions of text and voice, improves the accuracy and efficiency of resource recommendations, and reduces the rate of customer service error tags.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126510A_ABST
    Figure CN120126510A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a voice detection processing method and device.The voice detection processing method comprises the steps that in the call voice processing process, on one hand, a call text obtained by a voice call of resource recommendation conducted by a customer service on a marked user is obtained, anomaly detection is conducted on the call text, and a detection result is obtained; on the other hand, a storage address of a user voice segment in call voice corresponding to the call text is obtained, the storage address and a user text segment corresponding to the user voice segment are input into an emotion recognition model, and the emotion recognition model is used for obtaining the user voice segment and performing emotion recognition to obtain probability distribution of each emotion category; under the condition that the detection result and the probability distribution meet the resource recommendation condition, the mark of the marked user is canceled, so that the marked user is detected from the text dimension and the voice dimension.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to the field of data processing technologies, and particularly relates to a voice detection and processing method and apparatus. Background Art

[0002] With the continuous development and popularization of Internet technologies, various online services provided based on Internet technologies have emerged as the times require. The application scope of online services is also becoming increasingly extensive, covering many fields. In order to promote to users, the service providers of these online services can recommend services to users through customer service. As the needs of users for service recommendations become more and more diversified, the recommendation methods for customer service to recommend services to users have also changed accordingly. For example, services are recommended to users through voice calls or through conversations. In this process, higher requirements are also put forward for customer service to recommend services, and certain challenges are also brought to the service providers of online services. Summary of the Invention

[0003] One or more embodiments of this specification provide a voice detection and processing method, including: obtaining a call text obtained from a voice call in which a customer service recommends resources to a marked user, and obtaining a storage address of a user voice segment in the call voice corresponding to the call text. Performing anomaly detection on the call text according to a text detection strategy matched with the call text to obtain a detection result. Inputting the storage address and the user text segment corresponding to the user voice segment into an emotion recognition model to obtain the user voice segment and perform emotion recognition, and obtaining a probability distribution of each emotion category. If the detection result and the probability distribution meet the resource recommendation condition, cancel the marking of the marked user.

[0004] One or more embodiments of this specification provide a voice detection and processing apparatus, including: an obtaining module configured to obtain a call text obtained from a voice call in which a customer service recommends resources to a marked user, and obtain a storage address of a user voice segment in the call voice corresponding to the call text. An anomaly detection module configured to perform anomaly detection on the call text according to a text detection strategy matched with the call text to obtain a detection result. An emotion recognition module configured to input the storage address and the user text segment corresponding to the user voice segment into an emotion recognition model to obtain the user voice segment and perform emotion recognition, and obtain a probability distribution of each emotion category. A marking cancellation module configured to cancel the marking of the marked user if the detection result and the probability distribution meet the resource recommendation condition.

[0005] One or more embodiments of this specification provide a voice detection and processing device, including: a processor; and a memory configured to store computer-executable instructions, which when executed cause the processor to: obtain the call text obtained from the voice call in which the customer service makes a resource recommendation to the marked user, and obtain the storage address of the user voice segment in the call voice corresponding to the call text. Perform anomaly detection on the call text according to the text detection strategy matched by the call text to obtain a detection result. Input the storage address and the user text segment corresponding to the user voice segment into an emotion recognition model to obtain the user voice segment and perform emotion recognition, and obtain the probability distribution of each emotion category. If the detection result and the probability distribution meet the resource recommendation conditions, cancel the marking of the marked user.

[0006] One or more embodiments of this specification provide a computer-readable storage medium for storing computer-executable instructions, which when executed implement the following steps: obtain the call text obtained from the voice call in which the customer service makes a resource recommendation to the marked user, and obtain the storage address of the user voice segment in the call voice corresponding to the call text. Perform anomaly detection on the call text according to the text detection strategy matched by the call text to obtain a detection result. Input the storage address and the user text segment corresponding to the user voice segment into an emotion recognition model to obtain the user voice segment and perform emotion recognition, and obtain the probability distribution of each emotion category. If the detection result and the probability distribution meet the resource recommendation conditions, cancel the marking of the marked user. Description of the Drawings

[0007] In order to more clearly illustrate the technical solutions in one or more embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings; Figure 1 It is a schematic diagram of the implementation environment of a voice detection and processing method provided by one or more embodiments of this specification; Figure 2 It is a processing flow chart of a voice detection and processing method provided by one or more embodiments of this specification; Figure 3 It is a schematic diagram of a strategy combination interface provided by one or more embodiments of this specification; Figure 4 It is a schematic diagram of the configuration page of a text detection strategy provided by one or more embodiments of this specification; Figure 5A processing flow chart of a voice detection processing method applied to a resource application recommendation scenario provided for one or more embodiments of this specification; Figure 6 A schematic diagram of an embodiment of a voice detection processing device provided for one or more embodiments of this specification; Figure 7 A schematic structural diagram of a voice detection processing device provided for one or more embodiments of this specification. Detailed implementation manners

[0008] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the following will clearly and completely describe the technical solutions in one or more embodiments of this specification with reference to the accompanying drawings in one or more embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this document.

[0009] The voice detection processing method provided by one or more embodiments of this specification is applicable to the implementation environment of voice detection. Referring to Figure 1 , this implementation environment at least includes: A server 101 for performing voice detection processing; Among them, the server 101 can be one or more servers, a server cluster composed of several servers, or a cloud server of a cloud computing platform; the server 101 can be deployed with a text detection module 102, an emotion recognition model 103, and a marking cancellation module 104; In this implementation environment, the server 101 obtains the call text obtained from the voice call in which the customer service recommends resources to the marked user, and obtains the storage address of the user voice segment in the call voice corresponding to the call text. The text detection module 102 deployed on the server 101 performs anomaly detection on the call text according to the text detection strategy matched by the call text to obtain a detection result. The server 101 inputs the storage address and the user text segment corresponding to the user voice segment into the emotion recognition model 103, and obtains the probability distribution of each emotion category through the emotion recognition model 103 for obtaining and emotion recognition of the user voice segment. When the detection result and the probability distribution meet the resource recommendation conditions, the marking cancellation module 104 deployed on the server 101 cancels the marking of the marked user, so as to realize the detection of the marked user from the text dimension and the voice dimension.

[0010] One or more embodiments of a voice detection processing method provided by this specification are as follows: Referring to Figure 2, the voice detection and processing method provided in this embodiment specifically includes steps S202 to S208.

[0011] Step S202: Obtain the call text obtained from the voice call in which the customer service recommends resources to the marked user, and obtain the storage address of the user voice segment in the call voice corresponding to the call text.

[0012] The customer service in this embodiment may include human customer service and / or intelligent customer service. A marked user refers to a user with a mark, and the mark may include an abnormal mark, that is, the marked user may include an abnormally marked user; optionally, the mark of the marked user is generated after the voice call, that is, the mark of the marked user can be created by the customer service after the voice call. Specifically, the customer service can initiate a voice call for resource recommendation to the user. After the customer service and the user have a voice call, the customer service can mark the user. Specifically, the customer service can mark the user as abnormal. After the marking is completed, the user becomes a marked user. The customer service's marking of the user can also include writing the user into the abnormal mark table; the mark of the marked user can be used to characterize that the marked user may have negative emotions towards the voice call for resource recommendation.

[0013] Resource recommendation refers to recommending resources to the user. Specifically, it can be a recommendation of a resource item to the user, such as a recommendation of a resource application item to the user or a recommendation of a resource purchase item to the user. That is, resource recommendation can include resource application recommendation or resource purchase recommendation.

[0014] The call text refers to the call text of the voice call between the customer service and the marked user. The call text may include the customer service call text and / or the user call text. The customer service call text refers to the call text of the customer service, and the user call text refers to the call text of the marked user; the call text can specifically be the call text obtained by performing voice recognition on the voice call for resource recommendation.

[0015] The user voice segment in the call voice refers to the voice segment of the marked user; the storage address of the user voice segment may include the storage link of the user voice segment in the data storage module. For example, the storage address of the user voice segment is the URL (Uniform Resource Locator) link of the user voice segment in the data storage module.

[0016] In specific implementation, on the one hand, obtain the call text obtained from the voice call in which the customer service recommends resources to the marked user, and perform subsequent text anomaly detection through the call text. On the other hand, obtain the storage address of the user voice segment in the call voice corresponding to the call text, and perform subsequent voice detection from the dimension of the call voice.

[0017] In practical applications, the capacity of the user voice segment in the call voice may be large. If the user voice segment is directly stored for subsequent emotion recognition, it may lead to a large storage volume of the user voice segment. In view of this, to reduce the storage volume of the user voice segment, the user voice segment can be stored in the data storage module, and subsequent emotion recognition can be performed through the storage address of the user voice segment. In an alternative implementation provided in this embodiment, the storage address of the user voice segment is generated by the data storage module and obtained by reading from the data storage module. The storage address of the user voice segment can be obtained after replacing the domain name of the initial storage address created based on the storage identifier of the user voice segment in the data storage module. The specific storage address can be generated in the following manner: Create an initial storage address based on the storage identifier of the user voice segment in the data storage module; Replace the domain name in the initial storage address with the target domain name to obtain the storage address of the user voice segment.

[0018] Among them, the data storage module refers to the data storage module that stores the user voice segment. The storage identifier may include the unique identifier of the user voice segment in the data storage module. Specifically, the storage identifier may be the storage path and / or name of the user voice segment in the storage bucket of the data storage module. The initial storage address refers to the initial storage address, which can be used for external access. For example, the initial storage address is a URL link based on HTTP (Hyper Text Transfer Protocol). The target domain name may include a custom domain name, such as xxx.xxx.com.

[0019] Specifically, an initial storage address can be created based on the storage path and / or storage name of the user voice segment in the storage bucket of the data storage module, and the domain name in the initial storage address is replaced with a custom domain name to obtain the storage address of the user voice segment. In this way, while ensuring the secure transmission of data through the storage address, it is ensured that external users can smoothly access the required user voice segment. Through domain name replacement, the storage address points to a specific service address, achieving public network accessibility.

[0020] In practical application scenarios, the call voice may include the customer service voice and the user voice, and the call voice may be long. In view of this, to perform subsequent emotion recognition more conveniently, the user voice segment can be extracted from the call voice, and subsequent emotion recognition can be performed through the user voice segment to improve the efficiency and convenience of emotion recognition. In an alternative implementation provided in this embodiment, the user voice segment is obtained in the following manner: Extract user text segments from the call text, determine the time offset between the user text segments and the call text, and perform channel separation on the call voice to obtain the user voice marked with the user; Perform voice segmentation on the user voice according to the time offset, and upload the user voice segments obtained by the voice segmentation to the data storage module.

[0021] Among them, the time offset can be the time offset of the user text segment relative to the start time of the call text. Optionally, the time offset includes the start time offset and / or the end time offset. The start time offset can include the time offset of the start time of the user text segment relative to the start time of the call text, and the end time offset can include the time offset of the end time of the user text segment relative to the start time of the call text.

[0022] Specifically, user text segments can be extracted from the call text, the start time offset and / or the end time offset of the user text segment can be determined based on the start time of the call text, the call voice can be subjected to channel separation to obtain the user voice marked with the user, and the user voice can be subjected to voice segmentation according to the start time offset and / or the end time offset. The user voice segments obtained by the voice segmentation are uploaded to the data storage module; after the above voice segmentation of the user voice according to the time offset is executed, the candidate user voice segments with a voice duration greater than the duration threshold in the candidate user voice segments obtained by the voice segmentation can be used as user voice segments to be uploaded to the data storage module; the voice duration here refers to the voice duration of the candidate user voice segment; Before the above channel separation of the call voice, the file format of the call voice can also be obtained based on the voice storage address of the call voice. If the file format is the preset file format, the call voice can be subjected to channel separation to obtain the user voice marked with the user. If the file format is not the preset file format, the file format of the call voice can be converted to the preset file format, and then the converted call voice can be subjected to channel separation to obtain the user voice marked with the user; the preset file format here can be any file format. For example, the preset file format is the wav format. In the case where the file format is the MP3 format, the file format of the call voice is converted from the MP3 format to the wav format; thus, user voice segments are obtained through methods such as channel separation and voice segmentation, the effectiveness of the user voice segments is improved, and further the convenience and efficiency of subsequent emotion recognition are improved.

[0023] It should be noted that the above implementation manner of how to obtain the user voice segments can be executed on the basis of the implementation manner of how to generate the storage address of the user voice segments.

[0024] In specific implementation, after the voice call for resource recommendation to the user by the customer service is completed, the service provider for resource recommendation may send a call completion message; optionally, the call text is obtained by querying the text based on the call parameters carried in the call completion message sent by the service provider for resource recommendation; among them, the call parameters may be parameters characterizing the uniqueness of the call, for example, the call parameter is the call number; specifically, the call text may take the call parameters carried in the call completion message sent by the service provider for resource recommendation as the interface call input, and obtain it after calling the text query interface for text query. Here, the text query interface may be the text query interface of the call system.

[0025] In the specific execution process, before extracting the user voice segment, it is necessary to obtain the call voice as the basis for extracting the user voice segment; optionally, the call voice is obtained by downloading the voice based on the voice storage address. Specifically, the call voice can be downloaded from the call system based on the voice storage address; optionally, the voice storage address is obtained by reading from the call completion message sent by the service provider for resource recommendation.

[0026] In addition, in actual applications, the call completion message may not carry the voice storage address. In view of this, in order to successfully extract the user voice segment from the call voice for subsequent emotion recognition; in an optional implementation manner provided in this embodiment, the voice storage address can be obtained by querying the address based on the call number information in the call details of the voice call. Specifically, the voice storage address can be obtained through the following method: Call the call query interface to query the call details of the voice call; Extract the call number information from the call details, and use the call number information and the call identifier as the interface call input to call the address query interface to query the voice storage address.

[0027] Among them, the call query interface and the address query interface can be provided by the call system; the call details refer to the detailed information of the voice call. The call number information may include the number pool information. For example, the call number information is the information in the numberPool (number pool) field in the call details; the call identifier may include the call identifier. For example, the call identifier is callId (Call Identifier, call identifier).

[0028] Specifically, the call query interface can be called to query the call details of the call voice, extract the number pool information from the call details, and use the number pool information and the call identifier as the interface call input to call the address query interface to query the voice storage address of the call voice.

[0029] It should be added that the above implementation of how to obtain the voice storage address can be executed based on the above implementation of how to obtain the user voice segment; the above implementation of how to obtain the voice storage address can also be executed on the basis that the address query interface of the call system fails to query the voice storage address of the call voice. In the case where the address query interface of the call system queries the voice storage address of the call voice, the call voice can be obtained by voice download based on the voice storage address.

[0030] During the specific execution process, after the customer service makes a voice call for resource recommendation to the user, the user can be written into the exception mark table, such as being written into the exception mark list. The users recorded in the exception mark table can be marked, indicating that the user may have negative emotions towards the voice call for resource recommendation. In this case, the call records of the marked users recorded in the exception mark table can be obtained, and subsequent text detection and emotion recognition can be performed; in an optional implementation provided by this embodiment, during the process of obtaining the call text of the voice call for resource recommendation made by the customer service to the marked user, the call records can be read from the data processing module based on the number of call records stored in the data processing module and the read limit number, and the call text can be queried based on the call parameters included in the target call record in the call records. The following specific operations can be performed: Determine the number of call records stored in the data processing module, and calculate the data read volume based on the read limit number and the number of the data processing module; Read the call records from the data processing module according to the data read volume, and query the call text based on the call parameters included in the target call record in the call records.

[0031] Among them, the call records stored in the data processing module can be historical call records. The call records stored in the data processing module can be the call records of all users, or the call records of the marked users in the exception mark table; the read limit number of the data processing module can include the number limited for each time of reading data from the data processing module; the data read volume can include the actual data volume read from the data processing module each time; the target call record can include the target call record of the marked user in the call records. The call parameter can be a parameter representing the uniqueness of the call, such as the call number.

[0032] Specifically, the number of call records of a user marked in the exception flag table stored in the data processing module can be determined. If the number does not exceed the reading limit number of the data processing module, the call records can be read from the data processing module, and the call text can be queried based on the call parameters included in the target call record in the call records; if the number exceeds the reading limit number of the data processing module, the data reading amount can be calculated based on the reading limit number of the data processing module and the number, and the call records can be read from the data processing module according to the data reading amount, and the call text can be queried based on the call parameters included in the target call record in the call records; in the process of reading the call records from the data processing module according to the data reading amount, a query statement can be generated based on the data reading amount, and the call records can be read from the data processing module by executing the query statement; among them, the query statement can include a statement for call records, such as the query statement is a limit statement.

[0033] In the above process of generating a query statement based on the data reading amount and reading the call records from the data processing module by executing the query statement, in order to improve the efficiency of data reading, the data reading can be performed in the way of an asynchronous task. Specifically, asynchronous tasks can be generated according to the concurrency number of the data processing module, and the call records can be queried through the query statements of the asynchronous tasks; the concurrency number here can be the number of concurrencies that the data processing module can bear.

[0034] In addition, call records may continue to be generated, so the newly added call records can be obtained regularly; in an optional implementation manner provided in this embodiment, in the process of obtaining the call text obtained from the voice call in which the customer service recommends resources to the marked user, the time stamp of the newly added data partition can be detected according to the detection period, and when the time stamp meets the reading condition, the call text can be queried according to the target call record in the call records read from the newly added data partition; specifically, the following operations can be performed: Detect the time stamp of the newly added data partition according to the detection period, and determine whether the time stamp meets the reading condition; If it meets the condition, determine the number of call records stored in the newly added data partition. If the number does not exceed the preset number threshold, query the call text according to the target call record in the call records read from the newly added data partition.

[0035] Among them, the detection period can include the time period for detecting the time stamp of the newly added data partition, such as days, 1 hour, 2 hours; the time stamp of the newly added data partition can include the partition time stamp of the newly added data partition, such as the partition time stamp of the newly added dt partition; the reading condition can include that the time stamp is the time stamp of the newly added data partition, that is, the time stamp of the newly generated data partition; the call records stored in the newly added data partition can be newly added call records, which can be the newly added call records of all users, or the newly added call records of the users marked in the exception flag table.

[0036] Specifically, the timestamp of the newly added data partition can be detected according to the detection period. If the timestamp does not meet the reading condition, no processing may be performed or the timestamp of the newly added data partition may continue to be detected according to the detection period. If the timestamp meets the reading condition, the number of call records stored in the newly added data partition can be determined. When the number does not exceed the preset quantity threshold, the call text is queried according to the call parameters included in the target call record in the call records read from the newly added data partition. When the number exceeds the preset quantity threshold, the data reading quantity is calculated based on the read limit quantity and the number, and call records are read from the newly added data partition according to the data reading quantity, and the call text is queried based on the call parameters included in the target call record in the call records. Specifically, a query statement can be generated based on the data reading quantity, and call records are read from the newly added data partition by executing the query statement, and the call text is queried based on the call parameters included in the target call record in the call records.

[0037] It should be added that the operations of obtaining the call text obtained from the voice call for resource recommendation to the marked user by the customer service and obtaining the storage address of the user voice segment in the call voice corresponding to the call text can be replaced by obtaining the call text obtained from the voice call between the customer service and the marked user and obtaining the storage address of the user voice segment in the call voice corresponding to the call text; or can be replaced by obtaining the call voice and / or call text obtained from the voice call for resource recommendation to the marked user by the customer service; or can be replaced by obtaining the call text obtained from the voice call for resource recommendation to the marked user by the customer service, and / or, obtaining the storage address of the user voice segment in the call voice corresponding to the call text; or can be replaced by obtaining the call text obtained from the voice call for service recommendation to the marked user by the customer service, and / or, obtaining the user voice segment in the call voice corresponding to the call text; and form a new implementation manner with other processing steps provided in this embodiment.

[0038] Step S204, perform anomaly detection on the call text according to the text detection strategy matched by the call text to obtain a detection result.

[0039] On the one hand, the call text obtained from the voice call for resource recommendation to the marked user by the customer service is obtained above, and on the other hand, the storage address of the user voice segment in the call voice corresponding to the call text is obtained. In this step, starting from the call text dimension, anomaly detection is performed on the call text according to the text detection strategy matched by the call text to obtain a detection result, so as to achieve anomaly detection from the text dimension.

[0040] The text detection strategy in this embodiment refers to the detection strategy for detecting abnormalities in call texts. Specifically, the configuration and addition of the text detection strategy can be performed through the voice detection system, and the deletion or offline of the text detection strategy can also be carried out through the voice detection system. The text detection strategy can be a single text detection strategy or a strategy combination composed of multiple text detection strategies. For example Figure 3 As shown in the strategy combination interface, multiple text detection strategies for combination can be selected in the configuration area of the strategy combination, and the name of the strategy combination can be entered in the configuration area of the strategy combination name. After triggering the "OK" control, multiple text detection strategies can be combined into a strategy combination. Each text detection strategy can be a keyword detection strategy, a regular detection strategy, or a model detection strategy. Optionally, the text detection strategy is obtained by querying based on the strategy identifier carried in the call completion message sent by the service provider for resource recommendation. The detection result can include passing the detection and failing the detection, and the detection result can be the detection result of marking the user.

[0041] In practical applications, different resource types for resource recommendation and different abnormal levels of marked users may lead to different emotional categories of marked users during the voice call with the customer service. In view of this, in order to improve the effectiveness and comprehensiveness of the text detection strategy for call text matching; in an optional implementation manner provided in this embodiment, the text detection strategy for call text matching is obtained through the following method: Determine candidate text detection strategies in the text detection strategy set according to the resource type for resource recommendation; Based on the text length of the call text and / or the abnormal level of the marked user, screen out text detection strategies from the candidate text detection strategies.

[0042] Among them, the resource type for resource recommendation refers to the resource type for recommending resources to the marked user. For example, the resource type includes a credit resource type, a precious metal resource type, and / or an asset investment resource type. The text length of the call text can be determined based on the number of characters included in the call text. The abnormal level of the marked user can include the abnormal level marked by the marked user. For example, the marks of the marked user include three abnormal levels: a1, a2, and a3, and the relationship of the abnormal degrees of these three abnormal levels is a1 > a2 > a3. The candidate text detection strategies can include matching text detection strategies that match the resource type, that is, text detection strategies adapted to the resource type for resource recommendation. The text detection strategy set can be a set composed of one or more text detection strategies, and the one or more text detection strategies can be the text detection strategies of the user, that is, the text detection strategy set can be the text detection strategy set of the user.

[0043] Specifically, in the process of screening out a text detection strategy from candidate text detection strategies based on the text length of the call text and / or the abnormal level of the marked user, a text detection strategy can be screened out from candidate text detection strategies and the default text detection strategy based on the text length of the call text and / or the abnormal level of the marked user; the default text detection strategy here can be a text detection strategy that does not distinguish the resource types of resource recommendations and is adapted to all resource types, that is, under any resource type of resource recommendation, the default text detection strategy can be used to perform abnormal detection on the call text. In the process of screening out a text detection strategy from candidate text detection strategies and the default text detection strategy based on the text length of the call text and / or the abnormal level of the marked user, if the text length of the call text is greater than the length threshold, the candidate text detection strategy and the default text detection strategy can be used as the text detection strategy. If the text length of the call text is less than or equal to the length threshold, the default text detection strategy can be used as the text detection strategy; it is also possible that if the abnormal level of the marked user is lower than the preset abnormal level, the default text detection strategy can be used as the text detection strategy. If the abnormal level is not lower than the preset abnormal level, the candidate text detection strategy and the default text detection strategy can be used as the text detection strategy; it is also possible that if the text length of the call text is less than or equal to the length threshold and the abnormal level of the marked user is lower than the preset abnormal level, the default text detection strategy can be used as the text detection strategy. If the text length of the call text is greater than the length threshold or the abnormal level of the marked user is not lower than the preset abnormal level, the candidate text detection strategy and the default text detection strategy can be used as the text detection strategy.

[0044] The text detection strategy can be obtained through configuration by the voice detection system; for example, the default text detection strategy or the matching text detection strategy can be selected and configured by triggering the strategy configuration controls included in the strategy configuration page of the voice detection system; the matching text detection strategy refers to the text detection strategy that matches the resource type of the resource recommendation, such as Figure 4 As shown in the configuration page of the text detection strategy, the configuration page includes configuration controls for tenants, service lines. By triggering these configuration controls, tenants and service lines can be configured, the major and minor categories of the configuration policy type can be configured, and the policy name, policy object can be configured. The policy object can select the customer service or the user, select whether to hit all policies, and the keyword detection strategy, regular detection strategy, and / or model detection strategy can be configured. If the keyword detection strategy is selected, keywords can be set. If the regular detection strategy is selected, the hit regular expression can be set, or the non-hit regular expression can also be set. If the model detection strategy is selected, the algorithm model and the model score can be set.

[0045] In actual application scenarios, the call text may be relatively long. If the text detection strategy is directly used to perform anomaly detection on the call text itself, it may lead to low detection efficiency and low detection accuracy. In view of this, in order to improve the detection efficiency and detection accuracy of anomaly detection; in an alternative implementation provided in this embodiment, during the process of performing anomaly detection on the call text according to the text detection strategy matched by the call text and obtaining the detection result, the text segments included in the call text can be screened based on the regular detection strategy, and the screened abnormal text segments can be input into the text detection model for text anomaly detection to obtain the detection result. The specific operations can be as follows: Perform segmentation processing on the call text to obtain text segments, and screen the text segments based on the regular detection strategy to obtain abnormal text segments; Input the abnormal text segments into the text detection model for text anomaly detection to obtain the detection result.

[0046] Among them, the regular detection strategy can be a regular detection strategy constructed based on one or more regular expressions, that is, the regular detection strategy can include one or more regular expressions. For example, the regular expression is to screen phone numbers in a specific format, and screen the text segments that meet the regular expression from the call text; the regular expression can be a regular expression that matches text segments or a regular expression that does not match text segments. The regular expression that matches text segments refers to screening out the abnormal text segments that match the regular expression from the text segments, and the regular expression that does not match text segments refers to ignoring the text segments that match the regular expression in the text segments, that is, screening may not be performed.

[0047] Specifically, the call text can be segmented according to punctuation marks to obtain text segments, and the abnormal text segments that meet one or more regular expressions included in the regular detection strategy are screened out from the text segments. The abnormal text segments are input into the text detection model for text anomaly detection to obtain the anomaly detection score. If the anomaly detection score exceeds the score threshold, it is determined that the detection result is a failed detection. If the anomaly detection score does not exceed the score threshold, it is determined that the detection result is a passed detection.

[0048] In addition, during the process of performing anomaly detection on the call text according to the text detection strategy matched by the call text and obtaining the detection result, anomaly detection can be performed on the call text according to the keyword detection strategy matched by the call text to obtain the detection result, or anomaly detection can be performed on the call text according to the regular detection strategy matched by the call text to obtain the detection result, or anomaly detection can be performed on the call text according to the model detection strategy matched by the call text to obtain the detection result; Among them, the keyword detection strategy can be a keyword detection strategy constructed based on one or more keywords, that is, the keyword detection strategy can include one or more keywords. For example, the keywords included in the keyword detection strategy include "complaint", "dissatisfaction", and "problem".

[0049] Specifically, in the process of performing anomaly detection on the call text according to the keyword detection strategy matched with the call text to obtain the detection result, the call text can be matched with the keywords included in the keyword detection strategy. If the call text matches all the keywords included in the keyword detection strategy, it can be determined that the detection result is a failed detection. If the call text does not match all the keywords included in the keyword detection strategy, it can be determined that the detection result is a passed detection; or, the call text can be matched with the keywords included in the keyword detection strategy. If the call text matches any one of the keywords included in the keyword detection strategy, it can be determined that the detection result is a failed detection. If the call text does not match any of the keywords included in the keyword detection strategy, it can be determined that the detection result is a passed detection. In the process of performing anomaly detection on the call text according to the regular expression detection strategy matched with the call text to obtain the detection result, the call text can be matched with the regular expressions included in the regular expression detection strategy. If the call text matches all the regular expressions included in the regular expression detection strategy, it can be determined that the detection result is a failed detection. If the call text does not match all the regular expressions included in the regular expression detection strategy, it can be determined that the detection result is a passed detection; or, the call text can be matched with the regular expressions included in the regular expression detection strategy. If the call text matches any one of the regular expressions included in the regular expression detection strategy, it can be determined that the detection result is a failed detection. If the call text does not match any of the regular expressions included in the regular expression detection strategy, it can be determined that the detection result is a passed detection. In the process of performing anomaly detection on the call text according to the model detection strategy matched with the call text to obtain the detection result, the call text can be input into the text detection model for text anomaly detection to obtain the anomaly detection score. If the anomaly detection score exceeds the score threshold, it is determined that the detection result is a failed detection. If the anomaly detection score does not exceed the score threshold, it is determined that the detection result is a passed detection.

[0050] It should be noted that in the process of performing anomaly detection on the call text according to the text detection strategy matched with the call text and obtaining the detection result, any one or more of the above keyword detection strategy, regular detection strategy, and model detection strategy can also be used for anomaly detection. When any one or more of the anomaly detection results in the anomaly detection result pass the detection, it is determined that the detection result passes the detection. For example, when using the keyword detection strategy, regular detection strategy, and model detection strategy for anomaly detection, when any one or more of the three anomaly detection results pass the detection, it is determined that the detection result passes the detection.

[0051] In addition, in the process of performing anomaly detection on the call text according to the text detection strategy matched with the call text and obtaining the detection result, the call text can be subjected to anomaly detection according to the combination of text detection strategies matched with the call text. If the anomaly detection results of all the text detection strategies included in the text detection strategy combination pass the detection, it is determined that the detection result passes the detection. If any one of the anomaly detection results of the text detection strategies included in the text detection strategy combination fails to pass the detection, it is determined that the detection result fails to pass the detection; it is also possible to perform anomaly detection on the call text according to the combination of text detection strategies matched with the call text. If any one of the anomaly detection results of the text detection strategies included in the text detection strategy combination passes the detection, it can be determined that the detection result passes the detection. If the anomaly detection results of all the text detection strategies included in the text detection strategy combination fail to pass the detection, it can be determined that the detection result fails to pass the detection.

[0052] It should be supplemented that the operation of performing anomaly detection on the call text according to the text detection strategy matched with the call text and obtaining the detection result can be replaced by performing anomaly detection on the call text to obtain the detection result and forming a new implementation method with other processing steps provided in this embodiment.

[0053] Step S206: Input the storage address and the user text segment corresponding to the user voice segment into the emotion recognition model to obtain the user voice segment and perform emotion recognition, and obtain the probability distribution of each emotion category.

[0054] In the above, anomaly detection is performed on the call text according to the text detection strategy matched with the call text to obtain the detection result. In this step, the storage address of the user voice segment and the user text segment corresponding to the user voice segment are input into the emotion recognition model to obtain the user voice segment and perform emotion recognition, and obtain the probability distribution of each emotion category, so as to perform emotion recognition on the marked user from the voice dimension and realize the emotion analysis of the marked user.

[0055] The user text segment corresponding to the user voice segment in this embodiment may be the text description information of the user voice segment. The probability distribution of each emotion category may include a probability distribution marked by the probabilities of the user for each emotion category, such as a probability distribution marked by the probabilities of the user for positive emotions, negative emotions, and neutral emotions.

[0056] Optionally, the emotion recognition model is determined based on the anomaly level of the marked user and / or the text length of the user text segment. Specifically, the emotion recognition model can be determined in the emotion recognition model set based on the anomaly level of the marked user and / or the text length of the user text segment; in this way, by setting different emotion recognition models for marked users with different anomaly levels and / or user text segments with different text lengths, the flexibility of emotion recognition is improved.

[0057] In specific implementation, in order to improve the comprehensiveness and accuracy of emotion recognition, multi-modal emotion recognition can be combined with user voice segments and user text segments to improve the accuracy of emotion recognition; in an optional implementation manner provided in this embodiment, the emotion recognition model obtains and recognizes user voice segments in the following manner: Obtain the user voice segment from the data storage module based on the storage address; Perform emotion recognition on the marked user according to the user voice segment and the user text segment.

[0058] On this basis, in order to improve the accuracy of the voice features of the user voice segment so that the voice features can more accurately represent the user voice segment, and in order to improve the accuracy of the text features of the user text segment, voice auxiliary features and text auxiliary features can be introduced, and emotion recognition is performed on the marked user according to the voice features, voice auxiliary features of the user voice segment, the text features and text auxiliary features of the user text segment; specifically, in an optional implementation manner provided in this embodiment, in the process of performing emotion recognition on the marked user according to the user voice segment and the user text segment, the following operations are performed: Concatenate the voice features and voice auxiliary features of the user voice segment to obtain a voice concatenation feature, and perform first emotion recognition based on the voice concatenation feature to obtain the first probability distribution of each emotion category; Concatenate the text features and text auxiliary features of the user text segment to obtain a text concatenation feature, and perform second emotion recognition based on the text concatenation feature to obtain the second probability distribution of each emotion category; Calculate the probability distribution based on the first probability distribution and the second probability distribution.

[0059] Among them, the voice auxiliary feature can be a learnable feature vector, and the text auxiliary feature can also be a learnable feature vector.

[0060] Specifically, the user voice segment can be input into a pre-trained model for voice feature extraction to obtain voice features. The voice features are concatenated with voice auxiliary features, and the concatenated voice features are input into a first classifier for emotion classification to obtain the first probability distribution of each emotion category. The user text segment is input into a pre-trained text feature extraction model for text feature extraction to obtain text features. The text features are concatenated with text auxiliary features, and the concatenated text features are input into a second classifier for emotion classification to obtain the second probability distribution of each emotion category. The first probability distribution and the second probability distribution are fused to obtain a probability distribution. Specifically, the first probability distribution and the second probability distribution can be weighted to obtain the probability distribution.

[0061] Among them, the pre-trained model can be a pre-trained voice feature extraction model, such as the pre-trained model wav2vec 2.0 (Waveform to Vector Version 2.0, a self-supervised speech representation learning model) using self-supervised learning; the pre-trained text feature extraction model can include a pre-trained language model. For example, the pre-trained text feature extraction model is BERT (Bidirectional Encoder Representations from Transformers, a pre-trained language model) or RoBERTa (Robustly Optimized BERT Approach, a BERT-based language model).

[0062] Optionally, the voice auxiliary feature is a feature vector of a first preset length, and the text auxiliary feature is a feature vector of a second preset length. The first preset length and the second preset length can be the same or different; the voice auxiliary feature can be obtained by training based on a pre-trained voice feature extraction model and a first classifier, and the text auxiliary feature can be obtained by training based on a pre-trained text feature extraction model and a second classifier, so as to reduce the number of parameters to be trained through the voice auxiliary feature and the text auxiliary feature and reduce the risk of overfitting.

[0063] It should be added that the operation of inputting the user text segment corresponding to the storage address and the user voice segment into the emotion recognition model to obtain the user voice segment and perform emotion recognition, and obtaining the probability distribution of each emotion category can be replaced by inputting the user voice segment and the corresponding user text segment into the emotion recognition model for emotion recognition to obtain the probability distribution of each emotion category; or it can also be replaced by inputting the user text segment corresponding to the storage address and the user voice segment into the emotion recognition model for emotion recognition to obtain the probability distribution of each emotion category; and it can form a new implementation method with other processing steps provided in this embodiment.

[0064] Step S208, if the detection result and the probability distribution meet the resource recommendation conditions, cancel the marking of the user.

[0065] In the above process, the storage address of the user voice segment and the user text segment corresponding to the user voice segment are input into the emotion recognition model to obtain the user voice segment and perform emotion recognition, and the probability distribution of each emotion category is obtained. In this step, when the detection result and the probability distribution meet the resource recommendation conditions, the marking of the user is cancelled.

[0066] The resource recommendation conditions described in this embodiment may include that the detection result is a pass and the probability of negative emotion in the probability distribution of each emotion category is less than the probability threshold.

[0067] In specific implementation, if the detection result and the probability distribution meet the resource recommendation conditions, cancel the marking of the user. Specifically, it can be to remove the marked user from the abnormal marking table; if the detection result or the probability distribution does not meet the resource recommendation conditions, no processing is required. The detection result or the probability distribution not meeting the resource recommendation conditions may include that the detection result is a fail or the probability of negative emotion in the probability distribution is greater than or equal to the probability threshold.

[0068] Specifically, the detection result and the probability distribution meeting the resource recommendation conditions may also include that the detection result is a pass and the probability of negative emotion in the probability distribution of the user voice segment in the call voice is less than the probability threshold. In the case where the probability distribution of the previous user voice segment in the call voice does not meet the resource recommendation conditions, the emotion recognition of the subsequent user voice segment of the previous user voice segment in the call voice can be terminated, and the marked user can be marked again for manual review or secondary verification; the subsequent user voice segment of the previous user voice segment refers to the user voice segment after the previous user voice segment in the call voice.

[0069] In practical applications, there is also a need for anomaly detection for customer service. To meet the diverse needs of anomaly detection for customer service, in an alternative implementation provided in this embodiment, anomaly detection can also be performed on the call text based on the text detection strategy of the customer service determined by the resource type of the resource recommendation, and the anomaly detection result of the customer service can be obtained. Specifically, the following operations can also be performed: Based on the resource type of the resource recommendation, query the text detection strategy of the customer service in the text detection strategy set of the customer service; Perform anomaly detection on the call text according to the queried text detection strategy to obtain the anomaly detection result of the customer service.

[0070] Among them, the text detection policy set of the customer service refers to a set composed of one or more text detection policies. The text detection policy set of the customer service may include keyword detection policies, regular detection policies, and / or model detection policies. The text detection policy set may include a single text detection policy or a policy combination composed of multiple text detection policies.

[0071] Specifically, based on the resource type recommended by the resource, a text detection policy matching the resource type can be queried in the text detection policy set of the customer service, and the call text can be abnormally detected according to the queried text detection policy to obtain the abnormal detection result of the customer service; specifically, the implementation process of abnormally detecting the call text according to the queried text detection policy to obtain the abnormal detection result of the customer service is similar to the implementation process of abnormally detecting the call text according to the text detection policy matching the call text to obtain the detection result of the marked user, which will not be elaborated here.

[0072] It should be noted that after obtaining the detection result and the probability distribution of each emotion category, the emotion detection result of the marked user can be determined based on the probability distribution. Specifically, if the probability of negative emotion in the probability distribution of each emotion type is greater than the probability threshold, it can be determined that the emotion detection result is that the emotion detection fails; if the probability of negative emotion is less than or equal to the probability threshold, it can be determined that the emotion detection result is that the emotion detection passes. If the emotion detection of any user voice segment in the call voice fails, it can be determined that the emotion detection of the call voice fails; if the emotion detection of the user voice segment (which can represent all user voice segments) in the call voice passes, it can be determined that the emotion detection of the call voice passes. Therefore, the detection result and the emotion detection result of the call voice or the user voice segment can be returned to the service provider recommended by the resource; the detection result and the emotion detection result can be accessed through the voice detection system.

[0073] For example, the detection result page of the call voice by the voice detection system displays the emotion detection result of the call voice. The detection result page includes the call number of the call voice, the detection time of the call voice, the playback control of the call voice, the emotion detection result of the call voice, and / or the access control of the voice details of the call voice. Different emotion detection results of the call voice are represented by different colors and / or texts. For example, red is used to represent that the emotion detection result is that the emotion detection fails, and green is used to represent that the emotion detection result is that the emotion detection passes. After triggering the access control of the voice details, the details interface of the call voice can be displayed. The details interface includes the call number of the call voice, the start time of the voice, the end time of the voice, and / or the resource type recommended by the resource. Similarly, the emotion detection result of the user voice segment in the call voice can be displayed by the voice detection system. The detection result page of the voice segment includes the number of the user voice segment, the detection time of the user voice segment, the playback control, the emotion detection result of the user voice segment, and / or the access control of the voice details of the user voice segment. After triggering the access control of the voice details of the user voice segment, the details interface of the user voice segment can be displayed. The details interface includes the number of the user voice segment, the start time offset of the user voice segment, the end time offset, the user text segment corresponding to the user voice segment, and / or the voice role. The voice role can be the user.

[0074] In this embodiment, the call text and call voice obtained from the voice call in which the customer service recommends resources to the user can also be obtained. Optionally, the call text is obtained by querying the text based on the call parameters carried in the call completion message sent by the service provider for resource recommendation, and the call voice is obtained by downloading the voice based on the voice storage address carried in the call completion message sent by the service provider for resource recommendation. On this basis, the text detection policy can be queried based on the policy identifier carried in the call completion message, the call text can be abnormally detected according to the text detection policy to obtain the detection result, and the user text segment can be extracted from the call text, the time offset between the user text segment and the call text can be determined, and the call voice can be separated into channels to obtain the user's voice. The user voice is segmented according to the time offset, and the segmented user voice segments are uploaded to the data storage module. The storage address of the user voice segment is read from the data storage module, and the storage address of the user voice segment and the user text segment are input into the emotion recognition model to obtain the user voice segment and emotion recognition, and the probability distribution of the user in each emotion category is obtained. Based on the probability distribution, the emotion detection result of the user is determined. Specifically, if the probability of the negative emotion in the probability distribution is less than the probability threshold, it can be determined that the emotion detection of the user passes. If the probability is greater than or equal to the probability threshold, it can be determined that the emotion detection of the user fails. Among them, the process of obtaining a detection result by performing anomaly detection on the call text according to the text detection policy may include: extracting a text segment of the detection object from the call text based on the detection object of the text detection policy, and performing anomaly detection on the text segment based on the text detection policy to obtain a detection result; the detection object here may be the customer service and / or the user; on this basis, the detection result and the emotion detection result may be returned to the service provider for resource recommendation.

[0075] It should be added that the operation of canceling the label of the labeled user if the detection result and the probability distribution meet the resource recommendation condition may be replaced by canceling the label of the user; and it forms a new implementation manner with other processing steps provided in this embodiment.

[0076] It should also be added that each alternative implementation manner and each feasible execution manner in steps S202 to S208 provided in this embodiment can be independently executed according to needs, or can be combined and referenced with each other. At the same time, each specific execution step in each alternative implementation manner or each feasible execution manner can also be independently executed and combined according to needs. This embodiment does not make specific limitations in this regard; for the operations executed under the conditions of "if" or "in a certain situation" in this embodiment, the conditions of "if" and "in a certain situation" can be deleted, and the subsequent operations can be directly executed; the limitations made by "in order to" in the execution steps can also be deleted according to needs.

[0077] In summary, for one or more voice detection processing methods provided in this embodiment, first, obtain the call text obtained by the customer service for resource recommendation of the labeled user, and obtain the storage address of the user voice segment in the call voice corresponding to the call text. Secondly, on the one hand, perform anomaly detection on the call text according to the text detection policy matched by the call text to obtain a detection result. On the other hand, input the storage address and the user text segment corresponding to the user voice segment into the emotion recognition model to obtain the user voice segment and emotion recognition, and obtain the probability distribution of each emotion category, so as to perform detection from two dimensions of the call text and the voice segment, improve the comprehensiveness and accuracy of the detection, and reduce the storage amount of the user voice segment by inputting the storage address of the user voice segment into the emotion recognition model; on this basis, if the detection result and the probability distribution meet the resource recommendation condition, cancel the label of the labeled user, so as to perform detection on the labeled user marked by the customer service from two dimensions of text and voice. When the resource recommendation condition is met, the error rate of the customer service label is reduced by canceling the label of the labeled user, and then the resource recommendation can be continued for the unlabeled user, improving the success rate and quantity of resource recommendation.

[0078] The following takes the application of a voice detection processing method provided in this embodiment in a resource application recommendation scenario as an example to further illustrate the voice detection processing method provided in this embodiment. SeeFigure 5 , a voice detection and processing method applied to the resource application recommendation scenario, specifically including the following steps.

[0079] Step S502, obtain the call text of the marked user in the exception mark table, and obtain the storage address of the user voice segment in the call voice corresponding to the call text.

[0080] Optionally, the mark of the marked user is generated after the customer service makes a voice call for resource application recommendation to the user. Specifically, the mark of the marked user can be generated by the customer service; the mark of the marked user can include an exception mark; the call text of the marked user can be obtained after the customer service makes a voice call for resource application recommendation to the user; the resource application recommendation can be a recommendation to the user for a resource application project.

[0081] Step S504, perform segmentation processing on the call text to obtain text segments, and perform screening processing on the text segments based on a regular detection strategy to obtain abnormal text segments.

[0082] Step S506, input the abnormal text segments into a text detection model for text anomaly detection to obtain a detection result.

[0083] Step S508, input the storage address and the user text segment corresponding to the user voice segment into an emotion recognition model to obtain the user voice segment and perform emotion recognition, and obtain the probability distribution of each emotion category.

[0084] Step S510, if the detection result is a pass and the probability of negative emotion in the probability distribution of each emotion category is less than the probability threshold, delete the marked user from the exception mark table to re-make a voice call for resource application recommendation to the marked user.

[0085] It should be noted that any one step or any combination of steps from Step S502 to Step S510 can be replaced by the corresponding technical means provided in the above Step S202 to Step S208 according to the needs of implementation and deployment. Also, any one step or any combination of steps from Step S502 to Step S510 can be combined into a new implementation manner according to the needs of implementation and deployment; and any one step or any combination of steps from Step S502 to Step S510 can also be combined with one or more steps provided in the above Step S202 to Step S208 to form a new implementation manner according to the actual deployment requirements, or combined with one or more optional implementation manners provided in Step S202 to Step S208 to form a new implementation manner, which will not be elaborated here one by one.

[0086] An embodiment of a voice detection and processing device provided in this specification is as follows: In the above embodiments, a voice detection and processing method is provided. Correspondingly, a voice detection and processing device is also provided, which will be described below with reference to the accompanying drawings.

[0087] Referring to Figure 6 , which shows a schematic diagram of an embodiment of a voice detection and processing device provided in this embodiment.

[0088] Since the device embodiment corresponds to the method embodiment, the description is relatively simple. For the relevant parts, please refer to the corresponding description of the method embodiment provided above. The device embodiments described below are merely illustrative.

[0089] This embodiment provides a voice detection and processing device, including: An acquisition module 602, configured to acquire the call text obtained from the voice call in which the customer service recommends resources to the marked user, and acquire the storage address of the user voice segment in the call voice corresponding to the call text; An anomaly detection module 604, configured to perform anomaly detection on the call text according to the text detection policy matched by the call text to obtain a detection result; An emotion recognition module 606, configured to input the storage address and the user text segment corresponding to the user voice segment into an emotion recognition model to obtain the user voice segment and perform emotion recognition, and obtain the probability distribution of each emotion category; A marking cancellation module 608, configured to cancel the marking of the marked user if the detection result and the probability distribution meet the resource recommendation conditions.

[0090] An embodiment of a voice detection and processing device provided in this specification is as follows: Corresponding to the above-described voice detection and processing method, based on the same technical concept, one or more embodiments of this specification also provide a voice detection and processing device, which is used to execute the above-provided voice detection and processing method. Figure 7 It is a schematic diagram of the structure of a voice detection and processing device provided in one or more embodiments of this specification.

[0091] A voice detection and processing device provided in this embodiment includes: Such as Figure 7As shown, voice detection processing devices can vary significantly due to differences in configuration or performance. They can include one or more processors 701 and a memory 702. One or more application programs or data can be stored in the memory 702. Among them, the memory 702 can be for transient storage or persistent storage. The application programs stored in the memory 702 can include one or more modules (not shown in the figure), and each module can include a series of computer-executable instructions in the voice detection processing device. Further, the processor 701 can be configured to communicate with the memory 702 and execute a series of computer-executable instructions in the memory 702 on the voice detection processing device. The voice detection processing device can also include one or more power supplies 703, one or more wired or wireless network interfaces 704, one or more input / output interfaces 705, one or more keyboards 706, etc.

[0092] In a specific embodiment, the voice detection processing device includes a memory and one or more programs. One or more of the programs are stored in the memory, and one or more of the programs can include one or more modules. Each module can include a series of computer-executable instructions in the voice detection processing device and is configured to be executed by one or more processors. The one or more programs include the following computer-executable instructions for: Obtain the call text obtained from the voice call in which the customer service makes a resource recommendation to the marked user, and obtain the storage address of the user voice segment in the call voice corresponding to the call text; Perform anomaly detection on the call text according to the text detection strategy matched with the call text to obtain a detection result; Input the storage address and the user text segment corresponding to the user voice segment into an emotion recognition model to obtain the user voice segment and emotion recognition, and obtain the probability distribution of each emotion category; If the detection result and the probability distribution meet the resource recommendation conditions, cancel the marking of the marked user.

[0093] An embodiment of a computer-readable storage medium provided in this specification is as follows: Corresponding to the above-described voice detection processing method, based on the same technical concept, one or more embodiments of this specification also provide a computer-readable storage medium.

[0094] The computer-readable storage medium provided in this embodiment is used to store computer-executable instructions, and when the computer-executable instructions are executed, the following steps are implemented: Obtain the call text obtained from the voice call in which the customer service makes resource recommendations to the marked user, and obtain the storage address of the user voice segment in the call voice corresponding to the call text; Perform anomaly detection on the call text according to the text detection strategy matched with the call text to obtain a detection result; Input the storage address and the user text segment corresponding to the user voice segment into an emotion recognition model to obtain the user voice segment and perform emotion recognition, and obtain the probability distribution of each emotion category; If the detection result and the probability distribution meet the resource recommendation conditions, cancel the marking of the marked user.

[0095] It should be noted that the embodiments of a computer-readable storage medium in this specification and the embodiments of a voice detection processing method in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the corresponding method described above, and the repeated parts will not be elaborated.

[0096] An embodiment of a computer program product provided in this specification is as follows: Corresponding to the above-described voice detection processing method, based on the same technical concept, one or more embodiments of this specification also provide a computer program product.

[0097] A computer program product includes computer programs / instructions, and when the computer programs / instructions are executed by a processor, the following steps are implemented: Obtain the call text obtained from the voice call in which the customer service makes resource recommendations to the marked user, and obtain the storage address of the user voice segment in the call voice corresponding to the call text; Perform anomaly detection on the call text according to the text detection strategy matched with the call text to obtain a detection result; Input the storage address and the user text segment corresponding to the user voice segment into an emotion recognition model to obtain the user voice segment and perform emotion recognition, and obtain the probability distribution of each emotion category; If the detection result and the probability distribution meet the resource recommendation conditions, cancel the marking of the marked user.

[0098] It should be noted that the embodiments of a computer program product in this specification and the embodiments of a voice detection processing method in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the corresponding method described above, and the repeated parts will not be elaborated.

[0099] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments. For example, the device embodiment, the equipment embodiment, and the computer-readable storage medium embodiment are all similar to the method embodiment, so the description is relatively simple. Please refer to the relevant content in the method embodiment for the description of the device embodiment, the equipment embodiment, and the computer-readable storage medium embodiment.

[0100] The specific embodiments of this specification are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0101] In the 1930s, it was obvious to distinguish whether an improvement in technology was an improvement in hardware (e.g., improvements in circuit structures such as diodes, transistors, switches, etc.) or an improvement in software (improvements in method flows). However, with the development of technology, many improvements in method flows today can be regarded as direct improvements in hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structures by programming the improved method flows into the hardware circuits. Therefore, it cannot be said that an improvement in a method flow cannot be implemented with hardware entity modules. For example, a programmable logic device (PLD) (e.g., a field programmable gate array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program by themselves to "integrate" a digital system on a piece of PLD without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development writing. The original code before compilation also has to be written in a specific programming language, which is called a hardware description language (HDL), and there is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be clear that as long as the method flow is slightly logically programmed with the above-mentioned several hardware description languages and programmed into the integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.

[0102] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.

[0103] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0104] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing the embodiments of this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0105] Those skilled in the art should understand that one or more embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0106] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the specification. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable test processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable test processing device generate means for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 box or multiple boxes.

[0107] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable test processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 box or multiple boxes.

[0108] These computer program instructions can also be loaded onto a computer or other programmable test processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 box or multiple boxes.

[0109] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0110] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0111] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined in this article, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0112] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of features includes not only those features, but also other features not explicitly listed, or also includes features inherent to such process, method, commodity or device. In the absence of further restrictions, features defined by the phrase "including a ..." do not exclude the existence of other identical features in the process, method, commodity or device including the features.

[0113] One or more embodiments of the present specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of the present specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0114] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0115] The above are only examples of this document and are not intended to limit this document. For those skilled in the art, this document may have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this document shall be included within the scope of the claims of this document.

Claims

1. A speech detection and processing method, comprising: Obtaining a call text obtained from a voice call in which the customer service recommends resources to the marked user, and obtaining a storage address of a user voice segment in the call voice corresponding to the call text; Performing anomaly detection on the call text according to the text detection strategy matched by the call text to obtain a detection result; Inputting the storage address and the user text segment corresponding to the user voice segment into an emotion recognition model to obtain the user voice segment and perform emotion recognition, thereby obtaining a probability distribution of each emotion category; If the detection result and the probability distribution meet the resource recommendation condition, the marking of the marked user is cancelled.

2. According to the voice detection processing method of claim 1, the acquiring of the user voice segment and the emotion recognition comprises: Acquire the user voice segment from the data storage module based on the storage address; Emotion recognition is performed on the marked user according to the user voice segment and the user text segment.

3. The speech detection processing method according to claim 2, wherein the step of performing emotion recognition on the marked user according to the user speech segment and the user text segment comprises: Splicing the speech features and the speech auxiliary features of the user speech segment to obtain a speech splicing feature, performing a first emotion recognition based on the speech splicing feature, and obtaining a first probability distribution of each emotion category; Splicing the text features and the text auxiliary features of the user text fragment to obtain a text splicing feature, performing a second emotion recognition based on the text splicing feature, and obtaining a second probability distribution of each emotion category; The probability distribution is calculated based on the first probability distribution and the second probability distribution.

4. The voice detection and processing method according to claim 1, wherein the storage address of the user voice segment is obtained by reading from a data storage module; The storage address of the user voice segment is generated in the following manner: Creating an initial storage address based on a storage identifier of the user voice segment in the data storage module; The domain name in the initial storage address is replaced with the target domain name to obtain the storage address.

5. The voice detection and processing method according to claim 4, wherein the user voice segment is obtained by: Extracting the user text segment from the call text, determining a time offset between the user text segment and the call text, and performing channel separation on the call voice to obtain the user voice of the marked user; The user voice is segmented according to the time offset, and user voice segments obtained by the speech segmentation are uploaded to the data storage module.

6. The voice detection and processing method according to claim 5, wherein the call voice is obtained by downloading the voice based on the voice storage address; The voice storage address is obtained in the following manner: Calling a call query interface to query the call details of the voice call; The call number information is extracted from the call details, and the call number information and the call identifier are used as an interface call input to call an address query interface to query the voice storage address.

7. The voice detection processing method according to claim 1, wherein obtaining the call text obtained from the voice call in which the customer service recommends resources to the marked user comprises: Determining the number of call records stored in the data processing module, and calculating the data reading amount based on the read limit number of the data processing module and the number; The call record is read from the data processing module according to the data reading amount, and the call text is queried based on the call parameters contained in the target call record in the call record.

8. The voice detection processing method according to claim 1, wherein obtaining the call text obtained by the customer service during the voice call for recommending resources to the marked user comprises: Detect the timestamp of the newly added data partition according to the detection period, and determine whether the timestamp meets the reading condition; If satisfied, determine the number of call records stored in the newly added data partition, and if the number does not exceed a preset number threshold, query the call text according to the target call record in the call record read from the newly added data partition.

9. The speech detection processing method according to claim 1, wherein the performing abnormality detection on the call text according to the text detection strategy matched with the call text to obtain the detection result comprises: Segmenting the call text to obtain text segments, and screening the text segments based on a regular expression detection strategy to obtain abnormal text segments; The abnormal text segment is input into a text detection model to perform text anomaly detection to obtain the detection result.

10. The speech detection processing method according to claim 1, further comprising: Based on the resource type recommended by the resource, query the text detection strategy of the customer service in the text detection strategy set of the customer service; The call text is subjected to anomaly detection according to the queried text detection strategy to obtain anomaly detection results of the customer service.

11. The voice detection processing method according to claim 1, wherein the call text is obtained by performing a text query based on the call parameters carried in the call completion message sent by the service provider recommended by the resource; The text detection strategy is obtained by performing a strategy query based on the strategy identifier carried in the call completion message.

12. The speech detection processing method according to claim 1, wherein the text detection strategy for call text matching is obtained by: Determining a candidate text detection strategy in a text detection strategy set according to the resource type recommended by the resource; The text detection strategy is selected from the candidate text detection strategies based on the text length of the call text and / or the abnormality level of the marked user.

13. The speech detection processing method according to claim 12, wherein the emotion recognition model is determined based on the abnormality level of the marked user and / or the text length of the user text segment.

14. A speech detection processing device, comprising: An acquisition module is configured to acquire a call text obtained from a voice call in which the customer service recommends resources to the marked user, and to acquire a storage address of a user voice segment in the call voice corresponding to the call text; an anomaly detection module, configured to perform anomaly detection on the call text according to a text detection strategy matched by the call text, and obtain a detection result; An emotion recognition module is configured to input the storage address and the user text segment corresponding to the user voice segment into an emotion recognition model to obtain the user voice segment and perform emotion recognition, thereby obtaining a probability distribution of each emotion category; The marking cancellation module is configured to cancel the marking of the marked user if the detection result and the probability distribution meet the resource recommendation condition.

15. A speech detection and processing device, comprising: processor; and a memory configured to store computer executable instructions that, when executed, cause the processor to: Obtaining a call text obtained from a voice call in which the customer service recommends resources to the marked user, and obtaining a storage address of a user voice segment in the call voice corresponding to the call text; Performing anomaly detection on the call text according to the text detection strategy matched by the call text to obtain a detection result; Inputting the storage address and the user text segment corresponding to the user voice segment into an emotion recognition model to obtain the user voice segment and perform emotion recognition, thereby obtaining a probability distribution of each emotion category; If the detection result and the probability distribution meet the resource recommendation condition, the marking of the marked user is cancelled.

16. A computer-readable storage medium for storing computer-executable instructions, wherein the computer-executable instructions implement the steps of the method of claim 1 when executed.