Speech processing method and apparatus, storage medium, and electronic device
By performing semantic recognition on target speech data to obtain tags, and then performing speech matching based on the tags, the problem of cumbersome speech processing and poor timeliness in existing technologies is solved, and fast and real-time similar speech feedback is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
- Filing Date
- 2022-11-02
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies for searching for similar voices from large amounts of user voice data in user transaction scenarios involve cumbersome processing, high computational load, poor timeliness, and inability to achieve real-time feedback.
By performing semantic recognition processing on the target speech data to obtain target semantic tags, and then performing speech matching on the reference speech set based on these tags, the traditional speech-text clustering method is avoided, enabling fast matching and real-time feedback.
The speech processing flow has been optimized, reducing the amount of computation and improving timeliness, enabling real-time speech processing and real-time feedback of similar speech.
Smart Images

Figure CN115841810B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a voice processing method, apparatus, storage medium and electronic device. Background Technology
[0002] Speech processing technology is an important branch of information processing and a core technology in current speech recognition and evaluation systems. With technological advancements, its applications are becoming increasingly widespread. In various scenarios such as shopping, travel, and audiovisual experiences, user voice data can reveal valuable common questions, providing insights for improving the user experience in these scenarios. Summary of the Invention
[0003] This specification provides a voice processing method, apparatus, storage medium, and electronic device, the technical solutions of which are as follows:
[0004] Firstly, this specification provides a speech processing method, the method comprising:
[0005] Semantic recognition processing is performed on the target speech data to obtain at least one target semantic label corresponding to the target speech data;
[0006] Based on the at least one target semantic label, speech matching processing is performed on the reference speech set to obtain similar speech data corresponding to the target speech data.
[0007] Secondly, this specification provides a voice processing device, the device comprising:
[0008] The label determination module is used to perform semantic recognition processing on the target speech data to obtain at least one target semantic label corresponding to the target speech data.
[0009] The speech matching module is used to perform speech matching processing on the reference speech set based on the at least one target semantic label to obtain similar speech data corresponding to the target speech data.
[0010] Thirdly, this specification provides a computer storage medium storing a plurality of instructions adapted for loading by a processor and executing the above-described method steps.
[0011] Fourthly, this specification provides an electronic device that may include: a processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and to execute the above-described method steps.
[0012] Fifthly, this specification provides a computer program product that stores at least one instruction, which is loaded by a processor and executes the above-described method steps.
[0013] The beneficial effects of the technical solutions provided in some embodiments of this description include at least the following:
[0014] In one or more embodiments of this specification, at least one target semantic tag is determined by performing semantic recognition processing on the target speech data. Then, based on the at least one target semantic tag, speech matching processing is performed on several reference speech data in the reference speech set to obtain several similar speech data corresponding to the target speech data. Throughout the speech processing stage, the method of clustering a large amount of speech text is avoided. Based on the target semantic tag of the target speech data, rapid matching of the reference speech set can be achieved, optimizing the speech processing flow and reducing the computational load. Real-time speech processing can be achieved to provide real-time feedback on similar speech, improving the timeliness of speech processing. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in this specification or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a schematic diagram of a voice processing system provided in this manual;
[0017] Figure 2 This is a flowchart illustrating a speech processing method provided in this manual;
[0018] Figure 3 This is a flowchart illustrating a speech processing method provided in this manual;
[0019] Figure 4 This is a schematic diagram of the structure of a voice processing device provided in this specification;
[0020] Figure 5 This is a schematic diagram of the structure of a coefficient determination module provided in this manual;
[0021] Figure 6 This is a schematic diagram of the structure of a vector building unit provided in this specification;
[0022] Figure 7 This is a schematic diagram of the structure of a speech processing unit provided in this specification;
[0023] Figure 8 This is a schematic diagram of the structure of a coefficient determination module provided in this manual;
[0024] Figure 9 This is a schematic diagram of the structure of a result determination module provided in this specification;
[0025] Figure 10 This is a schematic diagram of the structure of an electronic device provided in this specification. Detailed Implementation
[0026] The technical solutions in this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0027] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. In the description of this application, it should be noted that, unless otherwise expressly specified and limited, "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances. Furthermore, in the description of this application, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist; for example, A and / or B can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.
[0028] In related technologies, when users are in transactional scenarios such as shopping, travel, or audiovisual experiences, there are situations where it is necessary to search for similar voices from a large amount of user voice data. For example, searching for similar voices of a target voice to further analyze and improve the user experience based on these similar voices and the target voice. Typically, both the target voice data and the large amount of user voice data are converted into speech-to-text, and this is achieved by performing pure text clustering on all speech-to-text data. Common text clustering algorithms cluster all speech-to-text data, resulting in multiple sets of similar speech-to-text data within at least one category. The original voice data corresponding to these multiple sets of similar speech-to-text data within a certain category can be considered as a group of similar voices. However, this text clustering approach is cumbersome in terms of voice processing, computationally intensive in semantic processing, and has poor timeliness. It typically requires offline computation and cannot reflect real-time dynamic similarity scenarios.
[0029] Please see Figure 1 This is a schematic diagram of a speech processing system provided in this specification. Figure 1 As shown, the voice processing system may include at least a client cluster and a service platform 100.
[0030] The client cluster may include at least one client, such as Figure 1 As shown, it specifically includes client 1 corresponding to user 1, client 2 corresponding to user 2, ..., client n corresponding to user n, where n is an integer greater than 0.
[0031] Each client in a client cluster can be an electronic device with communication capabilities, including but not limited to: wearable devices, handheld devices, personal computers, tablets, in-vehicle devices, smartphones, computing devices, or other processing devices connected to a wireless modem. Electronic devices may have different names in different networks, such as: user equipment, access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent or user device, cellular phone, cordless phone, personal digital assistant (PDA), and electronic devices in 5G networks or future evolved networks.
[0032] The service platform 100 can be a standalone server device, such as a rack-mount, blade, tower, or cabinet-type server device, or a workstation, mainframe, or other hardware device with strong computing power; or it can be a server cluster composed of multiple servers. The servers in the service cluster can be composed in a symmetrical manner, wherein each server is functionally and hierarchically equivalent in the transaction chain, and each server can provide services independently. The independent provision of services can be understood as not requiring the assistance of other servers.
[0033] In one or more embodiments of this specification, the service platform 100 can establish a communication connection with at least one client in the client cluster, and complete the interaction of voice data during voice processing based on this communication connection. Illustratively, client 1 in the client cluster can upload user voice data in a corresponding transaction scenario, such as target voice data or reference voice data, to the service platform 100 through the communication connection. The service platform 100 obtains the target voice data uploaded by client 1, executes the voice processing method, and then determines similar voice data corresponding to the target voice data.
[0034] It should be noted that the service platform 100 establishes a communication connection with at least one client in the client cluster for interactive communication via a network. This network can be a wireless network or a wired network. Wireless networks include, but are not limited to, cellular networks, wireless LANs, infrared networks, or Bluetooth networks. Wired networks include, but are not limited to, Ethernet, universal serial bus (USB), or controller area networks. In one or more embodiments of the specification, technologies and / or formats including Hyper Text Markup Language (HTML), Extensible Markup Language (XML), etc., are used to represent data exchanged over the network (such as target compressed packets). Furthermore, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), and Internet Protocol Security (IPsec) can be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can be used to replace or supplement the aforementioned data communication technologies.
[0035] The speech processing system embodiments provided in this specification and the speech processing methods described in one or more embodiments belong to the same concept. The execution entity corresponding to the speech processing method involved in one or more embodiments of this specification can be the aforementioned service platform 100; the execution entity corresponding to the speech processing method involved in one or more embodiments of this specification can also be the electronic device corresponding to the client, specifically determined based on the actual application environment. The implementation process of the speech processing system embodiments can be detailed in the following method embodiments, and will not be repeated here.
[0036] based on Figure 1 The following is a detailed description of the speech processing methods provided by one or more embodiments of this specification, as illustrated in the scene diagram.
[0037] Please see Figure 2 This document provides a flowchart illustrating a speech processing method according to one or more embodiments. This method can be implemented using a computer program and can run on a background investigation device based on the von Neumann architecture. The computer program can be integrated into an application or run as a standalone utility application. The speech processing device can be a service platform.
[0038] Specifically, the speech processing method includes:
[0039] S102, perform semantic recognition processing on the target speech data to obtain at least one target semantic label corresponding to the target speech data.
[0040] Speech data refers to the sound data of language. It is the carrier of the language symbol system. The target speech data acquired by electronic devices is actually a signal wave emitted by the user in the corresponding transaction scenario. Illustratively, in the target speech data acquisition stage, the acquired signal wave first needs to be preprocessed and framed. At this point, the speech becomes many small segments. Then, a time-domain transformation is performed on the speech signal waveform. A commonly used method for time-domain transformation is to extract Mel-Frequency Cepstral Coefficients (MFCC) feature information. Based on the physiological characteristics of the human ear, each frame of the signal waveform is transformed into a multi-dimensional vector, which can be simply understood as containing the content information of that frame of speech.
[0041] Optionally, the target voice data can be real user feedback data in a specific transaction scenario, in which a large amount of user voice data can be collected. These voice data and the currently collected target semantic data share common features in the voice information dimension. These common features are of reference significance for improving and optimizing the corresponding transaction functions in the transaction scenario. Based on this, in practical applications, there may be a need to query similar voice data from other users for the current target voice data, so as to improve and optimize the corresponding transaction functions in the transaction scenario based on a type of voice data (target voice data and similar voice data corresponding to the target voice data).
[0042] Optionally, the target voice data can be voice data of user feedback, voice data of transaction consultation, voice data of transaction evaluation, etc. in a certain transaction scenario.
[0043] The target semantic label is a semantic label corresponding to the semantic feature information obtained from the semantic recognition processing of the target speech data. This semantic feature information establishes a mapping relationship with the original target speech data in the form of a semantic label. In this embodiment, the semantic recognition processing can be understood as taking the target speech data as the recognition object. After acquiring the collected target speech data, the service platform automatically recognizes and understands the user's spoken speech in the transaction scenario by performing speech signal processing (such as preprocessing the target speech data) and semantic recognition on the target speech data, and identifies the semantic feature information in the speech to generate the corresponding target semantic label.
[0044] In one feasible implementation, the target speech data can be input into a label classification model, and the label classification model can perform semantic recognition processing on the target speech data to output at least one target semantic label for the target speech data.
[0045] Before semantic recognition processing, it is usually necessary to acquire a large amount of speech sample data to extract semantic feature information to train an initial label classification model. Semantic feature information is the unique semantic attribute of unstructured data expressed in natural language. Taking a paper as an example, semantic feature information includes semantic elements such as the author's creative intent, data theme description, and underlying feature meaning. The semantic feature information is a variety of features that can express the semantics of the object itself and its semantics in the environment. Taking scene speech data of a certain object as an example, the semantic feature information can be the order of constituent elements, word order, word sentiment information, mutual information, etc.
[0046] The constituent elements can be understood as the smallest constituent units that make up a paragraph. Taking the Chinese language as an example, the smallest constituent unit is the pronunciation of each character.
[0047] The word order is the sequence of words that a user uses to express a sentence (a single meaning).
[0048] The emotional information of a word is the emotional meaning that the user expresses in this sentence. The emotional meaning can be understood as whether the word is high or low, positive or negative, joyful or sad, etc.
[0049] Mutual information refers to the statistical independence relationship between a word or character and a category. Mutual information is often used to measure the reciprocity between two objects.
[0050] In this embodiment, the semantic tags generated from the semantic feature information extracted from the speech data include, but are not limited to, keyword tags corresponding to keyword information, word frequency distribution tags corresponding to word frequency distribution information, entity tags corresponding to grammatical entity information, and topic tags corresponding to semantic topics.
[0051] In this embodiment, for the entire target speech data, the target semantic tags include, but are not limited to, target keyword tags, target word frequency distribution tags, target entity tags, and target topic tags for the transaction scenario or transaction during the transaction function process in the user experience transaction scenario.
[0052] Specifically, the service platform first needs to train the initial label classification model. It can obtain a large amount of voice sample data (in some embodiments, voice sample data can also be called reference voice data) input by users in transaction scenarios from an existing voice database, and / or obtain voice sample data recorded by a recording device in the actual language environment. Then, it extracts key semantic segments through the initial label classification model, aggregates the key segments to extract sample semantic labels from the aggregated segments and outputs at least one sample semantic label.
[0053] In one or more embodiments of this specification, the label classification model is a neural network model. The neural network model can be implemented by fitting one or more of the following models: Convolutional Neural Network (CNN), Deep Neural Network (DNN), Recurrent Neural Networks (RNN), embedding model, Gradient Boosting Decision Tree (GBDT) model, Logistic Regression (LR) model, etc.
[0054] Understandably, after acquiring a large amount of speech sample data, key semantic segments are extracted through an initial label classification model to output several candidate semantic labels. Based on expert services, the candidate semantic labels corresponding to the speech sample data can be adjusted to obtain several standard semantic labels for the speech sample data. The label classification model is then trained based on the speech sample data with standard semantic labels. During model training, the model parameters and model architecture can be adjusted through the standard semantic labels to obtain a well-trained label classification model.
[0055] Optionally, adjusting the semantic labels of the candidate semantic labels corresponding to the speech sample data based on the expert service can be done by: rewriting the label names, merging and unifying the labels, attaching the label attributes, and confirming the label accuracy of several candidate semantic labels corresponding to the speech sample data, so as to obtain several standard semantic labels for the speech sample data, and then saving the standard semantic labels that conform to the specifications after associating them with the speech sample data (such as storing them in a database).
[0056] Optionally, tag renaming can be understood as modifying the tag name of a candidate semantic tag identified by the tag classification model to a tag name that conforms to the specification. For example, the candidate tag "QR code not found" can be modified to "QR code not found" according to the tag specification to generate a standard semantic tag that conforms to the tag naming specification.
[0057] Optionally, tag merging and unification can be understood as: multiple similar candidate semantic tags can be uniformly modified into a single standard semantic tag. For example, three candidate semantic tags: "health code turns red", "health code turns green", and "health code turns yellow" can all be merged into a single standard semantic tag: "health code color change".
[0058] Optionally, tag attribute mounting can be understood as: adjusting at least one attribute element contained in the candidate semantic tag, such as removing or adding a tag attribute from the candidate semantic tag.
[0059] Optionally, the label accuracy identifies labels with errors among the candidate semantic labels. These erroneous labels and the original speech sample data are then labeled and subsequently input into the model for training.
[0060] Understandably, in this application, semantic recognition can be directly performed on the speech data to obtain semantic tags. That is, it is not necessary to convert the speech data into speech text and then perform semantic recognition on the speech text. Furthermore, the aforementioned candidate semantic tags, target semantic tags, etc., can all be speech type tags or text type tags.
[0061] S104, based on the at least one target semantic label, perform speech matching processing on the reference speech set to obtain similar speech data corresponding to the target speech data.
[0062] In one feasible implementation, the reference speech set consists of several reference speech (data) sets, and each reference speech (data) can be associated with at least one reference semantic tag.
[0063] Optionally, the reference speech (data) refers to all or part of the speech data input by other users in similar scenarios, which the service platform acquires. This could include user feedback speech data, inquiry speech data, evaluation speech data, etc., in similar scenarios (such as payment, shopping, and health code scenarios). Furthermore, the service platform can pre-determine at least one reference semantic tag by performing semantic recognition on the reference speech (data) and associate these reference semantic tags with the reference speech data. This then forms a reference speech set consisting of several reference speech data points.
[0064] Optionally, the reference speech data can be input into the label classification model, which outputs at least one reference semantic label for the reference speech data.
[0065] In one feasible implementation, speech matching processing of the reference speech set based on the at least one target semantic label can be understood as calculating the semantic similarity between the target speech data and the reference speech data based on the semantic label. The semantic similarity is calculated based on the semantic labels corresponding to each pair of speech data, and speech data with semantic similarity greater than a threshold is taken as similar speech data to the target speech data.
[0066] In one feasible implementation, the service platform performs speech matching processing on the reference speech set based on the at least one target semantic label to obtain similar speech data corresponding to the target speech data. This can be achieved by the service platform determining at least one target semantic label corresponding to the target speech data and constructing target semantic query rules based on each of the target semantic labels.
[0067] The service platform can obtain reference semantic tags corresponding to at least one reference speech in the reference speech set, and perform speech matching processing based on the reference semantic tags corresponding to each reference speech using the target semantic query rules to obtain similar speech data corresponding to the target speech data.
[0068] Optionally, the target semantic query rule can be based on constructing a semantic query search formula based on each target semantic label, and then searching for similar speech with the same label in the reference speech set based on the semantic query search formula.
[0069] Optionally, the semantic query search formula can be to construct logical relationships based on multiple target semantic tags to generate a logical search formula for querying similar speech. For example, some semantic tags can be constructed as logical "AND" relationships, or some target semantic tags can be constructed as logical "OR" relationships, and so on.
[0070] Optionally, there can be multiple similar speech data corresponding to the target speech data. In the process of determining the similar speech data corresponding to the target speech data, the volume of the similar speech corresponding to the target speech data, that is, the number of similar speech data, can be calculated.
[0071] In one or more embodiments of this specification, at least one target semantic tag is determined by performing semantic recognition processing on the target speech data. Then, based on the at least one target semantic tag, speech matching processing is performed on several reference speech data in the reference speech set to obtain several similar speech data corresponding to the target speech data. Throughout the speech processing stage, the method of clustering a large amount of speech text is avoided. Based on the target semantic tag of the target speech data, rapid matching of the reference speech set can be achieved, optimizing the speech processing flow and reducing the computational load. Real-time speech processing can be achieved to provide real-time feedback on similar speech, improving the timeliness of speech processing.
[0072] Please see Figure 3 , Figure 3 This is a flowchart illustrating another embodiment of a speech processing method proposed in one or more embodiments of this specification. Specifically:
[0073] S202: Perform semantic recognition processing on the target speech data to obtain at least one target semantic label corresponding to the target speech data;
[0074] For details, please refer to one or more embodiments in this specification, which will not be repeated here.
[0075] Understandably, topic matching can be performed on the target speech data based on at least one target semantic label and at least one topic template rule to obtain the target topic corresponding to the target speech data. Then, similar speech data corresponding to the target topic can be obtained from the reference speech set based on the target topic. See the steps below.
[0076] S204: Obtain at least one semantic tag rule corresponding to a topic template rule;
[0077] Understandably, topics correspond to corresponding topic template rules, and topic template rules at least correspond to semantic tag rules under that topic. For example, if there are n topics, n topic template rules are pre-established for each of the n topics, and each topic template rule corresponds to a semantic tag rule for that topic.
[0078] Understandably, the target speech data is matched with the semantic label rules in the topic template rules set for several reference topics in order to determine the matching target label rules.
[0079] In one or more embodiments of this specification, the semantic tagging rule is a tag logic rule, that is, a tag logic rule composed of multiple semantic tags. The tag logic rule can be represented in the form of a tag logic expression, such as a logical expression composed of multiple tags (which can also be understood as a logical operation formula). For example, it can be set for a certain topic: constructing certain semantic tags related to the topic as a logical "AND" relationship, or constructing certain target semantic tags as a logical "OR" relationship, and so on. Schematic, for topic A, the semantic tagging rule under the topic template rule can be set as: semantic tag A "AND" semantic tag B, that is, the voice data of topic A corresponds to both semantic tag A and semantic tag B.
[0080] Optionally, the above-mentioned tag logic rules can be logical expressions based on at least one semantic tag and logical operators.
[0081] In a specific implementation scenario, several reference topics are set, and semantic tag rules are then set for each of these reference topics under topic template rules. Each semantic tag rule is based on at least one reference semantic tag and logical operators. The reference topics and their corresponding topic template rules can then be stored.
[0082] Optionally, some or all of each reference topic can be created by pre-acquiring a large amount of reference speech data;
[0083] Optionally, some or all of each reference topic can be created manually based on expert services;
[0084] Optionally, some or all of each reference topic can be obtained by classifying several reference semantic tags and extracting topics (or rewriting topics) based on the classified semantic categories. For example, several reference semantic tags can be divided into scenario obstacle categories, scenario demand categories, scenario consultation categories, scenario status categories, etc., based on actual business scenarios. Then, under any tag category, there are several semantic tags belonging to that category. Then, topic extraction can be performed based on a certain tag category to determine the reference topic, and then the semantic tag rules in the topic template rules can be constructed based on the several semantic tags under that tag category.
[0085] S206: Determine the target label rule that matches the at least one target semantic label from the semantic label rules.
[0086] Understandably, each reference topic corresponds to at least one semantic labeling rule. Semantic labeling rules are used to characterize the element relationships between semantic labels corresponding to speech data under a certain topic. Semantic labeling rules can at least reflect the number of semantic labels that a reference topic must have, at least one of the number of semantic labels that a reference topic can have, the semantic labels that a reference topic should not have, and the number of semantic labels under a reference topic.
[0087] Understandably, by identifying several target semantic tags corresponding to the target speech data, and then, based on these tags, determining whether each of the "several target semantic tags" collectively satisfies the semantic tag rules corresponding to the reference topic; illustratively, if the "several target semantic tags" do not satisfy the semantic tag rules corresponding to a certain reference topic, then the target speech data can usually be determined not to belong to "that reference topic" from the perspective of topic semantic tags. Illustratively, if the "several target semantic tags" satisfy the semantic tag rules corresponding to a certain reference topic, then the target speech data can usually be determined to belong to "that reference topic" from the perspective of topic semantic tags.
[0088] Optionally, a target semantic data is matched with several target semantic tags and multiple target tag rules to determine whether the target tag rules are met; several target semantic tags can simultaneously satisfy at least one of the "multiple target tag rules", that is, from the dimension of topic semantic tags, it can be determined that it belongs to or satisfies the semantic tag rules corresponding to multiple "reference topics".
[0089] In a specific implementation scenario, the semantic tagging rule can be a tag logic rule. Multiple reference topics can be set in advance, for example, the number of reference topics is n. Tag logic rules under the topic template rules corresponding to the reference topics are set. The tag logic rules corresponding to at least one topic template rule are obtained. Then, it is detected whether the "at least one target semantic tag" matches each of the tag logic rules, that is, whether these target semantic tags satisfy the tag logic rules corresponding to each reference topic. Then, the tag matching result can be obtained.
[0090] For illustration purposes, the label matching result can be the label matching result for each reference topic. If the i target semantic labels of the target speech data satisfy the label logic rules corresponding to the reference topic x, then the label matching result for the reference topic x is: satisfying the reference topic x; if the i target semantic labels of the target speech data do not satisfy the label logic rules corresponding to the reference topic x, then the label matching result for the reference topic x is: not satisfying the reference topic x.
[0091] Indicatively, the target label rule for matching the at least one target semantic label can then be determined based on the label matching result.
[0092] Indicatively, the target label rule is the semantic label rule satisfied by "the at least one target semantic label".
[0093] S208: Obtain the first topic corresponding to the target tag rule;
[0094] The first topic is also the reference topic corresponding to the target label rule. For example, if the reference topic corresponding to the target label rule is scenario obstacle topic A, then scenario obstacle topic A is also the first topic.
[0095] In one feasible implementation, the target topic corresponding to the target speech data can be determined based on a first topic. That is, the first topic corresponding to the target tagging rule is directly taken as the target topic.
[0096] In one feasible implementation, the following steps may also be performed to determine a second topic and a third topic, and to determine the target topic corresponding to the target speech data based on at least one of the first topic, the second topic, and the third topic.
[0097] S210: Obtain at least one key information rule and / or mixed information rule corresponding to a topic template rule;
[0098] The key information rule refers to the key characteristics of unstructured data expressed in natural language, specifically related to specific topics. A key information rule is a topic rule composed of at least one key piece of information (such as key text or key speech signal). Key information includes, for example, the topic intent corresponding to the reference topic, topic decision information, topic suggestion information, topic central words, reference key speech signals, and key text. Illustratively, a key information rule may consist of at least one key piece of information and a logical relation. In some embodiments, the key information rule may also include word order, key speech signal order, keyword / sentence spacing rules, and spacing rules between key speech signals.
[0099] Understandably, each topic corresponds to a specific topic template rule, and each topic template rule at least corresponds to a key information rule for that topic. For example, if there are n topics, n topic template rules are pre-established for each of the n topics, and each topic template rule corresponds to a key information rule for that topic. Optionally, the aforementioned key information rule can be a logical expression based on at least one key piece of information (element) and logical operators.
[0100] The hybrid information rule can be a template rule composed of semantic tags and key information, that is, a template rule composed of several semantic tags and several key information. The hybrid information rule can at least reflect the topic characteristics between semantic tags and key information under a reference topic, such as semantic tags and key information (such as keywords and key sentences) that must exist under a reference topic, semantic tags that must exist under a reference topic and key information (such as keywords and key sentences) that do not exist, etc.
[0101] Understandably, each topic corresponds to a specific topic template rule, and each topic template rule corresponds to at least one mixed information rule for that topic. For example, if there are n topics, n topic template rules are pre-established for each of the n topics, and each topic template rule corresponds to a mixed information rule for that topic. Optionally, the aforementioned mixed information rule can be a logical expression based on at least one semantic tag (element), at least one key information (element), and logical operators, such as the logical relationship between the semantic tag (element) and the key information (element) (e.g., logical AND, logical NOT, logical OR, etc.).
[0102] In one feasible implementation, at least one key information rule corresponding to a topic template rule can be obtained, and each topic template rule corresponds to a reference topic.
[0103] In one feasible implementation, at least one mixed information rule corresponding to a topic template rule can be obtained, and each topic template rule corresponds to a reference topic.
[0104] In one feasible implementation, at least one key information rule and mixed information rule corresponding to a topic template rule can be obtained, and each topic template rule corresponds to a reference topic.
[0105] Understandably, the topic template rules for reference topics include one or more fitting methods such as key information rules, mixed information rules, and semantic tag rules. Illustratively, after obtaining the corresponding topic template rules, the topic corresponding to the relevant information rules is determined.
[0106] S212: Determine the target key information rule matching the target speech data from each of the key information rules, and obtain the second topic corresponding to the target key information rule;
[0107] Understandably, after obtaining at least one key information rule corresponding to a topic template rule, the target key information rule matching the target speech data can be determined from each of the key information rules, thereby determining the second topic corresponding to the target key information rule.
[0108] Indicatively, each reference topic corresponds to at least one key information rule. The key information rule is used to characterize the element relationship between key information (elements) corresponding to the speech data under a certain topic. The key information rule can at least provide feedback on one or more of the following types of key information (elements) that must be present under the reference topic, at least one of the key information (elements) that can be present under the reference topic, key information (elements) that should not be present, the number of key information (elements), word order, key speech signal order, spacing rules between keywords / sentences, and spacing rules between key speech signals.
[0109] In one feasible implementation, the target speech data can be processed by information detection based on each of the key information rules to obtain information detection results; then, the target key information rules matching the target speech data can be determined based on the information detection results.
[0110] In illustrative terms, key information rules can be constructed from key information logical expressions. By constructing these expressions, it's possible to detect whether corresponding speech data conforms to key information rules, such as whether it possesses at least one of several key information elements, whether it lacks the corresponding key information, whether the quantity of key information elements is satisfied, whether the word order is satisfied, whether the order of key speech signals is satisfied, whether the spacing rules between keywords / sentences are satisfied, and whether the spacing rules between key speech signals are satisfied. All of the above can be implemented based on constructing key information logical expressions to achieve rule matching of target speech data in transaction scenarios.
[0111] Understandably, by using key information rules set for multiple reference topics, the target speech data is processed one by one to determine whether it meets the key information rules. For example, if the target speech data does not meet the key information rules corresponding to a certain reference topic, then the target speech data can usually be determined not to belong to "that reference topic" from the perspective of topic key information. Conversely, if the target speech data meets the key information rules corresponding to a certain reference topic, then the target speech data can usually be determined to belong to "that reference topic" from the perspective of topic key information.
[0112] In a specific implementation scenario, multiple reference topics can be pre-set, for example, the number of reference topics is n, and key information rules under the topic template rules corresponding to the reference topics are set; by obtaining at least one key information rule corresponding to the topic template rule; information detection processing is performed on the target speech data one by one to determine whether the key information rules are met, and then the key information detection result can be obtained;
[0113] For illustration purposes, if the target speech data satisfies the key information rules corresponding to the reference topic x, then the key information detection result for the reference topic x is: satisfies the reference topic x; if the target speech data does not satisfy the key information rules corresponding to the reference topic x, then the key information detection result for the reference topic x is: does not satisfy the reference topic x.
[0114] In a specific implementation scenario, key information rules can be information logic rules, information sequence rules, or information spacing rules;
[0115] The information logic rules can be understood as the logical relationships between several key information (elements) under the reference topic. Key information (elements) can be keywords, key texts, key sentences, etc., and logical relationships can be logical AND, logical OR, logical NOT, etc.
[0116] The information order rule can be understood as the order or temporal relationship between several key information (elements) under the reference topic. For example, for a certain reference topic, the information order rule can be that the order of keyword A is before the order of key text B, and the order of keyword B is after the order of keyword C.
[0117] The information spacing rule can be understood as the interval or spacing relationship between several key information (elements) under the reference topic. For example, for a certain reference topic, the information spacing rule can be that the interval between keyword A and key text B is x character units, and the interval between keyword B and keyword C should be within a certain unit range.
[0118] In illustrative terms, the target speech data is processed for key information detection based on each of the key information rules to obtain the information detection result. Specifically, the detection can be performed on at least one pair of the target speech data based on the information logic rules, information order rules, and information spacing rules included in the key information rules.
[0119] Optionally, the target speech data can be subjected to key information logic detection processing based on the information logic rules corresponding to each of the key information rules to obtain the logic detection results;
[0120] Specifically, multiple reference topics can be set in advance, for example, the number of reference topics is n. Key information rules under the topic template rules corresponding to the reference topics can be set. Key information rules can include information logic rules. By obtaining the information logic rules corresponding to several key information rules, and then detecting whether the target speech data matches each information logic rule, that is, whether the target speech data satisfies the information logic rules corresponding to each reference topic, the logic detection result can be obtained.
[0121] For illustrative purposes, the logic detection result can be the matching result based on several key information (elements) logic dimensions for each reference topic. If the target speech data satisfies the information logic rules corresponding to the reference topic x, then the logic detection result for the reference topic x is: satisfies the reference topic x; if the target speech data does not satisfy the information logic rules corresponding to the reference topic x, then the logic detection result for the reference topic x is: does not satisfy the reference topic x.
[0122] Optionally, the target speech data can be processed by key information sequence detection based on the information sequence rules corresponding to each key information rule to obtain the sequence detection result;
[0123] Specifically, multiple reference topics can be set in advance, for example, the number of reference topics is n. Key information rules under the topic template rules corresponding to the reference topics can be set. Key information rules can include information order rules. By obtaining the information order rules corresponding to several key information rules, and then detecting whether the target speech data matches each information order rule, that is, whether the target speech data satisfies the information order rules corresponding to each reference topic, the order detection result can be obtained.
[0124] For illustration purposes, the logical detection result can be the matching result based on several key information (elements) logical dimensions for each reference topic. If the target speech data satisfies the information order rule corresponding to the reference topic x, then the order detection result for the reference topic x is: satisfies the reference topic x; if the target speech data does not satisfy the information order rule corresponding to the reference topic x, then the order detection result for the reference topic x is: does not satisfy the reference topic x.
[0125] Optionally, the target speech data can be processed by key information spacing detection based on the information spacing rules corresponding to each key information rule to obtain the spacing detection result.
[0126] Specifically, multiple reference topics can be set in advance, for example, the number of reference topics is n. Key information rules under the topic template rules corresponding to the reference topics can be set. Key information rules can include information spacing rules. By obtaining the information spacing rules corresponding to several key information rules, and then detecting whether the target speech data matches each information spacing rule, that is, whether the target speech data satisfies the information spacing rules corresponding to each reference topic, the spacing detection result can be obtained.
[0127] For illustration purposes, the logical detection result can be the matching result based on several key information (elements) logical dimensions for each reference topic. If the target speech data satisfies the information spacing rule corresponding to the reference topic x, then the spacing detection result for the reference topic x is: satisfies the reference topic x; if the target speech data does not satisfy the information spacing rule corresponding to the reference topic x, then the spacing detection result for the reference topic x is: does not satisfy the reference topic x.
[0128] Understandably, information detection results, depending on the specific key information detection method, can include at least one of the following types: key information logic detection results, information sequence detection results, and text spacing detection results.
[0129] S214: Determine the target mixed information rule that matches the target speech data from each of the mixed information rules, and obtain the third topic corresponding to the target mixed information rule;
[0130] In one feasible implementation, the target mixed information rule for matching the target speech data is determined from each of the mixed information rules. Specifically, the mixed information rule can be understood as the logical rule of information elements that correspond to both semantic tags and key information. The mixed information rule can at least reflect the topic characteristics between semantic tags and key information under a reference topic, such as semantic tags and key information (e.g., keywords, key sentences) that must exist under a reference topic, semantic tags that must exist under a reference topic, and key information (e.g., keywords, key sentences) that do not exist.
[0131] Specifically: the target speech data can be processed by element logic detection based on the logical rules of each information element to obtain the element logic detection result;
[0132] For example, given n topics, n topic template rules are pre-established for each topic. Each topic template rule corresponds to a mixed information rule for that topic, i.e., an information element rule. Optionally, these information element rules can be logical expressions based on at least one semantic tag (element), at least one key information (element), and logical operators. For example, the logical relationship between semantic tags (elements) and key information (elements) can be logically defined (e.g., AND, NOT, OR, etc.). For instance, there are also "AND" and "OR" logical configurations between semantic tags and text keywords.
[0133] Specifically, multiple reference topics can be set in advance, for example, the number of reference topics is n. The element logic rules under the topic template rules corresponding to the reference topics are set. The element logic rules corresponding to several reference topics are obtained. Then, it is detected whether the target speech data matches each element logic rule, that is, whether the target speech data satisfies the element logic rules corresponding to each reference topic. Then the element logic detection result can be obtained.
[0134] For illustration, the element logic detection result can be the matching result of logical dimensions between several key information (elements) and several semantic tags for each reference topic. If the target speech data satisfies the element logic rules corresponding to the reference topic x, then the element logic result for the reference topic x is: satisfies the reference topic x; if the target speech data does not satisfy the element logic rules corresponding to the reference topic x, then the element logic result for the reference topic x is: does not satisfy the reference topic x.
[0135] Specifically: Then, based on the element logic detection results, the target mixed information rule for matching the target speech data is determined. The target mixed information rule is also the mixed information rule for matching the target speech data, and the third topic corresponding to the target mixed information rule is obtained. For example, if the reference topic corresponding to the target mixed information rule is scene obstacle topic A, then scene obstacle topic A is also the third topic.
[0136] S216: Determine the target topic corresponding to the target speech data based on at least one of the first topic, the second topic, and the third topic.
[0137] As an illustration, the target topic corresponding to the target speech data can be determined based on the first topic.
[0138] As an illustration, the target topic corresponding to the target speech data can be determined based on at least one of the first topic and the second topic.
[0139] For example, the intersection of the first topic, the second topic, and the third topic can be taken, and the common topic can be used as the target topic. In other words, the target topic corresponding to the target speech data can be accurately determined through the above steps, so as to search for reference speech of the same type or the same topic based on the target topic, and these reference speech are regarded as similar speech.
[0140] S218: Perform speech matching processing on the reference speech set based on the target topic to obtain similar speech data corresponding to the target speech data.
[0141] In one feasible implementation, a reference topic corresponding to at least one reference speech data in the reference speech set can be obtained; this can be understood as: the reference speech set consists of several reference speech data, and each reference speech data corresponds to at least one reference topic; then, similar speech data corresponding to the target topic can be determined from the reference speech set based on the reference topics corresponding to each of the reference speech data.
[0142] Optionally, topic matching processing can be performed on at least one reference speech data in the reference speech set beforehand based on the topic template rules to obtain the reference topic corresponding to the reference speech data, so as to associate each reference speech data with the corresponding reference topic. It is understood that the execution steps for determining the reference topic corresponding to the reference speech data are similar to or the same as the execution steps for determining the target topic corresponding to the target speech data, but the processing objects in the execution steps are different.
[0143] Optionally, the reference topic corresponding to the reference speech data can also be determined by calling the expert-side service.
[0144] In one or more embodiments of this specification, the entire speech processing stage avoids clustering large amounts of speech text. Target semantic tags based on the target speech data enable rapid matching of the reference speech set, optimizing the speech processing flow and reducing computational load. Furthermore, topic template rules configured for each reference topic, combining key information dimensions and / or mixed information dimensions on top of semantic tag dimensions, can recall a large number of similar scattered original sounds, improving the recall rate of similar sounds. Additionally, combining at least one of semantic tag dimensions, key information dimensions, and mixed information dimensions allows for accurate determination of the target topic, achieving accurate processing based on the target topic. Since real-time speech processing only requires determining the target topic of the target speech data, saving the reprocessing of data in the reference speech set, real-time speech processing can be achieved to provide real-time feedback on similar speech, improving the timeliness of speech processing.
[0145] Please see Figure 4 , Figure 4 This is a flowchart illustrating another embodiment of a speech processing method proposed in one or more embodiments of this specification. Specifically:
[0146] In one or more embodiments of this specification, at least one topic template rule can be determined based on at least one reference speech data in a reference speech set. The topic template rule may include, but is not limited to, at least one of semantic tag rules, key information rules, and mixed information rules; the process of determining the topic template rule can refer to the following steps.
[0147] S302: Determine reference semantic tags and reference topics based on at least one reference speech data in the reference speech set;
[0148] Optionally, the reference topics can be determined by calling expert services to analyze real feedback data (such as reference voice data) in the corresponding transaction scenario.
[0149] Optionally, the reference topic may be determined by performing speech clustering processing on each reference speech data in the reference speech set to obtain at least one clustered topic, and then determining the reference topic based on the at least one clustered topic.
[0150] Indicatively, voice clustering processing provides topic generation for the user's original voice in the corresponding transaction scenario; voice clustering processing can be carried out by adopting a corresponding voice clustering algorithm. For example, a pure clustering method can automatically cluster several reference voice data to obtain multiple clusters through semantic similarity calculation, where each cluster contains a corresponding topic.
[0151] For example, the reference voice data (such as service chat logs, service dialogue logs, etc.) from different sources in the corresponding transaction scenarios can be standardized. For example, each reference voice data can be converted from speech to text to obtain reference voice text, and a summary can be extracted from each reference voice text. Based on each summary, a standard text format can be obtained and stored. The standard format text can be vectorized using a neural network model (such as the word2vec model) to obtain text vectors. Based on a vector clustering algorithm (such as the Hdbscan density clustering algorithm), text clustering calculations can be performed to obtain multiple clusters, and the highest quality core of each cluster can be extracted as the cluster topic. The cluster topic can be directly used as the reference topic, or the cluster topics can be manually reviewed by experts. The topic names of several unreasonable cluster topics can be rewritten, or some semantically similar cluster topics can be merged to obtain the processed reference topic.
[0152] In one or more embodiments of this specification, the reference topic may be determined by one of the above methods, or a combination of the above methods, or may not be limited to the above methods for determining the reference topic.
[0153] In one feasible implementation, reference semantic tags can be determined based on at least one reference speech data in the reference speech set. Specifically, at least one key semantic segment can be determined for each reference speech data in the reference speech set; then, segment aggregation processing is performed on each key semantic segment to obtain at least one aggregated segment; and the reference semantic tag corresponding to each aggregated segment is determined.
[0154] In a demonstrative sense, the service platform can determine the reference semantic labels corresponding to the reference speech data based on a label classification model. First, the initial label classification model needs to be trained. This can be done by acquiring all or part of the reference speech data input by users in transaction scenarios from an existing speech database, and / or acquiring reference speech data recorded in actual language environments using recording equipment. Then, key semantic segments are extracted using the label classification model. These key semantic segments are then aggregated to obtain at least one aggregated segment. Finally, candidate semantic labels are extracted from the aggregated segment using the label classification model, and at least one candidate semantic label is output. The candidate semantic labels corresponding to the reference speech data can be adjusted based on expert-side services to obtain several standard semantic labels for the speech sample data. These standard semantic labels are, in this case, the reference semantic labels.
[0155] Optionally, adjusting the semantic labels of the candidate semantic labels corresponding to the reference speech data based on the expert service can be done by: rewriting the label names, merging and unifying the labels, attaching the label attributes, and confirming the label accuracy of several candidate semantic labels corresponding to the reference speech data, so as to obtain several standard semantic labels for the reference speech data. The standard semantic labels that conform to the specifications (i.e., reference semantic labels) are associated with the reference speech data and then saved (e.g., stored in a database).
[0156] Optionally, tag renaming can be understood as modifying the tag name of a candidate semantic tag identified by the tag classification model to a tag name that conforms to the specification. For example, the candidate tag "QR code not found" can be modified to "QR code not found" according to the tag specification to generate a standard semantic tag that conforms to the tag naming specification.
[0157] Optionally, tag merging and unification can be understood as: multiple similar candidate semantic tags can be uniformly modified into a single standard semantic tag. For example, three candidate semantic tags: "health code turns red", "health code turns green", and "health code turns yellow" can all be merged into a single standard semantic tag: "health code color change".
[0158] Optionally, tag attribute mounting can be understood as: adjusting at least one attribute element contained in the candidate semantic tag, such as removing or adding a tag attribute from the candidate semantic tag.
[0159] Optionally, the label accuracy identifies labels with errors among the candidate semantic labels. These erroneous labels and the original reference speech data are then labeled and subsequently input into the model for training.
[0160] Understandably, in this application, semantic recognition can be directly performed on the speech data to obtain semantic tags. That is, it is not necessary to convert the speech data into speech text and then perform semantic recognition on the speech text. Furthermore, the aforementioned candidate semantic tags, target semantic tags, etc., can all be speech type tags or text type tags.
[0161] Optionally, the label classification model can be trained based on reference speech data that has been labeled with standard semantic tags. During model training, the model parameters and model architecture can be adjusted using standard semantic tags to obtain a trained label classification model.
[0162] S304: Construct topic template rules for the reference topic based on the reference semantic tags;
[0163] In one feasible implementation, semantic tag rules corresponding to the reference topic can be constructed based on the reference semantic tags. It is understood that after determining the reference topic and reference semantic tags of each reference speech data, tag analysis processing can be performed on several reference speech data of the same type or the same reference topic to determine the common logical information of the tags among the several reference speech data, thereby constructing semantic tag rules corresponding to the reference topic.
[0164] For example, by analyzing the semantic tags of all reference speech data under a certain reference topic, we can determine common logical information such as the semantic tags that a reference topic must have, at least one of the semantic tags that a reference topic can have, the semantic tags that a reference topic should not have, and the number of semantic tags under a reference topic, and then construct semantic tag rules corresponding to the reference topic.
[0165] S306: Obtain key reference information for the reference topic, and construct topic template rules for the reference topic based on the reference semantic tags and the key reference information.
[0166] The reference key information may include key information elements (such as keywords, key phrases, key sentences), the order of key information, the number of key information (elements), the spacing between keywords / sentences, etc.
[0167] In one feasible implementation, obtaining reference key information for the reference topic may specifically involve: setting reference key information for the reference topic; and / or, performing key information identification on the reference speech data corresponding to the reference topic to obtain reference key information for the reference topic.
[0168] In illustrative terms, the expert service can be used to set key reference information for the reference topic, such as setting key information that is not present in a certain reference topic, the number of key information (elements), word order, key speech signal order, spacing between keywords / sentences, spacing between key speech signals, etc.
[0169] Indicatively, key information can be identified from several reference speech data corresponding to a reference topic, and candidate key information corresponding to each reference speech data under the reference topic can be determined sequentially. Information analysis and processing can be performed based on several candidate key information under the same or similar reference topics to determine the key logical common information among several reference speech data, thereby obtaining several reference key information for the reference topic.
[0170] For example, by identifying several sets of candidate key information for all reference speech data under a certain reference topic (each set of candidate key information corresponds to a reference speech data), and based on the common information logic among several sets of candidate key information under the same or similar reference topics, several reference key information for the reference topic can be obtained.
[0171] Indicatively, the topic template rules include at least one of semantic tag rules, key information rules, and mixed information rules;
[0172] Optionally, semantic tag rules corresponding to the reference topic can be constructed based on the reference semantic tags; see S304 for details, which will not be repeated here.
[0173] Optionally, key information rules corresponding to the reference topic can be constructed based on the aforementioned key information. For example, by identifying several sets of candidate key information for all reference speech data under a certain reference topic (each set of candidate key information corresponds to one set of reference speech data), and based on the commonalities in information logic among several sets of candidate key information under the same type or the same reference topic, such as the requirement that a reference topic must have at least one of several key information (elements), the corresponding key information that a reference topic should not have, the number of key information (elements) under a reference topic, the word order under a reference topic, the order of key speech signals under a reference topic, the spacing rules between keywords / sentences under a reference topic, and the spacing between key speech signals under a reference topic, etc., key information rules corresponding to the reference topic can be constructed based on the commonalities in information logic corresponding to several key information. Typically, key information rules are represented in the form of logical expressions.
[0174] Optionally, a hybrid information rule corresponding to the reference topic can be constructed based on the reference semantic tag and the reference key information.
[0175] The hybrid information rule can be understood as the logical rule of information elements that correspond to both semantic tags and key information. Through the hybrid information rule, at least the topic characteristics between semantic tags and key information under the reference topic can be reflected, such as semantic tags and key information (such as keywords and key sentences) that must exist under a certain reference topic, semantic tags that must exist under a certain reference topic and key information (such as keywords and key sentences) that do not exist.
[0176] Indicatively, after identifying several sets of candidate key information (each set of candidate key information corresponds to one set of candidate key information) and reference semantic tags for all reference speech data under a certain reference topic, based on the commonalities of the mixed information logic between several sets of "candidate key information and reference semantic tags" (each set of "candidate key information and reference semantic tags" corresponds to one set of candidate key information and reference semantic tags) under the same or similar reference topic, the commonalities of the mixed information logic include the semantic tags that must exist under the reference topic and the key information that does not exist. For example, it can be: the semantic tags and key information that must exist simultaneously under the reference topic (i.e., the logical AND relationship between semantic tags and key information), or the semantic tags or key information that must exist under the reference topic (i.e., the logical OR relationship between semantic tags and key information), etc., to construct the mixed information rules corresponding to the reference topic. Usually, the mixed information rules are represented in the form of logical expressions.
[0177] Understandably, the topic template rules constructed for the reference topic include at least one of the following: semantic tag rules, key information rules, and mixed information rules.
[0178] S308: Perform semantic recognition processing on the target speech data to obtain at least one target semantic label corresponding to the target speech data and obtain the transaction scenario corresponding to the target speech data;
[0179] The transaction scenario can be understood as the transaction scenario characteristics corresponding to the current target voice data. For example, the target voice data is usually input by the user in a certain transaction scenario. By collecting the target voice data input by the user in a certain transaction scenario, the service platform can obtain the target voice data. In other words, the target voice data is associated with the transaction scenario, which can be of the following types: function name (such as the name of an application function), platform name (such as a comprehensive transaction platform), application scenario name (such as a shopping scenario ID, a communication scenario name), etc.
[0180] Understandably, the transaction scenario to which the target voice data belongs can be determined during the collection process; for example, if the target voice data is user feedback voice collected in a shopping platform scenario, then the transaction scenario corresponding to the target voice data is the shopping platform scenario.
[0181] In one or more embodiments of this specification, for each reference speech data in the reference speech set, its reference transaction scenario can be determined, and the reference transaction scenario corresponding to each reference speech data can be associated.
[0182] S310: Based on the at least one target semantic label and the transaction scenario, perform speech matching processing on the reference speech set to obtain similar speech data corresponding to the target speech data.
[0183] In one or more embodiments of this specification, the target speech data can be processed by topic matching based on at least one target semantic tag using at least one topic template rule to obtain the target topic corresponding to the target speech data; then, based on the target topic and the transaction scenario, a reference speech set can be processed by speech matching to obtain similar speech data corresponding to the target speech data. By accurately determining the target topic and combining it with the transaction scenario to determine similar speech, the accuracy of similar speech can be improved, speech interference in non-transaction scenarios can be avoided, and the recall accuracy of similar speech can be improved. Furthermore, since it is not necessary to re-recall and recalculate the reference speech data and target speech data in each round of speech processing, the computational load of real-time processing is greatly reduced, the computational resources in the similar speech processing process are saved, the processing efficiency is improved, and real-time feedback of similar speech corresponding to the target speech data can be achieved.
[0184] Understandably, after determining the target topic and transaction scenario corresponding to the target speech data, speech data belonging to the same target topic and transaction scenario can be searched in the reference speech set, and then these speech data can be used as the first similar speech data of the target speech data.
[0185] In illustrative terms, there can be multiple target topics. The target topics corresponding to the target voice data can be assembled into an "OR" or "AND" logical relationship. Furthermore, search queries can be constructed based on scenario topics to achieve real-time searching of similar voices, and even real-time calculation of voice volume or real-time acquisition and display of voice volume trends. For example, if target voice data 1 identifies two target topics, topic1 and topic2, then the SQL (i.e., search query) for similar voices in target voice data 1 can be: `select * from voice where topic in (topic1, topic2) and product_id = 'xxx'`, where "product_id = 'xxx'" represents a transaction scenario.
[0186] In one or more embodiments of this specification, the entire speech processing stage avoids clustering large amounts of speech text. Target semantic tags based on the target speech data enable rapid matching of the reference speech set, optimizing the speech processing flow and reducing computational load. Furthermore, topic template rules configured for each reference topic, combining key information dimensions and / or mixed information dimensions on top of semantic tag dimensions, can recall a large number of similar scattered original audio points, improving the recall rate of similar sounds. Additionally, combining at least one of semantic tag dimensions, key information dimensions, and mixed information dimensions allows for accurate determination of the target topic, achieving accurate processing based on the target topic. Since real-time speech processing processes such as tag determination and template matching only require determining the target topic based on a single piece of target speech data, reprocessing of data in the reference speech set is saved. Real-time speech processing can be achieved to provide real-time feedback on similar speech, improving the timeliness of speech processing.
[0187] The following will combine Figure 5 This specification provides a detailed description of the speech processing apparatus provided in one or more embodiments. It should be noted that... Figure 5 The speech processing apparatus shown is used to execute the present application. Figures 2-3 The methods of the embodiments shown are illustrated only in the parts relevant to this specification for ease of explanation. For specific technical details not disclosed, please refer to one or more embodiments of this specification.
[0188] Please see Figure 5 This diagram illustrates the structure of a voice processing device according to one or more embodiments of this specification. The voice processing device 1 can be implemented as all or part of a user terminal through software, hardware, or a combination of both. According to some embodiments, the voice processing device 1 includes a tag determination module 11 and a voice matching module 12, specifically used for:
[0189] The label determination module 11 is used to perform semantic recognition processing on the target speech data to obtain at least one target semantic label corresponding to the target speech data;
[0190] The speech matching module 12 is used to perform speech matching processing on the reference speech set based on the at least one target semantic label to obtain similar speech data corresponding to the target speech data.
[0191] Optional, such as Figure 6 As shown, the voice matching module 12 includes:
[0192] The topic determination unit 121 is used to perform topic matching processing on the target speech data based on the at least one target semantic label and at least one topic template rule to obtain the target topic corresponding to the target speech data.
[0193] The speech matching unit 122 is used to perform speech matching processing on the reference speech set based on the target topic to obtain similar speech data corresponding to the target speech data.
[0194] Optional, such as Figure 7 As shown, the topic determination unit 121 includes:
[0195] Rule acquisition subunit 1211 is used to acquire at least one semantic tag rule corresponding to a topic template rule;
[0196] The rule matching subunit 1212 is used to determine the target label rule that matches the at least one target semantic label from each of the semantic label rules, and to obtain the first topic corresponding to the target label rule;
[0197] The topic determination subunit 1213 is used to determine the target topic corresponding to the target speech data based on the first topic.
[0198] Optionally, the topic determination unit 121 is specifically used for:
[0199] Obtain at least one tag logic rule corresponding to a topic template rule;
[0200] Detect whether the at least one target semantic tag matches each of the tag logical rules to obtain the tag matching result;
[0201] Based on the tag matching results, the target tag rule for matching the at least one target semantic tag is determined.
[0202] Optionally, the rule acquisition subunit 1211 is specifically used for:
[0203] Obtain at least one key information rule and / or mixed information rule corresponding to a topic template rule, wherein the mixed information rule is a template rule composed of semantic tags and key information;
[0204] The topic-determining subunit 1213 is specifically used for:
[0205] Determine the target key information rule matching the target speech data from each of the key information rules, and obtain the second topic corresponding to the target key information rule; determine the target topic corresponding to the target speech data based on at least one of the first topic and the second topic; or...
[0206] Determine the target mixed information rule matching the target speech data from each of the mixed information rules, and obtain the third topic corresponding to the target mixed information rule; determine the target topic corresponding to the target speech data based on at least one of the first topic and the third topic; or...
[0207] Determine the target key information rule matching the target speech data from each of the key information rules, and obtain the second topic corresponding to the target key information rule; determine the target mixed information rule matching the target speech data from each of the mixed information rules, and obtain the third topic corresponding to the target mixed information rule; determine the target topic corresponding to the target speech data based on at least one of the first topic, the second topic, and the third topic.
[0208] Optionally, the topic determination subunit 1213 is specifically used for:
[0209] Based on the key information rules, the target speech data is processed for information detection to obtain information detection results.
[0210] Based on the information detection results, the target key information rules for matching the target speech data are determined.
[0211] Optionally, the topic determination subunit 1213 is specifically used for:
[0212] Based on the information logic rules corresponding to each of the aforementioned key information rules, the target speech data is subjected to key information logic detection processing to obtain logic detection results; and / or,
[0213] Based on the information order rules corresponding to each of the aforementioned key information rules, the target speech data is subjected to key information order detection processing to obtain order detection results; and / or,
[0214] Based on the information spacing rules corresponding to each of the key information rules, the target speech data is subjected to key information spacing detection processing to obtain the spacing detection result.
[0215] Optionally, the rule matching subunit 1212 is specifically used for:
[0216] The hybrid information rule is a logical rule for information elements that jointly correspond to semantic tags and key information;
[0217] Based on the logical rules of each information element, the target speech data is subjected to element logic detection processing to obtain the element logic detection result;
[0218] The target mixed information rule for matching the target speech data is determined based on the element logic detection results.
[0219] Optional, such as Figure 8 As shown, the device 1 further includes:
[0220] Rule determination module 13 is used to determine at least one topic template rule based on at least one reference speech data in the reference speech set.
[0221] Optionally, the rule determination module 13 is specifically used for:
[0222] Based on at least one reference speech data from the reference speech set, a reference semantic label and a reference topic are determined; based on the reference semantic label, a topic template rule is constructed for the reference topic; or,
[0223] Based on at least one reference speech data in the reference speech set, a reference semantic label and a reference topic are determined; reference key information for the reference topic is obtained, and a topic template rule for the reference topic is constructed based on the reference semantic label and the reference key information.
[0224] Optionally, the topic template rules include at least one of semantic tagging rules, key information rules, and mixed information rules; such as Figure 9 As shown, the rule determination module 13 includes:
[0225] The semantic tag rule determination unit 131 is used to construct semantic tag rules corresponding to the reference topic based on the reference semantic tags;
[0226] The key information rule determination unit 132 is used to construct key information rules corresponding to the reference topic based on the reference key information;
[0227] The mixed information rule determination unit 133 is used to construct the mixed information rule corresponding to the reference topic based on the reference semantic tag and the reference key information.
[0228] Optionally, according to the method of claim 10, the rule determination module 13 is specifically used for:
[0229] Set key reference information for the reference topic; and / or,
[0230] Key information is identified from the reference speech data corresponding to the reference topic to obtain key reference information for the reference topic.
[0231] Optionally, the rule determination module 13 is specifically used for:
[0232] Determine at least one key semantic segment corresponding to each reference speech data in the reference speech set; perform segment aggregation processing on each key semantic segment to obtain at least one aggregated segment; determine the reference semantic tag corresponding to each aggregated segment;
[0233] Perform speech clustering processing on each reference speech data in the reference speech set to obtain at least one cluster topic, determine a reference topic based on the at least one cluster topic, and / or call an expert service to set the reference topic.
[0234] Optionally, the voice matching module 12 is specifically used for:
[0235] Obtain the reference topic corresponding to at least one reference speech data in the reference speech set;
[0236] Based on the reference topics corresponding to each of the reference speech data, similar speech data corresponding to the target topic are determined from the reference speech set.
[0237] Optionally, the voice matching module 12 is specifically used for:
[0238] Based on the topic template rules, at least one reference speech data in the reference speech set is subjected to topic matching processing to obtain the reference topic corresponding to the reference speech data.
[0239] Optionally, the label determination module 11 is specifically used for:
[0240] The target speech data is input into the label classification model, and at least one target semantic label is output for the target speech data.
[0241] Optionally, the tag determination module 11 is specifically used to: obtain the transaction scenario corresponding to the target voice data;
[0242] The speech matching module is specifically used to: perform speech matching processing on the reference speech set based on the at least one target semantic label and the transaction scenario to obtain similar speech data corresponding to the target speech data.
[0243] It should be noted that the speech processing device provided in the above embodiments is only illustrated by the division of the above functional modules when executing the speech processing method. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the speech processing device and the speech processing method embodiments provided in the above embodiments belong to the same concept, and the implementation process can be found in the method embodiments, which will not be repeated here.
[0244] The serial numbers in this specification are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0245] In one or more embodiments of this specification, the entire speech processing stage avoids clustering large amounts of speech text. Target semantic tags based on the target speech data enable rapid matching of the reference speech set, optimizing the speech processing flow and reducing computational load. Furthermore, topic template rules configured for each reference topic, combining key information dimensions and / or mixed information dimensions on top of semantic tag dimensions, can recall a large number of similar scattered original sounds, improving the recall rate of similar sounds. Additionally, combining at least one of semantic tag dimensions, key information dimensions, and mixed information dimensions allows for accurate determination of the target topic, achieving accurate processing based on the target topic. Since real-time speech processing only requires determining the target topic of the target speech data, saving the reprocessing of data in the reference speech set, real-time speech processing can be achieved to provide real-time feedback on similar speech, improving the timeliness of speech processing.
[0246] This specification also provides a computer storage medium that can store multiple instructions adapted to be loaded and executed by a processor as described above. Figures 2-4 The speech processing method described in the illustrated embodiment can be found in the following document for a detailed execution process. Figures 2-4 The specific details of the illustrated embodiments will not be elaborated here.
[0247] This specification also provides a computer program product that stores at least one instruction, which is loaded and executed by the processor as described above. Figures 2-4 The speech processing method described in the illustrated embodiment can be found in the following document for a detailed execution process. Figures 2-4 The specific details of the illustrated embodiments will not be elaborated here.
[0248] Please see Figure 10 This specification provides a schematic diagram of the structure of an electronic device for one or more embodiments. Figure 10 As shown, the electronic device 1000 may include: at least one processor 1001, at least one network interface 1004, a user interface 1003, a memory 1005, and at least one communication bus 1002.
[0249] The communication bus 1002 is used to realize the connection and communication between these components.
[0250] The user interface 1003 may include a display screen and a camera. Optionally, the user interface 1003 may also include a standard wired interface and a wireless interface.
[0251] The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).
[0252] The processor 1001 may include one or more processing cores. The processor 1001 connects to various parts within the server 1000 using various interfaces and lines, and performs various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 1005, and by calling data stored in the memory 1005. Optionally, the processor 1001 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 1001 may integrate one or a combination of several of the following: a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), and a modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 1001 and may be implemented as a separate chip.
[0253] The memory 1005 may include random access memory (RAM) or read-only memory. Optionally, the memory 1005 may include a non-transitory computer-readable storage medium. The memory 1005 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 1005 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 1005 may also be at least one storage device located remotely from the aforementioned processor 1001. Figure 10 As shown, the memory 1005, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and application programs.
[0254] exist Figure 10In the illustrated electronic device 1000, the user interface 1003 is mainly used to provide an input interface for the user and to obtain user input data; while the processor 1001 can be used to call the application program stored in the memory 1005 and specifically perform the following operations:
[0255] Semantic recognition processing is performed on the target speech data to obtain at least one target semantic label corresponding to the target speech data;
[0256] Based on the at least one target semantic label, speech matching processing is performed on the reference speech set to obtain similar speech data corresponding to the target speech data.
[0257] In one embodiment, when the processor 1001 performs the speech matching process based on the at least one target semantic label to obtain similar speech data corresponding to the target speech data, it specifically performs the following steps:
[0258] Based on the at least one target semantic tag, at least one topic template rule is used to perform topic matching processing on the target speech data to obtain the target topic corresponding to the target speech data;
[0259] Based on the target topic, speech matching processing is performed on the reference speech set to obtain similar speech data corresponding to the target speech data.
[0260] In one embodiment, when the processor 1001 performs topic matching processing on the target speech data based on the at least one target semantic tag and using at least one topic template rule to obtain the target topic corresponding to the target speech data, it specifically performs the following steps: obtaining the semantic tag rule corresponding to at least one topic template rule;
[0261] From the semantic tag rules, determine the target tag rule that matches the at least one target semantic tag, and obtain the first topic corresponding to the target tag rule;
[0262] The target topic corresponding to the target speech data is determined based on the first topic.
[0263] In one embodiment, when the processor 1001 executes the semantic tag rule as a tag logic rule, and obtains the semantic tag rule corresponding to at least one topic template rule, and determines the target tag rule matching the at least one target semantic tag from each of the semantic tag rules, it specifically performs the following steps: obtaining the tag logic rule corresponding to at least one topic template rule;
[0264] Detect whether the at least one target semantic tag matches each of the tag logical rules to obtain the tag matching result;
[0265] Based on the tag matching results, the target tag rule for matching the at least one target semantic tag is determined.
[0266] In one embodiment, after executing the semantic tag rule corresponding to the rule for obtaining at least one topic template, the processor 1001 performs the following steps:
[0267] Obtain at least one key information rule and / or mixed information rule corresponding to a topic template rule, wherein the mixed information rule is a template rule composed of semantic tags and key information;
[0268] Determining the target topic corresponding to the target speech data based on the first topic includes:
[0269] Determine the target key information rule matching the target speech data from each of the key information rules, and obtain the second topic corresponding to the target key information rule; determine the target topic corresponding to the target speech data based on at least one of the first topic and the second topic; or...
[0270] Determine the target mixed information rule matching the target speech data from each of the mixed information rules, and obtain the third topic corresponding to the target mixed information rule; determine the target topic corresponding to the target speech data based on at least one of the first topic and the third topic; or...
[0271] Determine the target key information rule matching the target speech data from each of the key information rules, and obtain the second topic corresponding to the target key information rule; determine the target mixed information rule matching the target speech data from each of the mixed information rules, and obtain the third topic corresponding to the target mixed information rule; determine the target topic corresponding to the target speech data based on at least one of the first topic, the second topic, and the third topic.
[0272] In one embodiment, when the processor 1001 executes the target key information rule for determining the target speech data matching from each of the key information rules, it specifically performs the following steps:
[0273] Based on the key information rules, the target speech data is processed for information detection to obtain information detection results.
[0274] Based on the information detection results, the target key information rules for matching the target speech data are determined.
[0275] In one embodiment, when the processor 1001 performs key information detection processing on the target speech data based on each of the key information rules to obtain information detection results, it specifically performs the following steps:
[0276] Based on the information logic rules corresponding to each of the aforementioned key information rules, the target speech data is subjected to key information logic detection processing to obtain logic detection results; and / or,
[0277] Based on the information order rules corresponding to each of the aforementioned key information rules, the target speech data is subjected to key information order detection processing to obtain order detection results; and / or,
[0278] Based on the information spacing rules corresponding to each of the key information rules, the target speech data is subjected to key information spacing detection processing to obtain the spacing detection result.
[0279] In one embodiment, when the processor 1001 executes the target mixed information rule for determining the target speech data matching from each of the mixed information rules, it specifically performs the following steps:
[0280] The hybrid information rule is a logical rule for information elements that jointly correspond to semantic tags and key information;
[0281] Based on the logical rules of each information element, the target speech data is subjected to element logic detection processing to obtain the element logic detection result;
[0282] The target mixed information rule for matching the target speech data is determined based on the element logic detection results.
[0283] In one embodiment, before executing the determination of at least one target semantic label corresponding to the target speech data, the processor 1001 further performs the following steps: determining at least one topic template rule based on at least one reference speech data in the reference speech set.
[0284] In one embodiment, when the processor 1001 executes the process of determining at least one topic template rule based on at least one reference speech data in the reference speech set, it specifically performs the following steps:
[0285] Based on at least one reference speech data from the reference speech set, a reference semantic label and a reference topic are determined; based on the reference semantic label, a topic template rule is constructed for the reference topic; or,
[0286] Based on at least one reference speech data in the reference speech set, a reference semantic label and a reference topic are determined; reference key information for the reference topic is obtained, and a topic template rule for the reference topic is constructed based on the reference semantic label and the reference key information.
[0287] In one embodiment, when the processor 1001 executes the topic template rule for the reference topic based on the reference semantic tag and the reference key information, it specifically performs the following steps:
[0288] The topic template rules include at least one of semantic tag rules, key information rules, and mixed information rules;
[0289] Based on the reference semantic tags, construct semantic tag rules corresponding to the reference topic;
[0290] Based on the aforementioned key reference information, construct key information rules corresponding to the reference topic;
[0291] Based on the reference semantic tags and the reference key information, construct the hybrid information rules corresponding to the reference topic.
[0292] In one embodiment, when the processor 1001 performs the step of acquiring key reference information for the reference topic, it specifically executes the following steps:
[0293] Set key reference information for the reference topic; and / or,
[0294] Key information is identified from the reference speech data corresponding to the reference topic to obtain key reference information for the reference topic.
[0295] In one embodiment, when the processor 1001 performs the steps of determining reference semantic labels and determining reference topics based on at least one reference speech data in the reference speech set, it specifically executes the following steps:
[0296] Determine at least one key semantic segment corresponding to each reference speech data in the reference speech set; perform segment aggregation processing on each key semantic segment to obtain at least one aggregated segment; determine the reference semantic tag corresponding to each aggregated segment;
[0297] Perform speech clustering processing on each reference speech data in the reference speech set to obtain at least one cluster topic, determine a reference topic based on the at least one cluster topic, and / or call an expert service to set the reference topic.
[0298] In one embodiment, when the processor 1001 performs speech matching processing on the reference speech set based on the target topic to obtain similar speech data corresponding to the target speech data, it specifically performs the following steps: obtaining a reference topic corresponding to at least one reference speech data in the reference speech set; determining similar speech data corresponding to the target topic from the reference speech set based on the reference topics corresponding to each of the reference speech data.
[0299] In one embodiment, before executing the step of obtaining the reference topic corresponding to at least one reference speech data in the reference speech set, the processor 1001 further performs the following steps: performing topic matching processing on at least one reference speech data in the reference speech set based on each of the topic template rules to obtain the reference topic corresponding to the reference speech data.
[0300] In one embodiment, when the processor 1001 executes the process of determining at least one target semantic label corresponding to the target speech data, it specifically performs the following steps: inputting the target speech data into a label classification model and outputting at least one target semantic label for the target speech data.
[0301] In one embodiment, when the processor 1001 performs the speech matching processing on the reference speech set based on the at least one target semantic label to obtain similar speech data corresponding to the target speech data, it specifically performs the following steps: obtaining the transaction scenario corresponding to the target speech data; performing speech matching processing on the reference speech set based on the at least one target semantic label and the transaction scenario to obtain similar speech data corresponding to the target speech data.
[0302] In one embodiment, after obtaining the similar speech data corresponding to the target speech data, the processor 1001 further performs the following step: determining the similar volume corresponding to the target speech data based on the similar speech data.
[0303] In one or more embodiments of this specification, the entire speech processing stage avoids clustering large amounts of speech text. Target semantic tags based on the target speech data enable rapid matching of the reference speech set, optimizing the speech processing flow and reducing computational load. Furthermore, topic template rules configured for each reference topic, combining key information dimensions and / or mixed information dimensions on top of semantic tag dimensions, can recall a large number of similar scattered original sounds, improving the recall rate of similar sounds. Additionally, combining at least one of semantic tag dimensions, key information dimensions, and mixed information dimensions allows for accurate determination of the target topic, achieving accurate processing based on the target topic. Since real-time speech processing only requires determining the target topic of the target speech data, saving the reprocessing of data in the reference speech set, real-time speech processing can be achieved to provide real-time feedback on similar speech, improving the timeliness of speech processing.
[0304] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory, or random access memory, etc.
[0305] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.
Claims
1. A speech processing method, the method comprising: Semantic recognition processing is performed on the target speech data to obtain at least one target semantic label corresponding to the target speech data; Based on the at least one target semantic label, speech matching processing is performed on the reference speech set to obtain similar speech data corresponding to the target speech data; The step of performing speech matching processing on the reference speech set based on the at least one target semantic label to obtain similar speech data corresponding to the target speech data includes: Based on the at least one target semantic tag, at least one topic template rule is used to perform topic matching processing on the target speech data to obtain the target topic corresponding to the target speech data; Based on the target topic, speech matching processing is performed on the reference speech set to obtain similar speech data corresponding to the target speech data.
2. The method according to claim 1, wherein the step of performing topic matching processing on the target speech data based on the at least one target semantic tag using at least one topic template rule to obtain the target topic corresponding to the target speech data includes: Obtain at least one semantic tag rule corresponding to a topic template rule; From the semantic tag rules, determine the target tag rule that matches the at least one target semantic tag, and obtain the first topic corresponding to the target tag rule; The target topic corresponding to the target speech data is determined based on the first topic.
3. The method according to claim 2, wherein the semantic tagging rule is a tag logic rule, and the step of obtaining the semantic tagging rule corresponding to at least one topic template rule, and determining the target tagging rule matching the at least one target semantic tag from the semantic tagging rules, includes: Obtain at least one tag logic rule corresponding to a topic template rule; Detect whether the at least one target semantic tag matches each of the tag logical rules to obtain the tag matching result; Based on the tag matching results, the target tag rule for matching the at least one target semantic tag is determined.
4. The method according to claim 2, after obtaining the semantic tag rule corresponding to at least one topic template rule, further comprising: Obtain at least one key information rule and / or mixed information rule corresponding to a topic template rule, wherein the mixed information rule is a template rule composed of semantic tags and key information; Determining the target topic corresponding to the target speech data based on the first topic includes: Determine the target key information rule matching the target speech data from each of the key information rules, and obtain the second topic corresponding to the target key information rule; determine the target topic corresponding to the target speech data based on at least one of the first topic and the second topic; or... Determine the target mixed information rule matching the target speech data from each of the mixed information rules, and obtain the third topic corresponding to the target mixed information rule; determine the target topic corresponding to the target speech data based on at least one of the first topic and the third topic; or... Determine the target key information rule matching the target speech data from each of the key information rules, and obtain the second topic corresponding to the target key information rule; determine the target mixed information rule matching the target speech data from each of the mixed information rules, and obtain the third topic corresponding to the target mixed information rule; determine the target topic corresponding to the target speech data based on at least one of the first topic, the second topic, and the third topic.
5. The method according to claim 4, wherein determining the target key information rule for matching the target speech data from each of the key information rules comprises: Based on the key information rules, the target speech data is processed for information detection to obtain information detection results. Based on the information detection results, the target key information rules for matching the target speech data are determined.
6. The method according to claim 5, wherein the step of performing key information detection processing on the target speech data based on each of the key information rules to obtain information detection results includes: Based on the information logic rules corresponding to each of the key information rules, the target speech data is subjected to key information logic detection processing to obtain the logic detection result. And / or, Based on the information order rules corresponding to each of the aforementioned key information rules, the target speech data is subjected to key information order detection processing to obtain order detection results; and / or, Based on the information spacing rules corresponding to each of the key information rules, the target speech data is subjected to key information spacing detection processing to obtain the spacing detection result.
7. The method according to claim 4, wherein determining the target mixed information rule for matching the target speech data from each of the mixed information rules comprises: The hybrid information rule is a logical rule for information elements that jointly correspond to semantic tags and key information; Based on the logical rules of each information element, the target speech data is subjected to element logic detection processing to obtain the element logic detection result; The target mixed information rule for matching the target speech data is determined based on the element logic detection results.
8. The method according to claim 1, before determining at least one target semantic tag corresponding to the target speech data, further comprising: At least one topic template rule is determined based on at least one reference speech data in the reference speech set.
9. The method according to claim 8, wherein determining at least one topic template rule based on at least one reference speech data in the reference speech set comprises: Based on at least one reference speech data in the reference speech set, a reference semantic label and a reference topic are determined. Construct topic template rules for the reference topic based on the reference semantic tags; or, Based on at least one reference speech data in the reference speech set, a reference semantic label and a reference topic are determined. Obtain key reference information for the reference topic, and construct topic template rules for the reference topic based on the reference semantic tags and the key reference information.
10. The method according to claim 9, wherein constructing topic template rules for the reference topic based on the reference semantic tags and the reference key information includes: The topic template rules include at least one of semantic tag rules, key information rules, and mixed information rules; Based on the reference semantic tags, construct semantic tag rules corresponding to the reference topic; Based on the aforementioned key reference information, construct key information rules corresponding to the reference topic; Based on the reference semantic tags and the reference key information, construct the hybrid information rules corresponding to the reference topic.
11. The method according to claim 9, wherein obtaining key reference information for the reference topic includes: Set key reference information for the aforementioned topic; And / or, Key information is identified from the reference speech data corresponding to the reference topic to obtain key reference information for the reference topic.
12. The method according to claim 9, wherein determining the reference semantic label and determining the reference topic based on at least one reference speech data in the reference speech set comprises: Identify at least one key semantic segment corresponding to each reference speech data in the reference speech set; Perform fragment aggregation processing on each of the key semantic fragments to obtain at least one aggregated fragment; determine the reference semantic tag corresponding to each of the aggregated fragments; Speech clustering processing is performed on each reference speech data in the reference speech set to obtain at least one cluster topic, and a reference topic is determined based on the at least one cluster topic; And / or, invoke expert services to set reference topics.
13. The method according to claim 1, wherein the step of performing speech matching processing on the reference speech set based on the target topic to obtain similar speech data corresponding to the target speech data includes: Obtain the reference topic corresponding to at least one reference speech data in the reference speech set; Based on the reference topics corresponding to each of the reference speech data, similar speech data corresponding to the target topic are determined from the reference speech set.
14. The method according to claim 13, further comprising, before obtaining the reference topic corresponding to at least one reference speech data in the reference speech set: Based on the topic template rules, at least one reference speech data in the reference speech set is subjected to topic matching processing to obtain the reference topic corresponding to the reference speech data.
15. The method according to claim 1, wherein performing semantic recognition processing on the target speech data to obtain at least one target semantic tag corresponding to the target speech data includes: The target speech data is input into the label classification model, and at least one target semantic label is output for the target speech data.
16. The method according to claim 1, wherein the step of performing speech matching processing on the reference speech set based on the at least one target semantic label to obtain similar speech data corresponding to the target speech data includes: Obtain the transaction scenario corresponding to the target voice data; Based on the at least one target semantic label and the transaction scenario, speech matching processing is performed on the reference speech set to obtain similar speech data corresponding to the target speech data.
17. The method according to any one of claims 1-16, wherein after obtaining the similar speech data corresponding to the target speech data, it further comprises: Based on the similar speech data, the similar volume of the target speech data is determined.
18. A voice processing apparatus, the apparatus comprising: The label determination module is used to perform semantic recognition processing on the target speech data to obtain at least one target semantic label corresponding to the target speech data. The speech matching module is used to perform speech matching processing on the reference speech set based on the at least one target semantic label to obtain similar speech data corresponding to the target speech data; The step of performing speech matching processing on the reference speech set based on the at least one target semantic label to obtain similar speech data corresponding to the target speech data includes: Based on the at least one target semantic tag, at least one topic template rule is used to perform topic matching processing on the target speech data to obtain the target topic corresponding to the target speech data; Based on the target topic, speech matching processing is performed on the reference speech set to obtain similar speech data corresponding to the target speech data.
19. A computer storage medium storing a plurality of instructions, the instructions being loaded by a processor and executed as the method steps of any one of claims 1 to 17.
20. An electronic device, comprising: A processor and a memory; wherein the memory stores a computer program, which is loaded by the processor and executes the method steps as claimed in any one of claims 1 to 17.
21. A computer program product storing at least one instruction, the at least one instruction being loaded by a processor and executing the method steps as claimed in any one of claims 1 to 17.
Citation Information
Patent Citations
Intelligent voice interaction method and system and storage medium
CN112489645A