Voice navigation method and apparatus, and computer device and storage medium

By configuring business lines and semantic recognition strategies in the voice navigation system and combining semantic slots, the problem of poor universality of the existing IVR system is solved, and flexible adaptation and efficient voice navigation are achieved for a variety of business scenarios.

WO2025092791A1PCT designated stage expired Publication Date: 2025-05-08MASHANG CONSUMER FINANCE CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/128405
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-30
Filing Date
2024-10-30
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

The existing IVR systems are poorly versatile in voice navigation and cannot be applied to multiple business scenarios and specific business needs.

Method used

By configuring different business lines and semantic recognition strategies in the voice navigation system, combined with pre-configured semantic slots, semantic recognition and matching of user voice inputs can be achieved, thereby improving the universality of voice navigation.

Benefits of technology

It realizes the flexible adaptation of the voice navigation system to different business scenarios, and improves the universality of the system and the satisfaction of business needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024128405_08052025_PF_FP_ABST
    Figure CN2024128405_08052025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to a voice navigation method and apparatus, and a computer device and a storage medium. The method comprises: performing text conversion on acquired first audio, so as to obtain audio text, wherein the first audio is input by means of an incoming access number; determining a service line associated with the access number, and determining a semantic recognition strategy and reference semantics which are configured for the service line; performing semantic recognition processing on the audio text on the basis of the semantic recognition strategy, and obtaining N pieces of slot semantics on the basis of the recognized semantics and N pre-configured semantic slots, wherein N is a positive integer greater than or equal to 2; and determining a semantic matching result on the basis of the N pieces of slot semantics and the reference semantics, and performing voice broadcasting on second audio which matches the semantic matching result. By using the present application, the universality of voice navigation can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Voice navigation method, device, computer equipment and storage medium

[0001] This disclosure claims priority to the Chinese patent application filed with the China Patent Office on October 30, 2023, with application number 202311424157.2, and invention name “Voice navigation method, device, computer equipment and storage medium”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of speech recognition technology, and in particular to a speech navigation method, apparatus, computer equipment, and storage medium. Background Art

[0003] With the continuous advancement of communication technology, more and more customers are dialing a phone number to access an interactive voice response (IVR) system. The IVR system then plays voice guidance to guide users through relevant business operations. Therefore, how to accurately determine the voice guidance and improve its versatility is a hot topic in the field of speech recognition.

[0004] Summary of the Invention

[0005] Based on this, it is necessary to provide a voice navigation method, device, computer equipment and storage medium to address the above technical problems, which can improve the versatility of voice navigation.

[0006] In a first aspect, the present application provides a voice navigation method, comprising: performing text conversion on an acquired first audio to obtain an audio text; wherein the first audio is input through an incoming call access number; determining a business line associated with the access number, and determining a semantic recognition strategy and reference semantics configured for the business line; performing semantic recognition processing on the audio text according to the semantic recognition strategy, and obtaining N slot semantics based on the recognized semantics and pre-configured N semantic slots; wherein N is a positive integer greater than or equal to 2; determining a semantic matching result based on the N slot semantics and the reference semantics, and voice broadcasting a second audio that matches the semantic matching result.

[0007] In the second aspect, the present application provides a voice navigation device, comprising: a conversion module for converting the acquired first audio into text to obtain an audio text; wherein the first audio is input through the incoming access number; an identification module for determining the business line associated with the access number, and determining the semantic recognition strategy and reference semantics configured for the business line; a matching module for performing semantic recognition processing on the audio text according to the semantic recognition strategy, and obtaining N slot semantics based on the recognized semantics and pre-configured N semantic slots; wherein N is a positive integer greater than or equal to 2; a navigation module for determining the semantic matching result based on the N slot semantics and the reference semantics, and voice broadcasting the second audio that matches the semantic matching result.

[0008] In a third aspect, the present application provides a computer device, which includes a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method when executing the computer program.

[0009] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the steps in the above method when executed by a processor.

[0010] In a fifth aspect, the present application provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps in the above method. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments described in this specification. Those skilled in the art can also derive other drawings based on these drawings without inventive effort.

[0012] FIG1 is a diagram illustrating an application environment of a voice navigation method provided by an embodiment of the present application;

[0013] FIG2 is a flow chart of a voice navigation method provided in an embodiment of the present application;

[0014] FIG3 is a schematic diagram of the overall flow of a voice navigation method provided by an embodiment of the present application;

[0015] FIG4 is a schematic diagram of a service configuration interface provided in an embodiment of the present application;

[0016] FIG5 is a schematic diagram of an interface for publishing a voice navigation version provided in an embodiment of the present application;

[0017] FIG6 is a schematic diagram of an interface for configuring a semantic recognition strategy and a semantic recognition model according to an embodiment of the present application;

[0018] FIG7 is a schematic diagram of a semantic configuration page provided in an embodiment of the present application;

[0019] FIG8 is a schematic diagram of an interactive node configuration page provided in an embodiment of the present application;

[0020] FIG9 is a schematic diagram of a service routing configuration page provided in an embodiment of the present application;

[0021] FIG10 is a schematic diagram of a road sign voice configuration page provided in an embodiment of the present application;

[0022] FIG11 is a schematic diagram of an interface for configuring a service line according to an embodiment of the present application;

[0023] FIG12 is a schematic diagram of a model configuration page provided in an embodiment of the present application;

[0024] FIG13 is a schematic diagram of a flow chart of a single identification strategy provided in an embodiment of the present application;

[0025] FIG14 is a flow chart of a combined identification strategy provided in an embodiment of the present application;

[0026] FIG15 is a schematic diagram of a flow chart of a pre-positioned single strategy provided in an embodiment of the present application;

[0027] FIG16 is a flow chart of a pre-emptive backup strategy provided in an embodiment of the present application;

[0028] FIG17 is a flow chart of a diversion identification strategy provided in an embodiment of the present application;

[0029] FIG18 is a flowchart illustrating a keyword matching process according to an embodiment of the present application;

[0030] FIG19 is a schematic diagram of a process of standard word matching provided in an embodiment of the present application;

[0031] FIG20 is a schematic diagram of a process of entity recognition provided by an embodiment of the present application;

[0032] FIG21 is a schematic diagram of an interface for configuring categories, terms, and corpora provided in an embodiment of the present application;

[0033] FIG22 is a schematic diagram of an interface for configuring standard words and spoken words provided in an embodiment of the present application;

[0034] FIG23 is a schematic diagram of a semantic wildcarding process provided by an embodiment of the present application;

[0035] FIG24 is a flow chart of a semantic processing logic provided in an embodiment of the present application;

[0036] FIG25 is a schematic diagram of a flow chart of semantic result processing provided in an embodiment of the present application;

[0037] FIG26 is a structural block diagram of a voice navigation method and apparatus provided in an embodiment of the present application;

[0038] Figure 27 is an internal structure diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0039] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0040] The voice navigation method provided in the embodiment of the present application can be applied in the application environment shown in Figure 1. In which, the terminal 102 communicates with the server 104 through a communication network. The data storage system can store data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. The user dials the corresponding access number through the terminal 102, for example, by phone, and makes a voice call with the IVR system running on the server 104 after the call is successful. The IVR system captures the first audio of the user during the voice call. The IVR system converts the first audio into text to obtain an audio text; determines the business line associated with the access number, and determines the semantic recognition strategy and reference semantics configured for the business line; performs corresponding semantic recognition processing on the audio text according to the semantic recognition strategy; obtains N slot semantics based on the recognized semantics and the pre-configured N semantic slots; determines the semantic matching result based on the N slot semantics and the reference semantics; and voice broadcasts the second audio that matches the semantic matching result.

[0041] It should be understood that the server 104 and terminal 102 in FIG. 1 are merely illustrative. Any number of servers, networks, and terminal devices may be provided, depending on implementation requirements. For example, the server 104 may be a physical server or a server cluster consisting of multiple servers, and the terminal 102 may be a mobile phone, telephone, tablet, desktop computer, laptop computer, or the like. It should be understood that embodiments of the present application may also allow multiple terminals 102 to access the server 104 simultaneously.

[0042] In some embodiments, terminal 102 provides a dialing function for the user, allowing the user to dial an access number for accessing the IVR system through terminal 102. After terminal 102 successfully accesses the IVR system, it can record the user's voice to obtain the user's first audio, and send the user's first audio to the IVR system so that the IVR system processes the first audio using the voice navigation method described in the embodiment of the present application. After terminal 102 successfully accesses the IVR system, the IVR system can also record the user's voice to obtain the user's first audio so that the IVR system processes the first audio using the voice navigation method described in the embodiment of the present application.

[0043] It should be noted that the IVR system in traditional technology can only be designed and developed for a certain business need, which makes it unsuitable for other business scenarios and unable to meet specific business needs, resulting in poor versatility of voice navigation through the IVR system.

[0044] Based on this, the embodiments of the present application propose a new voice navigation method. Before voice navigation, different business lines are configured for different business scenarios, and different semantic recognition strategies are configured for different business lines. This can effectively meet different business needs. Moreover, in the embodiments of the present application, regardless of which semantic recognition strategy is executed, the semantics identified by the semantic recognition strategy can be converted into N slot semantics in a fixed format to achieve compatibility with the semantic recognition strategy, thereby ensuring that the fixed format N slot semantics can be directly used subsequently to perform semantic matching and voice navigation in a universal manner, thereby improving the versatility of voice navigation.

[0045] In one embodiment, as shown in FIG2 , a voice navigation method is provided. The method is described by applying the method to a server. The method can be implemented by the server alone or through interaction between the server and a terminal. In this embodiment, the voice navigation method includes but is not limited to the following steps:

[0046] Step S202: convert the acquired first audio into text to obtain an audio text.

[0047] The first audio can be input via an incoming call access code. The access code can be a specific number that a user dials through their terminal to access the IVR system. After dialing the access code, the user's terminal will connect to the IVR system, and then they can interact with the IVR system through keystrokes or voice. The first audio can be the audio of the user speaking during voice interaction with the IVR system after successfully dialing the access code through their terminal.

[0048] For IVR systems, converting the first audio to text can enable automated voice interaction. For example, if a user dials a customer service hotline and uses voice input to generate the first audio during the call, the IVR system can convert the first audio to the corresponding audio text, thereby understanding the user's business needs and providing the user with the corresponding service.

[0049] Specifically, a user dials an access number through a terminal to access an IVR system running on a server and initiates a conversation with the IVR system. The IVR system then captures the user's spoken audio during the conversation via a microphone or telephone line to obtain a first audio message. After receiving the first audio message, the IVR system converts the first audio message into corresponding audio text.

[0050] In some embodiments, the acquired first audio is converted into text to obtain audio text, including: identifying the emotion value of the first audio based on the tone of the first audio; and converting the first audio into text through a text conversion model corresponding to the emotion value to obtain audio text.

[0051] The emotion value may include positive emotion values, negative emotion values, etc. If the user expresses an emotion such as excitement, the corresponding emotion value may be set to a positive emotion value; if the user expresses an emotion such as frustration, the corresponding emotion value may be set to a negative emotion value. The tone of the first audio is matched with a preset tone to determine a matching preset tone, and the emotion value corresponding to the matched preset tone is determined as the emotion value of the first audio.

[0052] Different sentiment values ​​correspond to different text conversion models, and different text conversion models correspond to different network parameters. The network parameters corresponding to the text conversion model can be adjusted according to the model training results.

[0053] This embodiment can accurately determine the emotion value by combining the tone of the first audio, and then convert the first audio into text based on the emotion value, so as to accurately identify the audio text.

[0054] Step S204: determining the service line associated with the access number, and determining the semantic recognition strategy and reference semantics configured for the service line.

[0055] In an embodiment of the present application, a specific access number can be bound to a specific business line. Each business line generally represents a specific business function or service. For example, a bank may set up multiple business lines, such as account inquiry, transfer business, and loan consultation. Different business lines can use different access numbers to establish contact with users. For example, a user can dial the access number "111" to establish contact with a business line related to account inquiry business, a user can dial the access number "222" to establish contact with a business line related to transfer business, and a user can dial the access number "333" to establish contact with a business line related to loan consultation business.

[0056] Each business line has a specific configuration, including semantic recognition strategies and reference semantics. These configurations are used to meet the needs of that business line and provide the appropriate services to users. By binding an access number to a specific business line, the IVR system can automatically identify the business line the user dials based on the access number they dial, and then perform subsequent voice navigation based on the business line's specific configuration.

[0057] The semantic recognition strategy may represent a strategy for performing semantic recognition on audio text to understand user needs. The reference semantics may be used to provide reference sentences, keywords, or phrases for semantic recognition.

[0058] For example, for a business line providing hotel reservation services, a semantic recognition strategy might be configured to recognize semantics for keywords such as "booking," "hotel," and "date," with a corresponding reference semantic such as "I want to book a hotel." By configuring specific semantic recognition strategies and reference semantics for a business line, the IVR system can more accurately recognize and process user voice input, understand user needs, and provide services related to the business line, such as hotel reservations, based on specific business scenarios.

[0059] Specifically, before implementing voice navigation, administrators can configure different lines of business in the IVR system, as well as corresponding semantic recognition policies and reference semantics for each line of business. Since each line of business is associated with an access number, users can dial the appropriate access number to access the IVR system based on their actual business needs. The IVR system then determines the line of business associated with the incoming access number, along with the semantic recognition policy and reference semantics configured for that line of business.

[0060] Step S206 , performing semantic recognition processing on the audio text according to the semantic recognition strategy, and obtaining N slot semantics based on the recognized semantics and the pre-configured N semantic slots.

[0061] Different semantic recognition strategies are used to perform semantic recognition in different ways. The semantic recognition strategy may include: a model used in the semantic recognition process, a condition for calling the model, and the like.

[0062] The N preconfigured semantic slots can be a fixed set of slots predefined in the IVR system. Each semantic slot is used to place a different semantic. The N slot semantics placed in each semantic slot are combined to obtain the user's final intent. It is understood that N can be a positive integer greater than or equal to 2.

[0063] Specifically, the IVR system determines the model to be called based on the semantic recognition strategy and determines the conditions for calling the model. When the conditions are met, the model is called to perform semantic recognition processing on the audio text to obtain the recognized semantics. Since the formats of the semantics identified by different semantic recognition strategies may differ, for example, the results of semantic recognition using a certain semantic recognition strategy may only include two semantics, while the results of semantic recognition using another semantic recognition strategy may include five semantics. In order to eliminate the above-mentioned format differences, the recognized semantics can be matched with N pre-configured semantic slots, so that N slot semantics with a fixed format can be obtained. For example, if there are four semantic slots, the two recognized semantics can be matched to two of the semantic slots respectively, or four semantics can be selected from the five recognized semantics and matched to the four semantic slots to obtain four slot semantics with a fixed format, thereby achieving compatibility of the semantic recognition strategies and ensuring that the fixed format slot semantics can be directly used later, realizing semantic matching and voice navigation processing in a universal manner. Step S208: Determine a semantic matching result based on the N slot semantics and the reference semantics, and perform voice broadcasting on the second audio that matches the semantic matching result.

[0064] Among them, the semantic matching result can be an evaluation result obtained after comparing the slot semantics with the reference semantics. It is understandable that the evaluation result may include: at least one of the slot semantics that successfully matches the reference semantics or the default value. It should be understood that the default value is a preset value automatically returned by the IVR system when the slot semantics fails to match the reference semantics. The default value can serve as a placeholder, and the default value can indicate that the IVR system cannot obtain or match specific semantic information. In actual applications, the default value can be "default" or "*".

[0065] It is understood that the semantic matching result can be used to indicate the matching status between the slot semantics and the reference semantics. The semantic matching results can include a successful match or a failed match. A successful match means that one of the N slot semantics successfully matches the reference semantics; a failed match means that none of the N slot semantics successfully matches the reference semantics.

[0066] In some embodiments, administrators can configure corresponding dialogues for different matching scenarios in advance. When the slot semantics successfully matches the reference semantics, the configured dialogue can be a guidance dialogue, which can be used to guide the user to handle the corresponding business. When the slot semantics fail to match the reference semantics, the configured dialogue can be a follow-up dialogue, which can be used to guide the user to the next round of voice interaction to provide more detailed information.

[0067] Specifically, the IVR system can use different algorithms or models to sequentially match N slot semantics with reference semantics. When a slot semantic matches the reference semantics successfully, the successfully matched slot semantics are retained. When a slot semantic matches the reference semantics unsuccessfully, a default value can be returned and placed in the original semantic slot to indicate that the semantic slot did not match the appropriate reference semantics. After the IVR system completes the semantic matching, it combines the retained semantics or default values ​​in each semantic slot to obtain a semantic matching result. If the semantic matching result indicates that the slot semantics matches the reference semantics successfully, it means that the IVR system has successfully identified the user's final intent. In this case, the corresponding guidance speech can be determined, and the guidance speech can be converted into audio to obtain a second audio, and the second audio can be broadcast to guide the user to handle the corresponding business. If the semantic matching result indicates that the slot semantics fails to match the reference semantics, it means that the IVR system has not yet successfully identified the user's final intent. In this case, the corresponding follow-up question speech can be determined, and the follow-up question speech can be converted into audio to obtain a second audio, and the second audio can be broadcast to guide the user to the next round of voice interaction.

[0068] In some embodiments, as shown in FIG3 , the overall process of the voice navigation method includes the following: First, it is necessary to determine whether there is a situation in the IVR system where the next level returns to the current level. If so, it is necessary to clear the slots, that is, clear the values ​​in each semantic slot. This is because when the user returns to the first level, the context of the semantic slots in the previous level may no longer be valid. For example, in a voice dialogue for booking movie tickets, the user selects a movie in the movie selection level and then enters the seat selection level. However, when the user returns to the movie selection level, the previously selected seat information will no longer be valid. In this case, it is necessary to clear the values ​​of the semantic slots in the seat selection level to avoid using incorrect, expired information. If there is no situation where the next level returns to the current level, it is further determined whether the automatic speech recognition (ASR) has timed out or whether the speech recognition result is empty. If the ASR has timed out or the speech recognition result is empty, the timeout logic is jumped to, such as re-performing speech recognition or guiding the user to rephrase. If the ASR does not time out and the speech recognition result is not empty, semantic recognition is performed using the corresponding semantic recognition strategy, such as the corresponding natural language understanding (NLU) model strategy. If the semantic recognition result is empty, the system jumps to the rejection logic, such as providing a prompt to the user to indicate that the system cannot understand the user's intent and requesting the user to express it again. If the semantic recognition result is not empty, the value in each semantic slot is completely replaced with the newly recognized semantics, and the semantic recognition logic is entered to further analyze the user's intent and business needs.

[0069] In some embodiments, it is determined whether there is a situation in the IVR system where the next level returns to the current level.

[0070] In some embodiments, before step S202, the voice navigation method of the embodiment of the present application specifically also includes but is not limited to: responding to the business line configuration operation in the business configuration interface, determining the semantic recognition strategy and reference semantics for the business line configuration operation, and completing the configuration of the business line; responding to the version creation operation in the business configuration interface, binding the voice navigation version to be created to the configured business line to obtain the created voice navigation version; publishing the voice navigation module corresponding to the created voice navigation version, and binding the corresponding access number to the created voice navigation version.

[0071] The service configuration interface refers to the interface used for service configuration in the IVR system. Business line configuration operations refer to operations for configuring and managing information related to a specific business line. For example, operations such as adding a business line and customizing it to meet different business scenarios and needs. Version creation operations refer to the creation of a new voice navigation version, or IVR version, within the service configuration interface. A voice navigation module is the module used to implement the voice navigation functionality of the corresponding version.

[0072] Specifically, the administrator can add the business line that needs to be configured in the business configuration interface, and configure the corresponding semantic recognition strategy and reference semantics for the business line. The IVR system responds to the above-mentioned business line configuration operation and determines the semantic recognition strategy and reference semantics configured for the business line to be configured. The administrator can also create a new IVR version in the business configuration interface. The IVR system responds to the above-mentioned version creation operation and binds the created IVR version to the previously configured business line to obtain the created IVR version. Then, the administrator clicks on the control for triggering the version release in the business configuration interface and selects the access number corresponding to the created IVR version. The IVR system responds to the above-mentioned version release and access number selection operations to release the voice navigation module corresponding to the created voice navigation version, and binds the corresponding access number to the released voice navigation version.

[0073] As can be seen, in this embodiment, the above process achieves the binding of access numbers to IVR versions and subsequent voice navigation-related configuration, allowing for rapid adaptation to different business scenarios and needs. Furthermore, administrators can flexibly adjust configurations based on actual needs without the need for tedious code modifications and redeployment, saving development and maintenance costs.

[0074] In some embodiments, the business configuration interface is shown in Figure 4, including a text box control for configuring the business line, a text box control for configuring the semantic recognition strategy, and a text box control for configuring the reference semantics. After entering the corresponding values ​​in these text box controls, click OK to complete the relevant configuration of the business line.

[0075] In some embodiments, a schematic diagram of the interface for publishing a voice navigation version (i.e., IVR version) is shown in Figure 5. The left side illustrates the various published IVR versions, while the right side is used to publish new IVR versions. Before publishing a new IVR, it is necessary to determine the name of the current IVR version, the access number associated with the current IVR version, and the release time of the current IVR version. Furthermore, corresponding operations can be performed on the currently published IVR version, such as stopping service, creating a new release, rolling back a version, and updating a release. Stopping service refers to stopping the voice navigation service provided by the current IVR version, creating a new release refers to creating a new IVR version and releasing it, rolling back a version refers to rolling back the current IVR version to a previous IVR version, and updating a release refers to updating and releasing the current IVR version. Furthermore, the release history of each IVR version can be queried by entering a start time and end time. This release history includes the name of each IVR version, the associated access number, and the release time.

[0076] In some embodiments, step S206 specifically includes: determining a semantic recognition model configured by a semantic recognition strategy; performing semantic recognition on the audio text through the semantic recognition model to obtain recognized semantics.

[0077] Among them, the semantic recognition model refers to a model that converts audio text into structured semantic representation. The semantic recognition model is used to help the IVR system understand user intentions and business needs.

[0078] Specifically, the IVR system determines one or more semantic recognition models configured for the semantic recognition strategy. The IVR system calls one or more semantic recognition models according to the semantic recognition strategy to perform corresponding semantic recognition on the audio text through the called one or more semantic recognition models to obtain recognized semantics. It can be seen that in this embodiment, by determining the semantic recognition strategy corresponding to a specific business line and calling the semantic recognition model configured for the semantic recognition strategy to perform semantic recognition, compared with the traditional technology that can only call a single semantic recognition model for semantic recognition, it can be applied to more business scenarios and business needs and has higher flexibility.

[0079] In some embodiments, a schematic diagram of the interface for configuring semantic recognition strategies and semantic recognition models is shown in FIG6 . The administrator can enter the name of the corresponding business line in this interface and query to obtain the business line that needs to configure the semantic recognition strategy and semantic recognition model, such as the business line "Customer Service 229". At this time, the administrator can bind the corresponding semantic recognition strategy and semantic recognition model to the business line "Customer Service 229". Specifically, the administrator can select an appropriate semantic recognition strategy, such as a combined fallback strategy, which is also called a combined recognition strategy; thereafter, the administrator can select and bind multiple semantic recognition models, such as selecting and binding a first recognition model and a second recognition model. After the administrator completes the above configuration, click OK to complete the process of configuring the semantic recognition strategy and semantic recognition model for the business line "Customer Service 229". In this way, after subsequently obtaining the user's audio text, the IVR system can call the corresponding semantic recognition model according to the previously configured semantic recognition strategy to perform semantic recognition processing on the audio text.

[0080] In some embodiments, before implementing voice navigation, in addition to the business configuration shown in the above embodiments, semantic configuration, interaction node configuration, service routing configuration, signpost voice configuration, business line configuration and model configuration are also required.

[0081] Exemplarily, the administrator can perform semantic configuration in the semantic configuration page shown in FIG7 . Specifically, the administrator can query the corresponding semantic configuration information by entering the keywords of the reference semantics, the selected business line, and the selected signpost voice. Each piece of semantic configuration information includes reference semantics, signpost voice encoding, signpost voice text, interactive node, business line, whether it is valid, updater, and operation controls for editing or deleting the corresponding semantic configuration information. Among them, signpost voice refers to the combination of speech content and corresponding audio. Usually, a signpost voice is configured for each reference semantic, so that there will be content for dialogue with users for different semantics; signpost voice encoding refers to the encoding that can represent the voice features of the signpost voice; signpost voice text refers to the text content corresponding to the signpost voice; interactive node refers to the node used for interaction to handle the corresponding business; whether it is valid is used to indicate whether the semantic configuration information currently queried is valid; updater refers to the administrator who updates the semantic configuration information.

[0082] As another example, the administrator can configure the interactive node in the interactive node configuration page shown in Figure 8. Specifically, each interactive node configuration information includes the interactive node name, whether it is a business exit node, the business exit node type, the service routing target, the business line, the updater, the update time, and the operation control for editing or deleting the corresponding semantic configuration information. Among them, the interactive node name refers to the name of the configured interactive node; whether it is a business exit node is used to indicate whether the interactive node is configured with a business exit; the business exit node type refers to the type of business exit configured for the interactive node; the service routing target refers to the final destination of the call determined based on the voice command or key selection input by the user, and the logical rules preset by the IVR system; the update time refers to the time when the administrator updates the configuration information of the interactive node.

[0083] As another example, the administrator can configure the service routing in the service routing configuration page shown in FIG9. Specifically, the administrator can query the corresponding service routing configuration information by entering the service exit name, exit service channel, and the selected business line. The service exit name refers to the exit name for the user to exit the IVR system, and the exit service channel refers to the specific service channel or channel that directs the user to leave the IVR system. Each piece of service routing configuration information includes the service exit name, service exit description, business line, exit service channel, channel service identifier, service channel description, service channel routing type, service channel routing entry, and an operation control for editing or deleting the corresponding semantic configuration information. Among them, the service exit description is used to explain the corresponding service exit; the channel service identifier refers to the identifier corresponding to the specific service channel or channel that directs the user to leave the IVR system; the service channel routing type refers to the method or strategy for directing the user to the correct service channel; the service channel routing entry refers to the specific entrance used to enter the IVR system, that is, how the user accesses the IVR system to select the required service channel.

[0084] As another example, administrators can configure signpost voice on the Signpost Voice Configuration page shown in Figure 10. Specifically, administrators can query corresponding signpost voice configuration information by entering the signpost voice text, signpost voice code, and selected business line. Each signpost voice configuration information includes the signpost voice code, signpost voice text, business line, updater, update time, and control elements for editing or deleting the corresponding semantic configuration information.

[0085] For example, the administrator can configure the business line in the business line configuration interface shown in Figure 11. Each business line configuration information includes the business line, semantic recognition model name, updater, update time, and operation controls for editing, exporting, unbinding, and deleting the corresponding business line configuration information.

[0086] As another example, the administrator can configure the model in the model configuration page shown in Figure 12. Specifically, each model configuration information includes model identification, model name, algorithm type, data set, model effect, model status, model category, metadata business line, updater, update time, and operation controls for metadata download, refresh status, and offline for the model configuration information. Among them, the model identification refers to the identification used to represent the corresponding semantic recognition model, the model name refers to the name of the corresponding semantic recognition model, the algorithm type refers to the type of algorithm used by the corresponding semantic recognition model, the data set refers to the data set used when training the corresponding semantic recognition model, the model effect refers to the accuracy of semantic recognition using the corresponding semantic recognition model, the model status is used to indicate whether the corresponding semantic recognition model is successfully trained, the model category is used to indicate the way of semantic recognition using the corresponding semantic recognition model, and the metadata business line refers to the business line bound to the corresponding semantic recognition model.

[0087] In some embodiments, the semantic recognition strategy includes a single recognition strategy, also referred to as a single strategy, which means that the semantic recognition strategy calls upon only one semantic recognition model. As shown in Figure 13, the process for a single recognition strategy includes: determining a semantic recognition model, which is the model bound to the semantic recognition strategy, such as an NLU model; then performing semantic recognition on the audio text based on the NLU model type to obtain recognized semantics.

[0088] In some embodiments, the semantic recognition strategy includes a combined recognition strategy; the semantic recognition model configured by the semantic recognition strategy includes a first recognition model and a second recognition model. The step of "performing semantic recognition on the audio text using the semantic recognition model to obtain recognized semantics" specifically includes, but is not limited to: performing semantic recognition on the audio text using the first recognition model; and if the first recognition model fails to perform semantic recognition on the audio text, performing semantic recognition on the audio text using the second recognition model to obtain recognized semantics.

[0089] The combined recognition strategy, also known as the combined fallback strategy, refers to a strategy that uses multiple semantic recognition models to sequentially perform semantic recognition on audio text. It is understood that when one semantic recognition model fails to identify a semantic result, the IVR system can still perform semantic recognition on the audio text using other semantic recognition models. The first recognition model, also known as the primary model, is the model that prioritizes semantic recognition. The second recognition model, also known as the secondary model, is an alternative model used for semantic recognition.

[0090] Specifically, the IVR system calls the first recognition model (also called the main model) according to the combined recognition strategy to perform semantic recognition on the audio text through the first recognition model. If semantic recognition fails through the first recognition model, the second recognition model (also called the secondary model) is called to perform semantic recognition on the audio text through the second recognition model to obtain recognized semantics. If semantic recognition is successful through the first recognition model, the semantics recognized by the first recognition model are directly determined without using the second recognition model for recognition.

[0091] It can be seen that in this embodiment, it is possible to verify whether the first recognition model can recognize semantic processing, thereby achieving the purpose of verifying the effect of the first recognition model. When the first recognition model cannot recognize the semantics, the second recognition model can be used as a backup, so that the user's intention is not missed, thereby improving the accuracy of semantic recognition. In addition, after determining the specific semantic recognition model, the determined semantic recognition model is used for semantic recognition during this voice navigation interaction process, which can ensure that the determined semantic recognition model can recognize more intentions and further improve the accuracy of semantic recognition.

[0092] In some embodiments, as shown in FIG14 , the process of the combined recognition strategy includes: first determining whether to use the semantic recognition model used last time in this round of conversation; if used, jump to the rejection logic. If not used, first use the first recognition model for semantic recognition, and determine the semantic recognition result returned by the first recognition model. If the semantic recognition result returned by the first recognition model is not empty, it is determined that the first recognition model is used in this round of conversation, and the semantic recognition result of the first recognition model is returned. If the semantic result returned by the first recognition model is empty, the second recognition model is used for semantic recognition, and it is determined whether the semantic recognition result returned by the second recognition model is empty. If the semantic recognition result returned by the second recognition model is empty, the semantic recognition result of the first recognition model is returned. If the semantic recognition result returned by the second recognition model is not empty, it is determined that the second recognition model is used in this round of conversation, and the semantic recognition result of the second recognition model is returned.

[0093] In some embodiments, the semantic recognition strategy further includes a pre-recognition strategy and a single recognition strategy; the semantic recognition model configured by the pre-recognition strategy includes a third recognition model, and the third recognition model can also be referred to as a pre-model. The semantic recognition model configured by the single recognition strategy includes a fourth recognition model, and the fourth recognition model can also be referred to as a single semantic model. The step of "performing semantic recognition on the audio text through the semantic recognition model to obtain the recognized semantics" specifically further includes but is not limited to: performing sensitive word recognition on the audio text through the third recognition model; in the case where the third recognition model fails to recognize sensitive words in the audio file, performing semantic recognition on the audio text through the fourth recognition model, or performing semantic recognition on the audio text through the first recognition model to obtain the recognized semantics.

[0094] Among them, the pre-recognition strategy refers to the need to perform pre-sensitive word recognition before each semantic recognition to determine whether sensitive words appear.

[0095] Specifically, the IVR system calls the third recognition model according to the pre-recognition strategy to perform sensitive word recognition on the audio text through the third recognition model. In the case of recognizing sensitive words, the corresponding semantic result is directly returned. In the case of not recognizing sensitive words, the fourth recognition model is called according to the single recognition strategy to perform semantic recognition on the audio text, or the first recognition model is called according to the combined model strategy to perform semantic recognition on the audio text first. In the case where the first recognition model fails to recognize semantics, the second recognition model is then called to perform semantic recognition on the audio text.

[0096] It can be seen that in the embodiments of the present application, by using the third recognition model to recognize sensitive words in specific scenarios, when sensitive words are recognized, the recognized semantics can be directly returned without subsequent semantic recognition processing, which can improve the efficiency of semantic recognition.

[0097] For the convenience of understanding, an example is given below. Suppose the specific scenario is to transfer to an online customer service, and the sensitive word defined according to this specific scenario can be "online customer service". Before each semantic recognition, the third recognition model can be used to perform pre-recognition of keyword strong matching on the audio text to avoid missed scenarios. For example, if the third recognition model recognizes that the audio text contains the keyword "online customer service", it means that it has recognized the sensitive word. At this time, the slot semantics of "default#online customer service#default#default" can be directly returned, and after the slot semantics match the reference semantics successfully, the call is transferred to the artificial service without performing the subsequent semantic recognition process, which greatly improves the efficiency of semantic recognition.

[0098] In some embodiments, the pre-identification strategy may include a pre-single strategy and a pre-fallback strategy. The pre-single strategy refers to a strategy that uses a third recognition model for pre-identification and, if the third recognition model fails to perform pre-identification, calls a single recognition strategy for semantic recognition. The pre-fallback strategy refers to a strategy that uses a third recognition model for pre-identification and, if the third recognition model fails to perform pre-identification, calls a combined recognition strategy for semantic recognition.

[0099] In some embodiments, as shown in FIG15 , the process of the pre-single strategy includes: first, using the third recognition model (also referred to as the pre-single model) to perform pre-single recognition on the audio text, and if the pre-single recognition result is not empty, directly returning the pre-single recognition result. If the pre-single recognition result is empty, calling the fourth recognition model (also referred to as the single semantic model) to perform semantic recognition and obtain the recognized semantics.

[0100] In some embodiments, as shown in FIG16 , the process of the pre-emptive fallback strategy includes: first, performing pre-emptive recognition on the audio text using the third recognition model. If the pre-emptive recognition result is not null, the pre-emptive recognition result is directly returned. If the pre-emptive recognition result is null, determining whether to use the semantic recognition model previously used in the current conversation round is used. If so, the semantic recognition result is directly returned. If not, semantic recognition is first performed using the first recognition model, and the semantic recognition result returned by the first recognition model is determined. If the semantic recognition result returned by the first recognition model is not null, it is determined that the first recognition model is used in the current conversation round, and the semantic recognition result of the first recognition model is returned. If the semantic recognition result returned by the first recognition model is null, semantic recognition is performed using the second recognition model, and it is determined whether the semantic recognition result returned by the second recognition model is null. If the semantic recognition result returned by the second recognition model is null, the semantic recognition result of the first recognition model is returned. If the semantic recognition result returned by the second recognition model is not null, it is determined that the second recognition model is used in the current conversation round, and the semantic recognition result of the second recognition model is returned.

[0101] In some embodiments, the semantic recognition strategy also includes a diversion recognition strategy. The step of "determining a semantic recognition model configured by the semantic recognition strategy" specifically includes, but is not limited to: determining M semantic recognition strategies configured based on the diversion recognition strategy, and M diversion ratios configured for the M semantic recognition strategies; determining a target semantic recognition strategy from the M semantic recognition strategies based on the M diversion ratios, and determining a corresponding semantic recognition model according to the target semantic recognition strategy.

[0102] Among them, the diversion recognition strategy is a strategy for semantic recognition that treats other semantic recognition strategies as a whole and then diverts different semantic recognition strategies to achieve it. M is a positive integer greater than or equal to 1. For each semantic recognition strategy, the configured diversion ratio is the ratio of the number of times the semantic recognition strategy is called to the preset number of times when the semantic recognition strategy is used to perform a preset number of semantic recognitions. It should be noted that under the same diversion recognition strategy, the sum of the diversion ratios corresponding to each semantic recognition strategy is 1.

[0103] The target semantic recognition strategy refers to the semantic recognition strategy currently selected in the diversion recognition strategy.

[0104] Specifically, the IVR system determines multiple semantic recognition strategies configured for the diversion recognition strategy, as well as M diversion ratios configured for each of the multiple semantic recognition strategies. Based on the M diversion ratios, the IVR system preferentially selects the semantic recognition strategy with the largest diversion ratio from the multiple semantic recognition strategies as the target semantic recognition strategy. The IVR system calls the corresponding semantic recognition model according to the target semantic recognition strategy to perform semantic recognition on the audio text using the semantic recognition model to obtain recognized semantics.

[0105] As can be seen, in this embodiment, by designing a diversion recognition strategy, different semantic recognition strategies can be used in the same business line, which provides better flexibility. In addition, semantic recognition models that need to be used together can be bound together to meet the implementation needs of different business scenarios.

[0106] In an actual application scenario, suppose the administrator has launched two semantic recognition models in the IVR system. These two models include: front-end model 1 (mainly for screening sensitive words and the like) + semantic recognition model 1 (for main semantic recognition). In response to the above situation, you can consider using a front-end single strategy. That is, based on the original front-end model 1 and semantic recognition model 1, some new business scenarios are added to train the front-end model 2 and semantic recognition model 2. However, since the front-end model 2 and semantic recognition model 2 are both in the trial operation and verification stage, in order to ensure the smooth progress of the voice navigation process, not too much traffic will be allocated to the front-end model 2 and semantic recognition model 2 in the early stage. Therefore, you can consider allocating a diversion ratio of about 5% to the front-end model 2 and semantic recognition model 2. However, since the front-end model 2 and semantic recognition model 2 represent an overall semantic recognition strategy, the previous strategy does not meet the traffic distribution requirements. Therefore, a diversion recognition strategy can be used to configure different semantic recognition strategies separately. The specific configuration is as follows: select a specific diversion recognition strategy, add a pre-single strategy 1 to the diversion recognition strategy, configure pre-model 1 and semantic recognition model 1 for the pre-single strategy 1, and configure a corresponding diversion ratio of 95%; add a pre-single strategy 2 to the diversion recognition strategy, configure pre-model 2 and semantic recognition model 2 for the pre-single strategy 2, and configure a corresponding diversion ratio of 5%, thereby achieving the simplest diversion strategy configuration. In this way, it can ensure that the new semantic recognition model is trial-run and data is collected on the basis of the availability of the original semantic recognition model, and the new semantic recognition model is optimized based on the results of data collection, thereby ensuring that the semantic recognition model is continuously improved at a relatively low cost.

[0107] It should be noted that if you want to implement multiple semantic recognition strategies in parallel, you can add corresponding semantic recognition strategies based on the above configuration and configure the diversion ratio for multiple semantic recognition strategies. Moreover, the above semantic recognition strategy is not limited to a single semantic recognition strategy, but any semantic recognition strategy can be used, that is, it can be any one of the following: a single strategy, a pre-positioned single strategy, a combined backup strategy, or a pre-positioned combined backup strategy, as long as the total diversion ratio corresponding to each semantic recognition strategy is guaranteed to be 100%.

[0108] In some embodiments, as shown in FIG17 , the flow of the diversion recognition strategy includes: determining whether to use the semantic recognition model used in the previous round of conversation; if so, directly returning the semantic recognition result; if not, determining the current diversion semantic recognition strategy according to the diversion recognition strategy, using this semantic recognition strategy as the semantic recognition strategy for the current voice interaction, and using this semantic recognition strategy to perform semantic recognition on the audio text to obtain a semantic recognition result.

[0109] In some embodiments, the identified semantics include semantic parameters under multiple semantic attributes; each semantic slot is used to store semantic parameters under a slot semantic attribute. Step S206 specifically includes, but is not limited to, matching the slot semantic attribute corresponding to each of the N semantic slots with the multiple semantic attributes, and placing the semantic parameters under the successfully matched semantic attributes into the corresponding semantic slot to obtain the slot semantics.

[0110] Semantic attributes describe the meaning and intent of corresponding semantic parameters and can also be understood as describing the components and relationships within a sentence or text. Semantic attributes can be defined based on specific tasks and business scenarios. Slot semantic attributes describe the meaning and intent of the semantic parameters placed in the corresponding semantic slots.

[0111] Specifically, for each current semantic slot among the N semantic slots, the IVR system matches the slot semantic attributes corresponding to the current semantic slot with the N semantic attributes to see if the match is successful. In the case of attribute matching failure, the IVR system does not fill any semantic parameters into the current semantic slot, or fills the current semantic slot with default values. If the attribute matching is successful, the IVR system places the semantic parameters under the successfully matched semantic attributes into the current semantic slot to obtain the corresponding slot semantics.

[0112] It can be seen that in the embodiment of the present application, by matching the slot semantic attributes of each semantic slot with the identified semantic attributes, the corresponding semantic parameters are placed in the semantic slot only when the semantic match is successful. In this way, it is possible to ensure that the semantic parameters placed in each semantic slot are matched with the slot semantic attributes. Moreover, by performing semantic matching in fixed semantic slots, the slot semantics can be encapsulated into a unified result, thereby achieving the purpose of being compatible with the output of different semantic recognition models. In some embodiments, the multiple semantic slots include a subject semantic slot, an action semantic slot, a first parameter semantic slot, and a second parameter semantic slot; the slot semantic attributes corresponding one-to-one to the multiple semantic slots include a subject semantic attribute, an action semantic attribute, a first parameter attribute, and a second parameter attribute.

[0113] Among them, the subject semantic attribute refers to the entity or role that performs an action or is affected by the action in the text, such as a person, object, organization or concept, which is also the most basic component of semantic representation.

[0114] An action semantic attribute refers to the behavior or operation performed by a subject. It is used to describe the subject's behavior, activities, or events. An action semantic attribute can be a verb or verb phrase, indicating a specific operation or behavior performed by the subject in the context.

[0115] The first parameter attribute is the first additional information related to the action semantic attributes. It provides data or qualifications to further describe the action semantic attributes. The first parameter attribute can be an entity, attribute, or other related information, which varies depending on the context and task requirements.

[0116] The second parameter attribute is a second piece of additional information related to the action semantics. Similar to the first parameter attribute, the second parameter attribute also provides additional information to further describe the action semantics. The second parameter attribute can be an entity, attribute, or other related information, depending on the context and task requirements.

[0117] In some embodiments, for each of the N semantic slots, if there are multiple semantic parameters under the semantic attribute that successfully matches the slot semantic attribute corresponding to the semantic slot, the priority of the successfully matched semantic parameter is determined; the semantic parameter with the highest priority is placed in the corresponding semantic slot. Among them, the priority of the semantic parameter can be set according to actual needs. Usually, the priority of the sub-category semantic parameter is greater than the priority of the large-category semantic parameter. For example, "bill" belongs to the large-category semantic parameter, and "phone bill" belongs to the sub-category semantic parameter under "bill". Therefore, the priority of "phone bill" is higher than the priority of "bill".

[0118] For example, for the subject semantic slot, if the semantic parameters under the semantic attribute that successfully matches the slot semantic attribute corresponding to the subject semantic slot include: bill, phone bill, since the priority of "phone bill" is higher than the priority of "bill", "phone bill" is placed in the subject semantic slot.

[0119] When there are multiple semantic parameters under the successfully matched semantic attribute, this embodiment can accurately determine the semantic parameters in the semantic slot by determining the priority of each semantic parameter, thereby improving the accuracy of the slot semantics and facilitating the recognition of intent.

[0120] In some embodiments, for the subject semantic slot, the IVR system performs attribute matching on the subject semantic attribute with multiple semantic attributes respectively, so as to place the semantic attribute parameters under the successfully matched semantic attribute into the subject semantic slot, and obtain the slot semantics corresponding to the subject semantic slot. For the action semantic slot, the IVR system performs attribute matching on the action semantic attribute with multiple semantic attributes respectively, so as to place the semantic attribute parameters under the successfully matched semantic attribute into the action semantic slot, and obtain the slot semantics corresponding to the action semantic slot. For the first parameter semantic slot, the IVR system performs attribute matching on the first parameter attribute with multiple semantic attributes respectively, so as to place the semantic attribute parameters under the successfully matched semantic attribute into the first parameter semantic slot, and obtain the slot semantics corresponding to the first parameter semantic slot. For the second parameter semantic slot, the IVR system performs attribute matching on the second parameter attribute with multiple semantic attributes respectively, so as to place the semantic attribute parameters under the successfully matched semantic attribute into the second parameter semantic slot, and obtain the slot semantics corresponding to the second parameter semantic slot.

[0121] It should be noted that if the semantic attribute corresponding to the subject semantic attribute is not matched, the semantic parameter will not be placed in the subject semantic slot, or the default value will be placed in the subject semantic slot. If the semantic attribute corresponding to the action semantic attribute is not matched, the semantic parameter will not be placed in the action semantic slot, or the default value will be placed in the action semantic slot. If the semantic attribute corresponding to the first parameter attribute is not matched, the semantic parameter will not be placed in the first parameter semantic slot, or the default value will be placed in the first parameter semantic slot. If the semantic attribute corresponding to the second parameter attribute is not matched, the semantic parameter will not be placed in the second parameter semantic slot, or the default value will be placed in the second parameter semantic slot.

[0122] It can be seen that in the embodiment of the present application, a four-slot model is proposed, namely, the subject semantic attribute, the action semantic attribute, the first parameter semantic attribute and the second parameter semantic attribute. This is similar to the subject-predicate-object model, that is, the intention is subdivided and classified, and at the same time each semantic parameter can be placed in a different slot and combined into a semantic, that is, an intention result, which can improve the accuracy of intent recognition.

[0123] In some embodiments, the semantic recognition model includes a keyword recognition model, an entity recognition model and a regular recognition model. Based on the keyword recognition model, the audio text is matched with keywords to obtain recognized semantics; based on the entity recognition model, the audio text is recognized with entities to obtain recognized semantics; based on the regular recognition model, the audio text is regular matched to obtain recognized semantics.

[0124] In some embodiments, the multiple semantic slots of the present application are four semantic slots, that is, they include four slots for placing semantic parameters. As shown in Figure 18, the process of using the keyword recognition model for keyword matching includes: judging whether there is a major category in the previous round of dialogue. If the major category is determined in the previous round of dialogue, matching is performed according to the four semantic slots of the previous round of dialogue. If the value is matched, the result is returned. If the major category is determined in the previous round of dialogue and no value is matched, or if the major category is not determined in the previous round of dialogue, the self-training platform is called to query the major category. If the major category is empty, the result of the previous round of dialogue is returned. If the major category is successfully queried, the four semantic slots corresponding to the major category query and their intentions, matching subjects, and matching remaining slots are obtained to obtain a four-slot set, wherein each slot may have multiple values. The values ​​in the slots are screened, that is, the value with the highest priority is selected according to the priority of each value.

[0125] It is understood that before keyword matching, some relevant configuration is required, such as the configuration of major categories, terms and corpus, and the configuration of standard terms and spoken words. After each keyword match, the slot information from the previous round is saved. This allows the second round of matching to determine the major categories and slot information from the previous round. The default slots are then searched for based on the major categories, and matching is performed using the standard term matching flow chart based on the configured standard terms. Furthermore, each major category has a corresponding training corpus. Before formal keyword matching, the self-training platform needs to be called to train each major category and training corpus. After training is complete, the self-training platform is called to recognize the text and obtain the corresponding major category. After determining the major category, the corresponding terms and the standard terms configured for the terms are searched based on the major category. Standard term matching is performed using the standard terms and their configured spoken words. The user's spoken text is then matched to see if a corresponding spoken word exists. If so, the match is successful, and the current slot is assigned the corresponding standard term. The remaining slots are then matched until all four slots are matched.

[0126] It should be noted that the results returned after keyword matching through the keyword matching model are subject semantic attributes, action semantic attributes, first parameter semantic attributes, second parameter semantic attributes and major category semantic attributes, etc. Since the semantic attributes corresponding to the four semantic slots of this application only include subject semantic attributes, action semantic attributes, first parameter semantic attributes and second parameter semantic attributes, and do not include major category semantic attributes. In order to ensure the structural unity of the semantic parameters, only the semantic attributes corresponding to the subject semantic attributes, action semantic attributes, first parameter semantic attributes and second parameter semantic attributes will be placed in the semantic slots, and the semantic parameters corresponding to the major category semantic attributes will not be filled into the semantic slots.

[0127] In some embodiments, as shown in FIG19 , the process of standard word matching includes: generating a dictionary tree based on the content in the mapping table, matching the first layer of the dictionary tree in turn according to each word in the content of the speech, and then searching for subtrees in turn after matching, and successfully matching one indicates success. Among them, the mapping table includes multiple spoken words corresponding to each standard word, and these multiple spoken words will generate corresponding dictionary trees, with the purpose of quickly finding results that match the spoken words. It should be noted that if a spoken word is matched during the matching process, the matching is terminated. The corresponding spoken word is then returned, and then the upper layer will determine whether the current round of matching is successful based on the returned result. If the corresponding spoken word is to be stored, it is also convenient to store it.

[0128] In some embodiments, as shown in FIG20 , the entity recognition process includes: performing an entity recognition query, obtaining information containing four slots, and returning the recognition result if the call is normal. If the call times out or the call is wrong, an empty semantic recognition result is returned. Among them, the result returned by the entity recognition includes subject semantic attributes, action semantic attributes, and parameter semantic attributes. Since the embodiment of the present application requires two parameter semantic attributes, it is possible to consider adding a parameter semantic attribute to the result returned by the entity recognition for compatibility.

[0129] In some embodiments, the administrator can make some configurations related to keyword matching in the interface of major categories, entries and corpus configuration as shown in Figure 21. It can be seen that some pre-configured major category names are shown on the left, including the impact of overdue consultation, bill inquiry, withdrawal failure, settlement certificate, membership fees and overpayment. On the right, you can configure the basic information of the major category, such as the major category name, description and status, as well as the various attributes and update time in the major category entries. In addition, the configured corpus can also be managed, such as query, add, export and import.

[0130] In some embodiments, the administrator can configure standard word matching in the interface for configuring standard words and colloquial words as shown in Figure 22. For example, the administrator can configure specific standard words, colloquial words, creators, creation time, updaters, and update time, and edit or delete one or more configuration items.

[0131] In some embodiments, regular matching uses regular expressions to match audio text and returns the matching results, such as corresponding words. The returned words are then used as the main body and encapsulated into the semantics of "default#keyword#default#default" to make the results of different models compatible and unified.

[0132] It should be noted that this application uses a universal processing method for different semantic recognition models to mask the differences in results brought about by different semantic recognition models. On this basis, combined with the use of semantic recognition strategies, voice navigation processing can be implemented in multiple business scenarios, with greater versatility.

[0133] In some embodiments, step S208 specifically includes but is not limited to: extracting P slot semantics from N slot semantics according to a preset semantic matching strategy; matching the P slot semantics with reference semantics to obtain a semantic matching result.

[0134] Wherein, P is a positive integer greater than or equal to 2 and less than or equal to N.

[0135] Specifically, the IVR system extracts some or all slot semantics of the corresponding slot from the semantic slot according to the semantic matching strategy, i.e., P slot semantics. The IVR system matches the extracted P slot semantics with the reference semantics to obtain a semantic matching result.

[0136] It can be seen that in this embodiment, not all slot semantics are directly matched with reference semantics, but part or all of the slot semantics are extracted and matched with reference semantics according to the semantic matching strategy and combined with actual needs. In this way, the semantic matching results can be more accurate.

[0137] In some embodiments, the semantic importance of each slot semantic is determined; based on the importance threshold set in the semantic matching strategy, P slot semantics having a semantic importance greater than the importance threshold are extracted from the N slot semantics. The importance threshold can be set according to actual needs.

[0138] The semantic importance of each slot semantics can be determined based on the semantic slot corresponding to the slot semantics. For example, the semantic importance of the subject semantics slot is greater than the semantic importance of the action semantics slot. The semantic importance of each slot semantics can also be determined based on the frequency of the slot semantics in the audio text. The semantic importance of the slot semantics is proportional to the frequency of the slot semantics in the audio text. For example, the greater the frequency of the slot semantics in the audio text, the higher the semantic importance of the slot semantics.

[0139] By determining the semantic importance of each slot semantic, this embodiment can reasonably determine P slot semantics from N slot semantics, thereby reducing the number of slot semantics and improving the matching efficiency between slot semantics and reference semantics. In some embodiments, the four semantic slots of this application are shown in the following table:

[0140] Among them, the slot name corresponding to the subject semantic slot is subject, the slot name corresponding to the action semantic slot is action, the slot name corresponding to the first parameter semantic slot is parameter 1, and the slot name corresponding to the second parameter semantic slot is parameter 2. The slot semantics obtained by combining the above slot semantics can be "Bill#Query#Current Issue#Default".

[0141] In some embodiments, as shown in FIG23 , the semantic wildcarding process includes: performing a full match on the semantics of the four slots; if the match is successful, the process ends. If the full match fails, a second match is performed, i.e., the parameter 2 slot becomes * and the matching is performed last; if the match is successful, the process ends. If the second match fails, a third match is performed, i.e., the action slot becomes * and the matching is performed last; if the match is successful, the process ends. If the third match fails, a fourth match is performed, i.e., the parameter 1 slot becomes * and the matching is performed last; if the match is successful, the process ends. If the fourth match fails, a fifth match is performed, i.e., the parameter 2 and action slots both become * and the matching is performed last; if the match is successful, the process ends. If the fifth match fails, a sixth match is performed, i.e., the parameter 1 and action slots both become * and the matching is performed last; if the match is successful, the process ends. If the sixth match fails, a seventh match is performed, i.e., the parameter 1 and parameter 2 slots both become * and the matching is performed last; if the match is successful, the process ends. If the seventh match fails, the eighth match is performed, replacing Parameter 1, Parameter 2, and the action slot with * before matching. The process ends. Assuming the identified slot semantics are "Query#BankCard#Limit#Default" and the reference semantics are "Query#BankCard#*#*" and "*#BankCard#Limit#*," referring to the matching method above, *#BankCard#Limit#* is the preferred match for this slot semantic.

[0142] In some embodiments, the second audio includes a transfer audio. Step S208 specifically includes, but is not limited to, determining the interaction node configured for the business line if the semantic match result is successful; if a service outlet is configured on the interaction node, transferring the call to the business application that matches the service outlet and playing the transfer audio. The interaction node represents the operational steps involved in handling the service that matches the semantic match result, and the transfer audio is used to indicate the service transfer status and guide the caller to handle the corresponding service in the service application.

[0143] Specifically, if the semantic match result is successful, the IVR system determines the interaction node pre-configured by the administrator for the business line. If the IVR system recognizes that a business exit is configured on the interaction node, it transfers the call to the business application that matches the business exit and plays the transfer audio.

[0144] It can be seen that in the embodiment of the present application, whether the IVR system can accurately understand the intention expressed by the user is judged by whether the interactive node has a business exit. If the business exit is identified, it means that the IVR system can accurately express the intention expressed by the user. At this time, the transfer audio is directly played to guide the user to handle specific business.

[0145] In some embodiments, the second audio also includes a follow-up audio. Step S208 specifically includes but is not limited to: if the semantic matching result is a successful semantic match and no service outlet is configured on the interactive node, determining and playing the follow-up audio according to the multiple slot semantics and the interactive node.

[0146] The follow-up audio is used to guide the next round of voice navigation interaction. For example, if the semantic match result is "default#bankcard#default#default," the user's intent hasn't been understood at this point, so the follow-up audio is needed to identify the user's more detailed intent, such as whether the user wants to check their bank card or their balance.

[0147] Specifically, if the semantic match result is successful and no service exit is configured on the interaction node, it means that the IVR system has not yet fully understood the user's expressed intent. In this case, it is necessary to query the administrator's pre-configured follow-up audio in the interface based on the N slot semantics and interaction nodes, and play the follow-up navigation audio to more accurately identify the user's intended intent and ultimately jump to the corresponding service application.

[0148] In some embodiments, the voice navigation method specifically further includes: reading the first audio of the current round input through the access number; converting the first audio of the current round into text to obtain the audio text of the current round; calling the corresponding semantic recognition model according to the semantic recognition strategy used in the previous round, so as to perform corresponding semantic recognition on the audio text of each round through the semantic recognition model, and performing matching of the recognized semantics to multiple pre-configured semantic slots to obtain N slot semantics and subsequent steps. In this way, the same semantic recognition model can perform semantic recognition for the scenes and content of multiple rounds of conversations, and can more accurately identify the user's intentions, thereby achieving better results. Among them, the current round refers to the voice interaction of the current round.

[0149] In some embodiments, as shown in FIG24 , the process of the semantic processing logic includes: semantically matching N slot semantics and reference semantics, and jumping to the rejection logic when the recognition result is empty. When the result is recognized and there is a corresponding business exit in the result, the business exit is taken. When the result is recognized and there is no corresponding business exit in the result, if the system does not automatically fill the slot, the follow-up questioning words are used, and the follow-up audio corresponding to the follow-up questioning words are played. If the system needs to automatically fill the slot, the semantic matching continues. If no result is found, the follow-up questioning words are used, and the follow-up audio corresponding to the follow-up questioning words are played. If the result is found and there is a business exit in the result, the business exit is taken.

[0150] It's important to note that automatic slot filling means the IVR system automatically adds corresponding values ​​to one or more semantic slots. For example, if the system detects that a user wants to inquire about their bill for the current period, month, or current month, it will add the corresponding time to parameter 2 based on the current system time. This allows the system to more quickly identify the intended intent and redirect the user to the appropriate application.

[0151] In some embodiments, as shown in FIG25 , the semantic result processing process includes determining whether the semantics is a business exit. If so, the corresponding business exit information is returned and the process ends. If not, the follow-up questioning script configured for the corresponding road sign voice is returned, the follow-up audio corresponding to the follow-up questioning script is played, and the process ends.

[0152] In some specific embodiments, the voice navigation method of the present application includes the following steps:

[0153] 1. Business Configuration: The administrator adds the business line to be configured in the business configuration interface and binds the corresponding semantic reference policy to the business line, as well as the semantic reference model to be invoked when executing the semantic reference policy. The administrator configures the corresponding reference semantics, roadmap voice, and interaction nodes for the business line and creates an IVR version. The administrator binds the IVR version to the configured business line, publishes the created voice navigation version, and binds the corresponding access number to the published voice navigation version.

[0154] 2. The user dials the access number and waits for the IVR system to broadcast the following information:

[0155] The IVR system performs semantic recognition based on the user's first audio message according to the semantic recognition policy configured for the service line. The policy calls the corresponding semantic recognition model for semantic recognition and processes the slot information based on the results returned by the semantic recognition model to obtain the slot semantics corresponding to the four slots. The slot semantics corresponding to the four slots are then universally matched with the reference semantics configured by the administrator on the page to obtain a semantic match result.

[0156] If no semantics are matched, or the semantics corresponding to the four slots are empty, the previous conversation or globally configured dialogue is returned, and the corresponding follow-up audio is played. If the semantics are matched, and the corresponding interaction node has a service exit, the service exit information is returned, and based on this information, the user is redirected to a different business application and the forwarding audio is played. If the semantics are matched, but the corresponding interaction node has no service exit, the semantic information and interaction node configuration information are returned based on the semantic configuration, and the corresponding follow-up dialogue is converted into follow-up audio and played for the next round of interaction with the user.

[0157] In some embodiments, the voice navigation method of the present invention can be applied to the field of call center communications. A call center is a place where people interact and conduct transactions via telephone, such as service, sales, and emergencies, and can also include manual services and self-service services.

[0158] In actual applications, the components that implement the voice navigation method of the embodiment of the present application include an interactive voice response engine (ivr-engine), a free switch (Freeswitch), automatic speech recognition (ASR), interactive voice response dialogue management (ivr-dm), interactive voice response natural language processing (ivr-nlp), interactive voice response text-to-speech synthesis (ivr-tts), and an operator gateway. Among them, ivr-dm is used to accept the navigation configuration of the administrator for the IVR system (such as configuration semantics, roadmap voice and interactive nodes, etc.), and accept the call of ivr-engine. According to the user's spoken text and the context of multiple interactions, it performs corresponding semantic matching and processing to obtain the user's final intention and return the result. ivr-engine is used to accept calls from Freeswitch and return the text audio needed to interact with the user (for example, the audio that needs to be played for interactive responses with the user after the access number is called). According to the text passed when Freeswitch is called, ivr-dm is called and the corresponding speech audio is returned. ASR is used to accept calls from Freeswitch, pass the audio spoken by the user in real time, and return the audio text corresponding to the audio. ivr-tts is used to receive calls from ivr-engine, pass the input text, and return the corresponding audio. ivr-nlp is used to receive calls from ivr-dm, pass the user's spoken text, and return the corresponding slot information. Freeswitch is a cross-platform open source telephone exchange platform with strong scalability. In the embodiment of this application, it is responsible for providing routing and interconnected communications for audio, text or any other form of media, as well as providing self-service IVR services; the transmitted voice is converted into corresponding text by calling asr, and passed to ivr-engine for intent recognition based on the text, and the audio for interaction with the user is obtained for broadcast. The carrier gateway is used to connect to the lines of major carriers and serves as a bridge for call centers to establish communication with the Public Switched Telephone Network (PSTN).

[0159] It should be noted that the embodiment of the present application proposes a general method for semantic processing, which can call the corresponding semantic recognition model according to the user's audio text to convert it into slot information. The value of each slot is a small intention, which subdivides the large intention into multiple small intention parameters. Then, by configuring each semantic (i.e., a slot composed of several small intentions) and the user's dialogue words, multiple rounds of dialogue with the user are achieved. The intention parameters obtained from each round of dialogue are assembled to obtain a semantic that can represent the user's final intention, and jump to the corresponding semantic recognition model to achieve voice dialogue with the user, which can help users with inquiries, transfers to manual work, and other related matters, greatly simplifying the time and cost of user incoming calls. Among them, the semantic recognition logic is to find the recognition result after the user speaks and match it with all the semantics we configure to obtain a corresponding semantic, which also makes the human-computer interaction dialogue process clearer and clearer. If there is a need to modify the content of the dialogue, you only need to modify the configuration on the page, which does not affect the model algorithm related to the semantic recognition model at all.

[0160] It should be understood that, although the steps in the flowcharts of the above-mentioned embodiments are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts of the above-mentioned embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0161] Based on the same inventive concept, the present application also provides a voice navigation device. The solution provided by the device is similar to the solution described in the above method. Therefore, the specific limitations of one or more voice navigation device embodiments provided below can be referred to the limitations of the voice navigation method above and will not be repeated here.

[0162] As shown in FIG26 , an embodiment of the present application provides a voice navigation device, including:

[0163] The conversion module 2602 is configured to convert the acquired first audio into text to obtain an audio text; wherein the first audio is input through an incoming call access number;

[0164] Identification module 2604, used to determine the business line associated with the access number, and determine the semantic identification strategy and reference semantics configured for the business line;

[0165] Matching module 2606, configured to perform semantic recognition processing on the audio text according to the semantic recognition strategy, and obtain N slot semantics based on the recognized semantics and N pre-configured semantic slots; wherein N is a positive integer greater than or equal to 2;

[0166] The navigation module 2608 is configured to determine a semantic matching result based on the N slot semantics and the reference semantics, and to voice broadcast the second audio that matches the semantic matching result.

[0167] In some embodiments, the voice navigation device also includes a version release module, which is used to respond to the business line configuration operation in the business configuration interface, determine the semantic recognition strategy and reference semantics for the business line configuration operation, and complete the configuration of the business line; respond to the version creation operation in the business configuration interface, bind the voice navigation version to be created to the configured business line, and obtain the created voice navigation version; release the voice navigation module 2608 corresponding to the created voice navigation version, and bind the corresponding access number to the created voice navigation version.

[0168] In some embodiments, the matching module 2606 is further configured to determine a semantic recognition model configured by a semantic recognition strategy; and perform semantic recognition on the audio text through the semantic recognition model to obtain recognized semantics.

[0169] In some embodiments, the semantic recognition strategy includes a combined recognition strategy; the semantic recognition model configured by the combined recognition strategy includes a first recognition model and a second recognition model. Matching module 2606 is further configured to perform semantic recognition on the audio text using the first recognition model; if the first recognition model fails to recognize the audio text, the second recognition model is used to perform semantic recognition on the audio text to obtain recognized semantics.

[0170] In some embodiments, the semantic recognition strategy also includes a pre-recognition strategy and a single recognition strategy; the semantic recognition model configured by the pre-recognition strategy includes a third recognition model, and the semantic recognition model configured by the single recognition strategy includes a first recognition model and a fourth recognition model. Matching module 2606 is further configured to identify sensitive words in the audio text using the third recognition model; if the third recognition model fails to identify sensitive words in the audio text, semantic recognition is performed on the audio text using the fourth recognition model, or semantic recognition is performed on the audio text using the first recognition model to obtain recognized semantics.

[0171] In some embodiments, the semantic recognition strategy includes a diversion recognition strategy. The matching module 2606 is further configured to determine M semantic recognition strategies configured based on the diversion recognition strategy, and M diversion ratios configured for the M semantic recognition strategies; wherein M is a positive integer greater than or equal to 1; for each semantic recognition strategy, the configured diversion ratio is the ratio of the number of times the semantic recognition strategy is called to the preset number of times when semantic recognition is performed using the semantic recognition strategy; based on the M diversion ratios, a target semantic recognition strategy is determined from the M semantic recognition strategies, and a corresponding semantic recognition model is determined according to the target semantic recognition strategy.

[0172] In some embodiments, the identified semantics include semantic parameters under multiple semantic attributes; each semantic slot is used to store semantic parameters under a slot semantic attribute. Matching module 2606 is further configured to, for each of the N semantic slots, perform attribute matching on the slot semantic attribute corresponding to the semantic slot with the multiple semantic attributes, and place the semantic parameters under the successfully matched semantic attributes into the corresponding semantic slot to obtain the slot semantics.

[0173] In some embodiments, the navigation module 2608 is also used to extract P slot semantics from N slot semantics according to a preset semantic matching strategy; where P is a positive integer greater than or equal to 2 and less than or equal to N; and match the P slot semantics with the reference semantics to obtain a semantic matching result.

[0174] In some embodiments, the second audio includes a transfer audio. Navigation module 2608 is further configured to, if the semantic match result is a successful semantic match, determine an interaction node configured for the business line; the interaction node represents the operational steps involved in handling the business matching the semantic match result; if a business exit is configured on the interaction node, transfer the call to the business application matching the business exit and play a transfer audio; the transfer audio is used to indicate the business transfer status and guide the user to handle the corresponding business in the business application.

[0175] In some embodiments, the second audio also includes a follow-up audio. The navigation module 2608 is further configured to, if the semantic match result is successful and no service outlet is configured on the interaction node, determine and play the follow-up audio based on the multiple slot semantics and interaction nodes; the follow-up audio is used to guide the execution of the next round of voice navigation interaction.

[0176] Each module in the above-mentioned voice navigation device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0177] In some embodiments, a computer device is provided, which may be a server, and its internal structure diagram may be shown in Figure 27. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, the memory and the input / output interface are connected via a system bus, and the communication interface is connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data related to voice navigation. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the steps in the above-mentioned voice navigation method are implemented.

[0178] Those skilled in the art will understand that the structure shown in Figure 27 is merely a block diagram of a portion of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0179] In some embodiments, a computer device is provided. The computer device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the steps in the above method embodiments are implemented.

[0180] In some embodiments, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0181] In some embodiments, a computer program product is provided. The computer program product includes a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0182] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions.

[0183] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, and the like.

[0184] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0185] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A voice navigation method, comprising: Convert the acquired first audio into text to obtain an audio text; wherein the first audio is input through the incoming call access number; Determining a service line associated with the access number, and determining a semantic recognition strategy and reference semantics configured for the service line; Performing semantic recognition processing on the audio text according to the semantic recognition strategy, and obtaining N slot semantics based on the recognized semantics and the pre-configured N semantic slots; wherein N is a positive integer greater than or equal to 2; A semantic matching result is determined according to the N slot semantics and the reference semantics, and a second audio matching the semantic matching result is voice broadcasted.

2. The method according to claim 1, further comprising: In response to a business line configuration operation in a business configuration interface, determining a semantic recognition strategy and reference semantics for the business line configuration operation, and completing configuration of the business line; In response to the version creation operation in the service configuration interface, the voice navigation version to be created is bound to the configured service line to obtain the created voice navigation version; The voice navigation module corresponding to the created voice navigation version is released, and the corresponding access number is bound to the created voice navigation version.

3. According to the method of claim 1, the step of performing semantic recognition processing on the audio text according to the semantic recognition strategy comprises: Determining a semantic recognition model configured by the semantic recognition strategy; The semantic recognition model is used to perform semantic recognition on the audio text to obtain the recognized semantics.

4. According to the method of claim 3, the semantic recognition strategy comprises a combined recognition strategy; the semantic recognition model configured by the combined recognition strategy comprises a first recognition model and a second recognition model; The performing semantic recognition on the audio text by using the semantic recognition model to obtain the recognized semantics includes: The audio text is semantically recognized by the first recognition model; if the first recognition model fails to recognize the audio text, the audio text is semantically recognized by the second recognition model to obtain the recognized semantics.

5. According to the method of claim 4, the semantic recognition strategy further comprises a pre-recognition strategy and a single recognition strategy; the semantic recognition model configured by the pre-recognition strategy comprises a third recognition model, and the semantic recognition model configured by the single recognition strategy comprises the first recognition model and a fourth recognition model; The performing semantic recognition on the audio text by using the semantic recognition model to obtain the recognized semantics includes: Performing sensitive word recognition on the audio text by using the third recognition model; In the case that the third recognition model fails to recognize the sensitive words in the audio text, semantic recognition is performed on the audio text through the fourth recognition model, or semantic recognition is performed on the audio text through the first recognition model to obtain the recognized semantics.

6. The method according to claim 3, wherein the semantic recognition strategy comprises a diversion recognition strategy; The determining of the semantic recognition model configured by the semantic recognition strategy includes: Determine M semantic recognition strategies configured based on the diversion recognition strategy, and M diversion ratios configured for the M semantic recognition strategies; wherein M is a positive integer greater than or equal to 1; for each semantic recognition strategy, the configured diversion ratio is the ratio of the number of times the semantic recognition strategy is called to the preset number of times when the semantic recognition strategy is used to perform a preset number of semantic recognitions; According to the M diversion ratios, a target semantic recognition strategy is determined from the M semantic recognition strategies, and a corresponding semantic recognition model is determined according to the target semantic recognition strategy.

7. The method according to claim 1 or 3, wherein the identified semantics includes semantic parameters under multiple semantic attributes; each semantic slot is used to place a semantic parameter under a slot semantic attribute; The N slot semantics are obtained based on the identified semantics and the pre-configured N semantic slots, including: For each of the N semantic slots, the slot semantic attribute corresponding to the semantic slot is matched with the multiple semantic attributes, and the semantic parameters under the successfully matched semantic attributes are placed in the corresponding semantic slot to obtain slot semantics.

8. According to the method described in claim 7, the multiple semantic slots include a subject semantic slot, an action semantic slot, a first parameter semantic slot and a second parameter semantic slot; the slot semantic attributes corresponding one by one to the multiple semantic slots include a subject semantic attribute, an action semantic attribute, a first parameter attribute and a second parameter attribute.

9. The method according to claim 7, wherein for each of the N semantic slots, the slot semantic attribute corresponding to the semantic slot is matched with the multiple semantic attributes, and the semantic parameters under the successfully matched semantic attributes are placed in the corresponding semantic slot to obtain the slot semantics, including: For each of the N semantic slots, if there are multiple semantic parameters under the semantic attributes that successfully match the slot semantic attribute corresponding to the semantic slot, determine the priority of the successfully matched semantic parameters; The semantic parameter with the highest priority is placed in the corresponding semantic slot.

10. The method according to claim 1, wherein determining a semantic matching result according to the N slot semantics and the reference semantics comprises: According to a preset semantic matching strategy, P slot semantics are extracted from the N slot semantics; wherein P is a positive integer greater than or equal to 2 and less than or equal to N; The P slot semantics are matched with the reference semantics to obtain a semantic matching result.

11. The method according to claim 10, wherein extracting P slot semantics from the N slot semantics according to a preset semantic matching strategy comprises: Determine the semantic importance of each slot semantics; According to the importance threshold set in the semantic matching strategy, P slot semantics whose semantic importance is greater than the importance threshold are extracted from the N slot semantics.

12. The method according to any one of claims 1 to 11, wherein the second audio comprises a switched audio; The voice broadcasting of the second audio matching the semantic matching result includes: In the case where the semantic matching result is a successful semantic matching, determining an interactive node configured for the business line; the interactive node is an operation step for handling a business matching the semantic matching result; If a service outlet is configured on the interactive node, the call is transferred to the service application matching the service outlet, and the transfer audio is played; the transfer audio is used to prompt the service transfer status and to guide the processing of services in the service application.

13. The method according to claim 12, wherein the second audio further comprises a questioning audio; The voice broadcasting of the second audio matching the semantic matching result also includes: When the semantic matching result is a successful semantic matching and the service exit is not configured on the interactive node, the follow-up audio is determined and played according to the N slot semantics and the interactive node; the follow-up audio is used to guide the execution of the next round of voice navigation interaction.

14. The method according to claim 1, wherein converting the acquired first audio into text to obtain the audio text comprises: identifying an emotion value of the first audio according to the pitch of the first audio; The first audio is converted into text using a text conversion model corresponding to the emotion value to obtain the audio text.

15. A voice navigation device, comprising: A conversion module, configured to convert the acquired first audio into text to obtain an audio text; wherein the first audio is input through an incoming call access number; an identification module, configured to determine a service line associated with the access number, and to determine a semantic identification strategy and reference semantics configured for the service line; A matching module, used for performing semantic recognition processing on the audio text according to the semantic recognition strategy, and obtaining N slot semantics based on the recognized semantics and pre-configured N semantic slots; wherein N is a positive integer greater than or equal to 2; The navigation module is used to determine a semantic matching result according to the N slot semantics and the reference semantics, and to voice broadcast a second audio that matches the semantic matching result.

16. A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the voice navigation method according to any one of claims 1 to 14 when executing the computer program.

17. A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the voice navigation method according to any one of claims 1 to 14.

18. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the voice navigation method according to any one of claims 1 to 14 is implemented.

Citation Information

Patent Citations

  • Intelligent service method and system based on natural language interaction

    CN104199810A

  • Intelligent service method of chat robot, server and storage medium

    CN108829757A

  • Voice semantic information extraction method and device, intelligent terminal and storage medium

    CN111768766A

  • Interactive voice response method and device, electronic equipment and storage medium

    CN116153306A

  • Voice navigation method and device, computer equipment and storage medium

    CN117978921A