Telephone traffic flow configuration method, telephone traffic flow interaction method and device, electronic equipment and program product

By generating a speech node editing area and forming a path according to preset rules, the problem of users having difficulty configuring high-quality speech streams is solved, achieving efficient and flexible voice interaction configuration and improving user experience.

CN121597186APending Publication Date: 2026-03-03BAIRONG ZHIXIN (BEIJING) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411166457.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-23
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing graphical dialogue flow configuration technologies make it difficult for users to quickly and efficiently configure high-quality dialogue flows that meet their needs when faced with a large number of dialogue nodes, especially lacking flexibility when configuring complex voice interactions.

Method used

This paper provides a method for configuring a speech flow. By generating editing areas for multiple speech nodes, it traverses upstream and downstream speech nodes to form a path according to preset node association rules, and determines the start and end nodes. Combined with a visual display of the effective path, it provides cross-path copy and paste and node data analysis functions.

Benefits of technology

It enables users to conveniently and efficiently configure high-quality speech streams, improves flexibility and stability in complex voice interaction configurations, and enhances configuration efficiency and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121597186A_ABST
    Figure CN121597186A_ABST
Patent Text Reader

Abstract

The invention discloses a verbal-skill flow configuration method, a verbal-skill flow interaction method and device, electronic equipment and a program product, and the method comprises the steps: generating a verbal-skill flow editing region comprising a plurality of verbal-skill nodes which have node attributes and verbal-skill contents; according to a preset node association rule, traversing associated upstream verbal skill nodes and / or downstream verbal skill nodes to the upstream and / or the downstream from a target verbal skill node in the verbal skill flow editing area to form a verbal skill flow path; and if the verbal skill flow path has the starting verbal skill node and the ending verbal skill node, determining the verbal skill flow path as an effective path corresponding to the target verbal skill node. The invention further provides a verbal traffic interaction method and device, electronic equipment and a program product. According to the method disclosed by the invention, the user can conveniently and efficiently configure the high-quality verbal traffic flow meeting the requirements of the user in a plurality of verbal traffic flow nodes, and the flexibility of the user when the user configures complex voice interaction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, specifically to a speech flow configuration method, speech flow interaction method and device, electronic device, and program product. Background Technology

[0002] Artificial intelligence (AI) voice robots, as an emerging technology field in recent years, have shown broad application prospects in many fields such as customer service, education, and entertainment.

[0003] There are two main technical approaches for configuring the dialogue flow of AI voice robots: text-based dialogue flow configuration and graphics-based dialogue flow configuration.

[0004] Text-based script configuration technology requires users to write a large number of text scripts to describe the robot's interaction process, which not only requires high programming skills, but also makes it difficult to achieve complex voice interactions.

[0005] Graphical script configuration technology is easier for users to operate. However, the script configuration interface contains a large number of script flow nodes, making it difficult for users to configure a high-quality script flow that meets their needs. This problem is even more pronounced when users need to configure complex voice interactions. In addition, graphical script configuration technology often lacks flexibility.

[0006] Therefore, there is an urgent need for an improved speech flow configuration scheme for voice robots.

[0007] The background description is provided for the purpose of understanding the relevant technologies in this field and is not intended as an admission of prior art. Summary of the Invention

[0008] In response, this disclosure provides an improved speech flow configuration solution for voice robots, which aims to enable users to conveniently and efficiently configure a high-quality speech flow that meets their needs from a large number of speech flow nodes, and to increase the flexibility of users when configuring complex voice interactions.

[0009] In a first aspect, embodiments of this disclosure provide a speech flow configuration method, which includes:

[0010] Generate a speech flow editing area that includes multiple speech nodes, wherein each speech node has node attributes and speech content;

[0011] According to the preset node association rules, starting from the target speech node in the speech flow editing area, the associated upstream and / or downstream speech nodes are traversed upstream and / or downstream to form a speech flow path.

[0012] If the script flow path does not have a starting script node and an ending script node, then the script flow path is determined as the valid path corresponding to the target script node, wherein the node attribute of the starting script node is the starting node, and the node attribute of the ending script node is the ending node.

[0013] Secondly, embodiments of this disclosure provide a speech flow interaction method, which includes:

[0014] Displays a speech flow editing area that includes multiple speech nodes, wherein each speech node has node attributes and speech content;

[0015] In response to a user selecting a target speech node, at least one valid speech flow path containing the target speech node is displayed;

[0016] In each valid speech flow path, adjacent nodes conform to preset node association rules, and the adjacent nodes are connected by directed node lines.

[0017] Each valid speech flow path includes a starting speech node and an ending speech node. The starting speech node has the attribute of a starting node, and the ending speech node has the attribute of an ending node.

[0018] Thirdly, embodiments of this disclosure provide a speech flow configuration device, which includes:

[0019] The generation module is configured to generate a speech flow editing area that includes multiple speech nodes.

[0020] The traversal module is configured to traverse the associated upstream and / or downstream dialogue nodes from the target dialogue node in the dialogue flow editing area according to the preset node association rules, thereby forming a dialogue flow path.

[0021] The determination module is configured to determine the dialogue flow path as the valid path corresponding to the target dialogue node when there is a start dialogue node and an end dialogue node in the dialogue flow path, wherein the node attribute of the start dialogue node is a start node and the node attribute of the end dialogue node is an end node.

[0022] Fourthly, embodiments of this disclosure provide a speech stream interaction device, which includes:

[0023] The display module is configured to display a script flow editing area that includes multiple script nodes, wherein the script nodes have node attributes and script content;

[0024] The response module is configured to display at least one valid speech flow path containing the target speech node in response to a user selecting the target speech node; wherein, in each valid speech flow path, adjacent nodes conform to a preset node association rule, and the adjacent nodes are connected by directed node lines; wherein, each valid speech flow path includes a starting speech node and an ending speech node, the node attribute of the starting speech node is a starting node, and the node attribute of the ending speech node is an ending node.

[0025] Fifthly, embodiments of this disclosure provide an electronic device comprising: a processor and a memory storing a computer program, the processor being configured to implement the method as described in the first or second aspect when executing the computer program.

[0026] In a sixth aspect, embodiments of this disclosure provide a program product including a computer program that, when executed by a processor, implements the method described in the first or second aspect.

[0027] This disclosure provides a method for configuring a speech flow. It provides a speech flow editing area with multiple speech nodes. Based on preset node association rules, it starts from the target speech node in the speech flow editing area and traverses upstream and / or downstream associated speech nodes to form a speech flow path. It determines whether the speech flow path has a starting speech node and an ending speech node, thus identifying the speech flow path with both a starting and ending speech node as the valid path corresponding to the target speech node. This method facilitates efficient confirmation of the valid path for the target speech node and graphically displays the upstream and downstream relationships between speech nodes in the valid path. This allows users to easily and efficiently configure a high-quality speech flow that meets their needs from numerous speech flow nodes. Furthermore, it provides cross-path copy-paste and splicing functions for speech flows, combined with speech node completion reminders and node data analysis functions, further increasing the stability and flexibility for users configuring complex voice interaction speech flows.

[0028] This disclosure also provides a method for interactive speech flow, which displays a speech flow editing area including multiple speech nodes; in response to a user selecting a target speech node, it displays at least one valid speech flow path containing the target speech node. This method can display speech nodes and their valid paths in the speech flow configuration interface, allowing users to configure speech nodes and speech flow paths visually, providing intuitive, flexible, and efficient visual configuration, and improving user efficiency in configuring and managing speech flows. Other optional features and technical effects of the embodiments of this disclosure are described in part below, and in part will be apparent from reading this document. Attached Figure Description

[0029] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings. The elements shown are not limited to the scale shown in the drawings, and the same or similar reference numerals in the drawings denote the same or similar elements, wherein:

[0030] Figure 1 A flowchart illustrating a speech flow configuration method according to an embodiment of the present disclosure is shown;

[0031] Figure 2 A schematic diagram of a speech flow configuration interface according to a specific embodiment of the present disclosure is shown;

[0032] Figure 3 A flowchart illustrating a speech flow configuration method according to an embodiment of the present disclosure is shown;

[0033] Figure 4 A flowchart illustrating a speech flow configuration method according to an embodiment of the present disclosure is shown;

[0034] Figure 5 A schematic diagram of a speech flow configuration interface according to a specific embodiment of the present disclosure is shown;

[0035] Figure 6 A flowchart illustrating a speech flow configuration method according to an embodiment of the present disclosure is shown;

[0036] Figure 7 A flowchart illustrating a speech flow configuration method according to an embodiment of the present disclosure is shown;

[0037] Figure 8 A schematic diagram of a speech flow configuration interface according to a specific embodiment of the present disclosure is shown;

[0038] Figure 9 A flowchart illustrating a speech flow configuration method according to an embodiment of the present disclosure is shown;

[0039] Figure 10 A flowchart illustrating a speech flow configuration method according to an embodiment of the present disclosure is shown;

[0040] Figure 11 A flowchart illustrating a speech flow configuration method according to an embodiment of the present disclosure is shown;

[0041] Figure 12 A schematic diagram of a speech flow configuration interface according to a specific embodiment of the present disclosure is shown;

[0042] Figure 13 A flowchart illustrating a conversation flow interaction method according to an embodiment of the present disclosure is shown;

[0043] Figure 14 A schematic diagram of a speech flow visualization configuration interface according to a specific embodiment of the present disclosure is shown.

[0044] Figure 15 A schematic diagram of the module composition of a speech flow configuration device according to an embodiment of the present disclosure is shown;

[0045] Figure 16 A schematic diagram of the module composition of a speech flow interaction device according to an embodiment of the present disclosure is shown;

[0046] Figure 17 A schematic diagram of the structure of an electronic device for implementing a speech flow configuration method according to an embodiment of the present disclosure is shown. Detailed Implementation

[0047] To make the objectives, technical solutions, and advantages of this disclosure clearer, the disclosure will be further described in detail below with reference to specific embodiments and accompanying drawings. Here, the illustrative embodiments and descriptions of this disclosure are used to explain this disclosure, but are not intended to limit this disclosure.

[0048] In this embodiment of the disclosure, a speech node refers to the smallest unit that makes up the speech speech flow / speech speech flow path. Each node can be configured with different speech content and specific functions, and the node attributes of each node are also different.

[0049] In the embodiments of this disclosure, upstream and downstream have conventional meanings in the field of data stream processing. Upstream refers to the input party in a data stream or logical stream, while downstream refers to the receiver in a data stream or logical stream. For example, in a data processing system, if data enters node A and is then transmitted from node A to node B, then node A is the upstream node of node B, and node B is the downstream node of node A.

[0050] Currently, AI voice chatbot technology has broad application prospects, mainly based on pre-configured scripts to achieve natural and fluent voice interaction. However, in practical applications, as business complexity continues to increase, the number and complexity of scripts required for AI voice chatbot configuration are also rapidly increasing. Consequently, a graphical script configuration interface has been introduced, allowing users to configure the AI ​​voice chatbot's scripts through visual operations such as dragging and dropping or connecting script nodes.

[0051] Nevertheless, existing graphical configuration technologies still have shortcomings when dealing with a large number of dialogue nodes. Users often struggle to quickly and efficiently configure a high-quality dialogue flow that meets their needs when faced with a large number of disorganized dialogue flow nodes in the configuration interface, especially when configuring more complex voice interaction dialogue flows. Furthermore, graphical dialogue flow configuration also lacks flexibility, hindering users from subsequent editing and optimization of complex dialogue flows.

[0052] In response, this disclosure provides a method for configuring a speech stream, which enables users to conveniently and efficiently configure a complex and high-quality speech stream from numerous speech stream nodes according to their needs.

[0053] In some embodiments, reference is made to Figure 1 It illustrates a flowchart of a speech flow configuration method for an AI speech flow configuration system / program / device according to embodiments of the present disclosure. In some embodiments, the AI ​​speech flow configuration system / program / device may be located in an electronic device, a virtual device, or as part of a software program, and the present disclosure does not limit this. Figure 1 The script flow configuration method may include at least 110-130.

[0054] 110: Generate a speech flow editing area that includes multiple speech nodes.

[0055] In some embodiments of this disclosure, the script node has node attributes and script content. In some embodiments of this disclosure, the script flow editing area includes a visual graphical configuration interface, in which multiple script nodes are placed, and the display form of the script nodes includes, but is not limited to, icons or geometric shapes.

[0056] like Figure 2 As shown in a specific embodiment of this disclosure, an AI voice script flow configuration system is provided, which includes a main window interface for visually configuring the script flow path. The main window interface includes a rectangular script flow editing area, in which multiple script nodes (shown as circles) are placed and the node attribute display interface of the corresponding script nodes are displayed.

[0057] In this embodiment of the disclosure, for example in 110 above, the user can perform various visual operations on the script nodes in the script flow editing area, including but not limited to adding new script nodes, selecting / deleting existing script nodes, and moving / dragging existing script nodes in the script flow editing area. In a specific example, the user... Figure 2 Right-click at position A in the script editing area shown, expand the right-click menu, and select the "Add Node" option. A node configuration interface will pop up, including various selectable node types and a script content configuration area. The user selects "Jump Node" as the node type and configures the script content as "About to be transferred to an agent" in the script content configuration area. After configuration, the user clicks the "OK" button, and the new node is successfully added and displayed in the script editing area at position A. However, it is understood that the above process of adding a new script node is merely an example and is not considered a limitation of this disclosure. It is also understood that this disclosure does not limit the specific shape of the script editing area, the number of script nodes placed within it, or their specific arrangement. Furthermore, the script editing area described in this embodiment may also be reasonably combined with other visual configuration components or visual areas, and this disclosure also does not limit this.

[0058] 120: Based on the preset node association rules, start from the target speech node in the speech flow editing area, and traverse upstream and / or downstream associated upstream speech nodes to form a speech flow path.

[0059] In this embodiment of the disclosure, the target speech node refers to the speech node that the user currently intends to operate on, including but not limited to the speech node that the user has currently added, selected, moved, or dragged.

[0060] In the embodiments disclosed herein, such as Figure 2As shown, each dialogue node has node attributes, including but not limited to start nodes, jump nodes, request nodes, reinforcement nodes, response nodes, and termination nodes. The node attributes of dialogue nodes can also be reset in the dialogue flow editing area. In one example, in the dialogue flow, the start node can serve as the beginning of the dialogue flow, for example, to trigger an initial greeting; the jump node can be used to switch to other nodes, for example, to jump to other corresponding nodes based on user input or context conditions; the request node can be used to request information or confirm operations from the user, for example, to ask the user for information or confirm operation steps; the reinforcement node can be used for secondary confirmation or data verification, for example, to secondary confirm user requests or check the validity of user input; the response node can be used to return information or respond to the user, for example, to provide the user with the required information; the termination node can serve as the end of the dialogue flow, for example, to end the current conversation and perform some cleanup operations or record a dialogue log. It is understood that the above description of node attributes for relevant uses is illustrative and not a limitation of this disclosure. Other embodiments may include more / fewer node attributes corresponding to different specific uses.

[0061] In this embodiment of the disclosure, for example in 120 above, the preset node association rules include, but are not limited to, node attribute association rules between speech nodes, logical association rules of speech content of speech nodes, or a combination of the node attribute association rules and the logical association rules of speech content, etc., wherein the preset node association rules can be pre-configured manually or automatically generated by the system.

[0062] In some embodiments of this disclosure, the node attribute association rules are a set of rules that determine how speech nodes are associated with each other based on the node attributes of each node. For example, a speech node with a certain node attribute is allowed to be associated with a speech node with a matching node attribute. In one example, the node attribute association rules include upstream node association rules and downstream node association rules. The upstream node association rules are used to determine which speech nodes with certain node attributes can be upstream nodes of the current node; the downstream node association rules are used to determine which speech nodes with certain node attributes can be downstream nodes of the current node's data. In a specific embodiment, for a speech node with the node attribute "reply node," the upstream node association rules determine that the allowed upstream nodes for association with this speech node can have node attributes of "request node" and "reinforcement node"; for a speech node with the node attribute "request node," the downstream node association rules determine that the allowed downstream nodes for association with this speech node can have node attributes of "reply node" or "jump node."

[0063] In some embodiments of this disclosure, the logical association rules for the dialogue content are a set of rules that determine how dialogue nodes are related to each other based on the logical connections between the dialogue content configured in the dialogue nodes. In one example, the logical connection is called a logical association relationship, which may include causal relationships, conditional relationships, and sequential relationships between the dialogue content configured in the dialogue nodes, or that the dialogue content configured in the dialogue nodes is applicable to the same or similar contexts. In a specific example, the dialogue content corresponding to a certain dialogue node is "Whether to transfer to human service", and based on the causal and sequential relationships of the dialogue content, it can be determined that the dialogue content of the downstream node that the dialogue node is allowed to associate with is "You will be transferred to human service soon" or "Transferring in progress"; in another specific example, the dialogue content corresponding to a certain dialogue node is "Invalid order number", and based on the conditional relationships of the dialogue content, it can be determined that the dialogue content of the upstream node that the dialogue node is allowed to associate with can be "Please enter your order number" or "Looking up order number".

[0064] In some embodiments of this disclosure, the logical association rules can be generated based on the user's edited dialogue flow history. Accordingly, prior to step 120 above, the dialogue flow configuration method may further include the following steps: a. Obtaining a dialogue flow history set, wherein the dialogue flow history set includes at least one historical dialogue flow path and dialogue node configuration records corresponding to the historical dialogue flow path. In some embodiments of this disclosure, the dialogue node configuration records include the attributes of the dialogue nodes, the association relationships between dialogue nodes, and the dialogue content contained in the dialogue nodes; these records reflect the logic confirmed by the user when configuring the dialogue flow in the past. b. Generating logical association rules based on the dialogue flow history set.

[0065] In some embodiments of this disclosure, based on the historical dialogue flow record set, the association between the node attributes and node content of dialogue nodes can be analyzed to form logical association rules. In a specific example, the following configuration appears multiple times in the valid dialogue flow path in the historical record set: Node A (node ​​attribute: request node, node content: Do you need more help?) frequently connects to Node B (node ​​attribute: response node, node content: Do you need human assistance?) and Node C (node ​​attribute: redirect node, node content: transfer to human assistance), and Node B and Node C frequently appear as downstream nodes of Node A. Based on the above data analysis, the logical association rule is obtained: for "node attribute: request node, node content: Do you need more help?", the allowed downstream nodes include "node attribute: response node, node content: Do you need human assistance?" or "node attribute: redirect node, node content: transfer to human assistance". In some examples, the number of times Node A connects to Node B can be set as needed, and this disclosure does not impose any restrictions.

[0066] In some embodiments of this disclosure, such as Figure 3 As shown, the above 120 may include at least 121 to 122.

[0067] 121: Starting from the target speech node, determine the associated upstream speech nodes in sequence according to the preset node association rules until there are no more associated upstream speech nodes. The associated upstream speech node determined each time is used as the starting point for the next determination.

[0068] In some embodiments of this disclosure, such as in 121 above, the preset node association rule can be applied first to identify upstream nodes directly associated with the target speech node. After this identification, the preset node association rule is applied again, and starting from the newly identified upstream node, the process continues to traverse and identify more upstream related nodes. This traversal and identification process is repeated, with each identified upstream node serving as the starting point for the next traversal and identification, until no new related upstream node can be found by applying the preset node association rule, thereby identifying all related upstream nodes.

[0069] 122: Starting from the target speech node, determine the associated downstream speech nodes sequentially according to the preset node association rules until there are no more associated downstream speech nodes. The associated downstream speech node determined each time is used as the starting point for the next determination.

[0070] Here, the description of 122 above can be referred to the description of 121 in the above embodiments. The same principle can be used for 122 above to identify the associated downstream nodes, and will not be repeated here. In the embodiments of this disclosure, 121 and 122 above can be executed in parallel or in any order, and these situations all fall within the protection scope of this disclosure.

[0071] The above S121 to S122 will be further described below with reference to a specific embodiment:

[0072] Continue to refer to Figure 2 In one specific embodiment, a process for determining associated upstream / downstream nodes based on target speech nodes is illustrated. Figure 2In the script editing area, there is a target script node X (shown as a black circle) and several other script nodes (shown as white circles). First, starting with the target script node X, upstream script nodes A and B, and downstream script nodes C and D directly associated with the target script node are determined according to preset node association rules. Specifically, based on node attribute association rules, the node attribute "jump node" of the target script node X matches the node attribute "request node" of script node A. Simultaneously, it is confirmed that the script content "request to transfer to human operator" corresponding to the target script node X and the script content "soon to transfer to human operator" of script node A have a logical relationship. Therefore, script node A can be identified as the upstream script node of the target script node X. Similarly, upstream script node B and downstream script nodes C and D directly associated with the target script node can be determined accordingly. Then, taking dialogue nodes A, B, C, and D as starting nodes, the preset node association rules are repeatedly applied to determine the upstream dialogue node A1 associated with node A, the upstream dialogue node B1 associated with node B, the downstream dialogue node C1 associated with node C, and the downstream dialogue node D1 associated with node D. Similarly, the preset node association rules are repeatedly applied to determine the upstream dialogue node A2 associated with node A1, the upstream dialogue node B2 associated with node B, the downstream dialogue node C2 associated with node C, and the downstream dialogue node D2 associated with node D. At this point, applying the preset node association rules again will not find any new upstream / downstream dialogue nodes associated with A2 or B2, nor will it find any downstream dialogue nodes associated with C2 or D2. The process of determining the associated upstream / downstream nodes ends here.

[0073] In other embodiments of this disclosure, relative position information between multiple speech nodes and the target speech node may be introduced to determine the upstream / downstream speech nodes associated with the target speech node. In the embodiments of this disclosure, the relative position information is interpreted broadly, including but not limited to relative position information such as the straight-line distance between multiple speech nodes and the target speech node.

[0074] At this time, as Figure 4 As shown, the speech flow configuration method disclosed herein may further include 410 to 440:

[0075] 410: Obtain the relative position information of the target speech node and multiple adjacent speech nodes.

[0076] In this embodiment of the disclosure, the relative position information may include the coordinate information of the target speech node and the plurality of adjacent speech nodes or the distance information between the plurality of adjacent speech nodes, wherein the adjacent means that the speech node is close in spatial location.

[0077] In one example, a proximity threshold can be set. When the straight-line distance between a certain speech node and a target speech node is less than the proximity threshold, the speech node is determined to be an adjacent speech node of the target speech node, that is, the speech node and the target speech node are spatially close, and the spatial location of the speech node is located in the vicinity of the target speech node. In a specific example, such as Figure 2 As shown, for the target speech node X, the nodes that are spatially adjacent to it are speech nodes A, B, C and D. At this time, the relative position information of the target speech node X and speech nodes A, B, C and D can be obtained.

[0078] 420: Determine the first upstream speech node and / or the first downstream speech node associated with the target speech node based on the node association rules and relative position information.

[0079] In some embodiments of this disclosure, such as in 420 above, the node association rules may include node attribute association rules between the speech nodes and logical association rules for the speech content.

[0080] In some embodiments of this disclosure, such as in 420 above, the first upstream and / or first downstream speech nodes directly associated with the target node are determined by combining relative position information. This ensures that during the traversal of associated nodes, the spatial location of the first confirmed associated node is located in the area surrounding the target node, which helps the user quickly find associated nodes near the target node. Furthermore, when the user moves the target speech node, the first upstream and / or first downstream speech nodes directly associated with the target node are calculated and updated in real time. By limiting the spatial location of the first confirmed upstream and downstream nodes to the area near the target node, computational resources are saved, and the association relationship of a cluster of speech nodes near the target speech node is quickly determined. This effectively avoids the situation where too many associated nodes are displayed near the target node when it is moved, leading to slow confirmation and a cluttered speech flow editing interface. This helps the user's decision-making and selection, improving the efficiency of speech flow configuration and user experience. In one example, when the target node is moved, the associated nodes near the target node can be calculated in real time, and the association relationship between the target node and the associated nodes is shown by connecting lines.

[0081] In this embodiment of the disclosure, for example in 420 above, the node association rules include node attribute association rules between nodes and logical association rules of speech content.

[0082] 430: Starting from the first upstream verbal node, determine the associated second upstream verbal nodes in sequence according to the node association rules, until there are no more associated second upstream verbal nodes. The associated second upstream verbal node determined each time is used as the starting point for the next determination.

[0083] 440: Starting from the first downstream dialogue node, determine the associated second downstream dialogue nodes in sequence according to the node association rules until there are no more associated second downstream dialogue nodes. The associated second downstream dialogue node determined each time is used as the starting point for the next determination.

[0084] In this embodiment of the disclosure, for example, in steps 430-440 above, unlike step 420 above, the node association rules and relative position information between speech nodes are no longer combined. Instead, similar to steps 121-122 above, the associated second upstream speech node and / or second downstream speech node are identified based on the node attribute association rules between speech nodes and the logical association rules of speech content. In this embodiment of the disclosure, steps 430 and 440 above can be executed in parallel or sequentially in any order, and these situations all fall within the protection scope of this disclosure. The method of steps 410-440 above will be further described below with reference to specific embodiments:

[0085] refer to Figure 5 In one specific embodiment, upstream / downstream nodes associated with the target speech node are shown, wherein there is a target speech node X (shown as a black circle) and multiple other speech nodes (shown as white circles) in the speech stream editing area.

[0086] At this point, the entire traversal confirmation process is divided into two parts: (1) confirming the first upstream and / or first downstream phonics node of the target phonics node X, and (2) confirming the second upstream and / or second downstream phonics node associated with it. Specifically:

[0087] (1) First, obtain the location information of the speech nodes near the target speech node X, including the location coordinates of the nearby speech nodes A, B, C, D, E and F. Then, calculate the distances d1, d2, d3, d4, d5 and d6 between each node and the target speech node X. The distance order is d6 > d4 > d5 > d2 > d3 > d1. Based on the node attribute association rules and speech content logical association rules between the target speech node X and speech nodes A, B, C, D, E and F, as well as the relative position information, identify the first upstream speech node from speech nodes A, B, C, D, E and F: For speech nodes A, B and E located upstream of the target speech node X, it is confirmed that the node attribute "request node" of speech node A is the first upstream speech node. The node attributes "reinforcement node" of the dialogue node B match the node attribute "jump node" of the target dialogue node X, while the node attributes of the dialogue node E and the target dialogue node X do not match (not shown), so the dialogue node E is excluded. Meanwhile, it is confirmed that the node content "request transfer to human operator" of the dialogue node A and the node content "confirm transfer" of the dialogue node B match the node content "soon to be transferred to human operator" of the target dialogue node X, while the node content of the dialogue node E and the target dialogue node X do not match (not shown). Furthermore, since the distances between the dialogue nodes A and B and the target dialogue node X are d1 and d2 respectively, and d2 > d1, the dialogue node A is confirmed as the first upstream dialogue node associated with the target dialogue node. Similarly, the dialogue node C can be confirmed as the first downstream dialogue node associated with the target dialogue node.

[0088] (2) Subsequently, the relative position information of the target speech node X with speech nodes A, B, C and D is no longer considered. Instead, based on the node attribute association rules and the logical association rules of speech content, the second upstream speech node A associated with the first upstream speech node A is determined. 1、 And the second downstream dialogue node C1 associated with the first downstream dialogue node C. Then, the node attribute association rules and the dialogue content logical association rules are repeatedly applied to determine the second upstream dialogue node A2 associated with node A1, and the second downstream dialogue node C2 associated with node C1. This process is repeated until no new associated upstream / downstream dialogue nodes can be found using these rules.

[0089] In other embodiments of this disclosure, such as in 420 above, the relative position information may further include a distance threshold between speech nodes. In this case, an upstream / downstream speech node whose distance from the target speech node satisfies the distance threshold and satisfies the logical association rules of the node attribute association rules and the speech content can be identified as the first upstream speech node / first downstream speech node.

[0090] Continue to refer to Figure 5 In one specific embodiment, the relative position information may further include a distance threshold R between speech nodes (see...). Figure 5 (A dashed circle with radius R is defined). At this point, the location information of nearby speech nodes X can be obtained, including the coordinates of nearby speech nodes A, B, C, D, E, and F. Then, the distances d1, d2, d3, d4, d5, and d6 between each node and the target speech node X are calculated, where d4 > R, d6 > R, and d1 < R, d2 < R, d3 < R, and d5 < R. Based on the node attribute association rules and logical association rules of speech content between the target speech node X and speech nodes A, B, C, D, E, and F, as well as the relative position information, the first upstream speech node is identified from speech nodes A, B, C, D, E, and F: For speech nodes located upstream of the target speech node X... Since the distance d6 to the target script node X is greater than the threshold R, script node E is excluded. Script nodes A and B are located at distances d1 and d2 from the target script node X, respectively, where d1 < R and d2 < R. Furthermore, the node attributes "request node" for script node A and "reinforcement node" for script node B match the node attribute "jump node" for the target script node X. The script content "soon to be transferred to a human operator" for script node A and "confirm transfer" for script node B have a logical relationship with the script content "request transfer to a human operator" for the target script node X. Therefore, script nodes A and B are confirmed as the first upstream script nodes associated with the target script node. Similarly, script nodes C and D are confirmed as the first downstream script nodes associated with the target script node.

[0091] Subsequently, the relative position information of the target speech node X and speech nodes A, B, C and D is no longer considered. The second upstream / downstream speech nodes associated with speech nodes A, B, C and D are determined only according to the node attribute association rules and the speech content logical association rules, until no new associated upstream / downstream speech nodes can be found by applying the node attribute association rules and the speech content logical association rules. This process will not be elaborated further here.

[0092] In this embodiment of the disclosure, after confirming several speech flow paths for target speech nodes, the speech flow paths will be further confirmed and filtered to determine the valid paths for the target speech nodes. The method of this embodiment of the disclosure may include the following steps:

[0093] 130: Determine whether the script flow path has a starting script node and an ending script node, and then determine the script flow path with a starting script node and an ending script node as the valid path corresponding to the target script node.

[0094] As previously stated, each speech node in this embodiment has a corresponding speech node attribute. Here, for example, in step 130 above, the starting speech node and the ending speech node are not necessarily the speech nodes at the two ends of the speech flow path. In this embodiment, the starting speech node refers to a speech node with the node attribute of a starting node, and the ending speech node refers to a speech node with the node attribute of an ending node. In other words, whether a speech node is considered a starting / ending speech node depends on whether its node attribute is a starting / ending node.

[0095] Therefore, in some embodiments of this disclosure, the above-mentioned 130 may include the following steps:

[0096] (1) Determine whether the node attribute of one of the two ends of the speech flow path is the starting node.

[0097] (2) Determine whether the node attribute of the other of the two ends of the speech flow path is a termination node.

[0098] In some embodiments of this disclosure, it can be determined whether the speech flow path is a valid path by judging whether the node attributes of the speech nodes at both ends of the determined speech flow path are speech nodes that are the start node and / or the end node, respectively.

[0099] The methods of steps (1) to (2) above will be described in detail below with reference to a specific embodiment:

[0100] Continue to refer to Figure 2 In one specific embodiment of this disclosure, there is a target speech node X (shown as a gray circle) and multiple other speech nodes (shown as white circles) in the speech flow editing area. There are multiple speech flow paths including the target speech node X (only two are shown). The speech nodes at both ends of the speech flow path A2→A1→A→X→C→C1→C2 are speech nodes A2 and C2. Speech node A2 is determined to be the starting speech node (its node attribute is starting node), and speech node C2 is the ending speech node (its node attribute is ending node). Therefore, the speech flow path A2→A1→A→X→C→C1→C2 is confirmed as a valid path for the target speech node X (connected by a solid line and shown). In contrast, the two ends of the dialogue flow path B2→B1→B→X→D→D1→D2 are dialogue nodes B2 and D2. After judgment, dialogue node B2 is not the starting dialogue node (its node attribute is jump node), and dialogue node D2 is not the ending dialogue node (its node attribute is reply node). Therefore, the dialogue flow path B2→B1→B→X→D→D1→D2 is not a valid path to the target dialogue node X (connected by dashed lines and shown).

[0101] In other embodiments of this disclosure, the speech stream editing area may further include a start area and an end area, and the node attributes of speech nodes existing in / moved into the start area / end area will be automatically changed to start node / end node.

[0102] Specifically, such as Figure 6 As shown, the method in this embodiment may further include steps 610 to 620:

[0103] 610: In response to the introduction of a speech node in the starting region, the node attribute of the introduced speech node is defined as the starting node.

[0104] 620: In response to the introduction of a speech node in the termination region, the node attribute of the introduced speech node is defined as a termination node.

[0105] Therefore, in some embodiments of this disclosure, it can be determined whether there is a starting speech node and / or an ending speech node in the speech flow path by determining whether there are nodes in the speech flow path located in the starting region and / or the ending region.

[0106] Specifically, such as Figure 7 As shown, the speech flow configuration method disclosed herein can be further described in sections 710-720: 710: Determine the node located in the starting region of the speech flow path as the starting node; 720: Determine the node located in the ending region of the speech flow path as the starting node.

[0107] In some embodiments of this disclosure, when both the starting region and the ending region have a speech node in a speech flow path, the speech node in the starting region of the speech flow path is determined as the starting node, and the speech node in the ending region is determined as the ending node. The speech flow path with the starting node and the ending node is determined as a valid path.

[0108] The method for combining the start / end regions in the above steps will be described in detail below with reference to a specific embodiment:

[0109] Comparison Reference Figure 2 In one specific embodiment of this disclosure, such as Figure 8 As shown, in the script editing area, there is a target script node X (shown as a gray circle) and several other script nodes (shown as white circles). There are multiple script flow paths including the target script node X (only two are shown). The script nodes at both ends of the script flow path B2→B1→B→X→D→D1→D2 are script nodes B2 and D2, respectively, unlike... Figure 2In the illustrated embodiment, at this time, the speech node B2 is located in the starting region, and its speech node attribute is therefore modified to the starting node, and thus it is determined as the starting speech node; the speech node D2 is located in the ending region, and its speech node attribute is therefore modified to the ending node, and thus it is determined as the ending speech node. Therefore, the speech flow path B2→B1→B→X→D→D1→D2 is also confirmed as a valid path for the target speech node X (connected by solid lines and shown).

[0110] In some embodiments of this disclosure, such as Figure 9 As shown, after determining the effective path corresponding to the target speech node, for example in step S130 above, the determined effective path of the target speech node can be displayed using various visualization schemes.

[0111] Accordingly, the speech flow configuration method in this embodiment may include steps 910-940:

[0112] Among them, 910 to 930 are similar to 110 to 130 mentioned above, and will not be described again here.

[0113] 940: Displays the valid path of the target speech node.

[0114] In some embodiments of this disclosure, such as in 940 above, the effective path of the determined target speech node can be displayed using various visualization schemes, including but not limited to using lines to connect the speech nodes in the effective path in an upstream-downstream order.

[0115] In some embodiments of this disclosure, the identified multiple valid paths may be connected by lines of different colors.

[0116] In other embodiments of this disclosure, the determined multiple valid paths can be arranged sequentially according to the number of speech nodes they contain, so as to facilitate user viewing.

[0117] In other embodiments of this disclosure, the 940 described above may further include:

[0118] Highlight the verbal nodes in the effective path.

[0119] In some embodiments of this disclosure, the highlighting includes, but is not limited to, adjusting the color, shape, etc. of the speech nodes in the effective path.

[0120] Furthermore, in other embodiments of this disclosure, to help users track the configuration progress of specific script content at each node in the script flow in a timely manner, such as... Figure 10 As shown, the speech flow configuration method may further include steps 1010 to 1050:

[0121] Among them, 1010 to 1030 are similar to 110 to 130 mentioned above, and will not be described again here.

[0122] 1040: Determine the completion rate of the speech content of all speech nodes in the valid path.

[0123] In this embodiment of the disclosure, the completion degree of the speech content of each speech node in the effective path can be determined. The completion degree of the speech content can be determined by comparing the number of speech contents that have been configured in the current node with the number of speech contents that have been preset.

[0124] In some embodiments of this disclosure, when a user starts configuring a script node, the user can configure the preset number of script contents for that script node. This number can be set according to the type of the script node and the complexity of the script flow to be achieved. For example, a request node may need to be configured with 5 script contents, while a jump node may only need to be configured with 2 script contents.

[0125] In some embodiments of this disclosure, the completion rate of the script content can be represented by the percentage of script content completion. In a specific example, the percentage of script content completion can be calculated by the following formula: Script content completion percentage = (Number of configured script content / Number of preset script content) × 100%.

[0126] 1050: Identify the completion rate of at least a portion of the dialogue content in the valid path. In some embodiments of this disclosure, such as in 160 above, the completion rate of the dialogue content can be identified in various ways, including but not limited to badges, marker boxes, etc., and can simply indicate the current completion status, such as "completed" or "pending," where "completed" corresponds to a dialogue content completion percentage of 100%, and "pending" corresponds to a dialogue content completion percentage of less than 100%. In other embodiments, a specific completion percentage can also be identified, and this application does not limit this.

[0127] Continue to refer to Figure 8 In one specific embodiment of this disclosure, the specific percentage completion rate of the speech nodes in the effective path is shown in the form of marked boxes (some are not shown), wherein the percentage completion rate of the speech content of nodes A, B, C and D is 50%, 25%, 75% and 100%, respectively.

[0128] Furthermore, in some embodiments of this disclosure, such as Figure 11 As shown, after determining multiple valid paths (more than or equal to 2) corresponding to the target speech node, cross-path editing can be performed on any two valid paths (the first valid path and the second valid path). Specifically, this can include paths 1110 to 1160.

[0129] Among them, 1110 to 1130 are similar to the aforementioned 110 to 130, and will not be described again here.

[0130] 1140: In response to a copy or cut operation of multiple consecutive speech nodes of the first valid path, determine the segment to be spliced ​​to the second valid path.

[0131] In this embodiment of the disclosure, for example in 1140 above, the road segment including multiple consecutive speech nodes that is copied or cut will be determined as the road segment to be spliced, and the association relationship between the multiple consecutive speech nodes in the road segment to be spliced ​​will be temporarily stored and copied or cut together.

[0132] 1150: In response to the paste or move operation of the road segment to be spliced, determine the neighboring speech nodes in the second valid path that are adjacent to the endpoint speech node of the road segment to be spliced.

[0133] In some embodiments of this disclosure, such as in 1150 above, proximity refers to the spatial proximity of the speech nodes. In one example, a user can paste or move the road segment to be spliced ​​obtained by copying or cutting in 1140 above. When the road segment to be spliced ​​is close to the second effective path, the relative distance between the speech nodes on the second effective path near the two endpoints of the road segment to be spliced ​​and the two endpoints will be calculated and determined. Optionally, the speech node in the second effective path that is closest to the endpoint speech node of the road segment to be spliced ​​will be determined as the nearest speech node.

[0134] In other embodiments of this disclosure, such as in 1150 above, when the road segment to be spliced ​​is close to the second effective path, the relative distance between the speech nodes on the second effective path near the two endpoints of the road segment to be spliced ​​and the two endpoints can also be calculated and determined, and multiple speech nodes in the second effective path whose distance from the endpoint speech nodes of the road segment to be spliced ​​meets a preset distance threshold are determined as nearby speech nodes. The preset distance threshold can be set by the user as needed, and this disclosure does not limit it.

[0135] 1160: Based on the preset node association rules, splice the endpoint speech nodes and adjacent speech nodes that have a relationship to form a third valid path.

[0136] In some embodiments of this disclosure, such as in 1160 above, it may be optionally determined whether there is an association between the two nearest neighboring dialogue nodes of the endpoint dialogue node of the segment to be spliced ​​in the second valid path and the endpoint dialogue node. If there is an association according to the node attribute association rules and the logical association rules of the dialogue content, the endpoint dialogue node is associated with the nearest dialogue node to form a third valid path. If there is no association, the formation of a third valid path is rejected.

[0137] In other embodiments of this disclosure, such as in step 1160 above, it can also be determined, based on node attribute association rules and logical association rules of dialogue content, whether there is an association relationship between multiple neighboring dialogue nodes whose distance to the endpoint dialogue node of the segment to be spliced ​​in the second valid path meets a distance threshold and the endpoint dialogue node. If one neighboring dialogue node has an association relationship, a third valid path is formed; if more than one neighboring dialogue node has an association relationship, a prompt is made and a third valid path is formed based on further user specification; if no neighboring dialogue node has an association relationship, the formation of a third valid path is rejected.

[0138] The method for forming the third effective path in the above steps will be described in detail below with reference to a specific embodiment:

[0139] In one specific embodiment of this disclosure, such as Figure 12 As shown, in the script editing area, there are two valid paths: Valid Path 1: B2→B1→B→X→C→C1→C2, and Valid Path 2: A2→A1→A→Y→D→D1→D2. In this case, the segment C→C1→C2 in Valid Path 1 is selected, copied, and moved to the lower right corner of the script editing area. The calculation shows that the relative distance between script node D and endpoint script node C is the smallest. Therefore, the nearest script node in the second valid path to the endpoint script node C of the segment to be spliced ​​is identified as script node D. At this point, according to the section... The node attribute association rules and the logical association rules of the dialogue content confirm that the node attribute "request node" of dialogue node D and the node attribute "jump node" of endpoint dialogue node C conform to the node attribute association rules, and the node dialogue content "request to transfer to human agent" of dialogue node D and the node attribute "soon to transfer to human agent" of endpoint dialogue node C conform to the logical association rules of the dialogue content. Therefore, the node dialogue node D and the endpoint dialogue node C are associated (shown as a double line connection) to form a new valid path 3: A2→A1→A→Y→D→C→C1→C2.

[0140] The steps described above in this embodiment of the present disclosure can help users quickly copy and apply configured script paths or their script node segments across paths from a determined script flow (valid path), thereby improving the development efficiency of script flow and also helping users improve the consistency and maintainability of script flow.

[0141] In other embodiments of this disclosure, the speech flow configuration method may further include the following steps:

[0142] Obtain the search query, and confirm the required script node from multiple script nodes in the script flow editing area based on the search query.

[0143] In some embodiments of this disclosure, users are supported in using search queries to quickly retrieve desired target dialogues. The search queries support, but are not limited to, keywords, intents, and knowledge base matching. Through the methods described above, users can quickly find the required effective path or the desired target dialogue node from a large number of dialogue nodes, facilitating more refined management of dialogue flow paths and dialogue nodes.

[0144] In other embodiments of this disclosure, the retrieval method described above can also be used as described in steps 1110-1160, thereby quickly locating the required script node / script segment and copying and pasting across paths. Through the method described above, users can easily find the target intent key path by searching keywords and apply it repeatedly, quickly splitting different application scenario versions of a mature and effective script path. In some examples, the script flow editing area provides a retrieval function for finding the target intent key path by keywords. For example, the script flow editing area provides a retrieval function button; by entering keywords in the retrieval function box, the target intent key path matching the keywords can be found from the configured key paths; or, for example, the retrieval function box is activated when a preset shortcut key is detected. In one example, if the script content of any script node in a certain path matches the keywords entered by the user, then one or more effective paths including that script node are the target intent key paths.

[0145] In other embodiments of this disclosure, historical speech stream configuration error information can also be obtained, and error associations in the speech stream path can be located based on this information. Corresponding suggestions can then be provided based on preset response strategies. Specifically, in some embodiments of this disclosure, user configuration error information is recorded during the user's speech stream configuration process, the specific location of the error is given, and preset error handling suggestions are provided. This method can effectively avoid simple logical errors and various unexpected situations that occur during speech configuration.

[0146] In other embodiments of this disclosure, the configured speech flow can be analyzed based on historical speech flow configuration records to recommend associative speech nodes or pre-designed speech templates, thereby improving design efficiency. Specifically, in some embodiments of this disclosure, templates for creating the next node with one click can be recommended and supported based on user's historical operating habits and context node relationships.

[0147] In other embodiments of this disclosure, the actual operation data of the existing script flow can be statistically analyzed to obtain script conversion indicators such as node hang-up rate, node arrival rate, knowledge base hit rate, and process arrival rate, thereby assisting users in further optimizing the script.

[0148] In this embodiment, the node hang-up rate refers to the frequency at which a user hangs up the call at a certain dialogue node, i.e., node hang-up rate = number of hang-ups / number of arrivals, where the number of hang-ups is the number of times a user hangs up the call at that node, and the number of arrivals is the total number of times a user arrives at that node in the dialogue stream. The node hang-up rate can help identify which nodes may have problems, such as unclear dialogue content or poor user experience, and then optimize them.

[0149] Node arrival rate refers to the frequency with which a user reaches a specific node in the call flow. It is calculated as: Node arrival rate = Number of arrivals / Number of connected cases. Here, the number of arrivals refers to the number of times a user reaches a specific node in the call flow, and the number of connected cases refers to the total number of times the call flow containing that node has been successfully completed. Node arrival rate helps assess the access frequency of a node in the call flow, determining its importance and traffic volume.

[0150] Knowledge base hit rate refers to the frequency with which a user's request or query for information within the dialogue flow successfully hits the knowledge base. Specifically, knowledge base hit rate = number of knowledge base hits / total number of knowledge base visits. The knowledge base includes predefined answers, instructions, and knowledge points. Nodes within the dialogue flow can query the knowledge base based on user input to provide relevant information or solutions. By statistically analyzing and optimizing the knowledge base hit rate, the effectiveness and coverage of the knowledge base can be improved, increasing the probability of receiving effective feedback from user queries.

[0151] Process arrival rate refers to the frequency with which a user reaches a specific process within the conversation flow. It is calculated as: Process arrival rate = Number of calls reaching that process / Total number of connected calls. Here, "process" refers to a step or stage in the conversation flow path, and each step or stage must include at least two nodes. Process arrival rate helps in analyzing the path configuration of the conversation flow.

[0152] In one specific embodiment, there are three script flows 1, 2, and 3. Script flow 1 is A→B→C→D→E, script flow 2 is A→B→C→D→E→F, and script flow 3 is A→B→D→E→F. Each script flow runs twice, meaning the three script flows run a total of six times. During the execution of script flow 1, one user hangs up at node B, the script flow runs completely once, and node C hits the knowledge base (i) once. During the execution of script flow 2, two users hang up at node B, the script flow runs completely zero times, and the knowledge base is not hit. During the execution of script flow 3, all users hang up at node F, the script flow runs completely twice, and node D hits the knowledge base (j) twice. Then the hang-up rate of node B is 3 / 6 = 50%, the arrival rate of node B is 3 / 6 = 50%, the hit rate of knowledge base (i) is 1 / 2 = 16.7%, the hit rate of knowledge base (j) is 2 / 2 = 100%, the arrival rate of process C→D in script flow 1 is 1 / 2 = 50%, the arrival rate of process C→D→E in script flow 2 is 0 / 2 = 0%, and the arrival rate of process D→E in script flow 3 is 2 / 2 = 100%.

[0153] This disclosure provides a method for configuring a speech flow. It provides a speech flow editing area with multiple speech nodes. Based on preset node association rules, it starts from the target speech node in the speech flow editing area and traverses upstream and / or downstream associated speech nodes to form a speech flow path. It determines whether the speech flow path has a starting speech node and an ending speech node, thus identifying the speech flow path with both a starting and ending speech node as the valid path corresponding to the target speech node. This method facilitates efficient confirmation of the valid path for the target speech node and graphically displays the upstream and downstream relationships between speech nodes in the valid path. This allows users to easily and efficiently configure a high-quality speech flow that meets their needs from numerous speech flow nodes. It also provides further functions for cross-path copying and pasting and splicing of speech flows. Combined with speech node completion reminders and node data analysis functions, it further increases the stability and flexibility for users configuring complex voice interaction speech flows.

[0154] In some embodiments of this disclosure, reference is made to Figure 13 The diagram illustrates a flowchart of a speech flow interaction method for an AI voice speech flow configuration system according to an embodiment of the present application, wherein the speech flow interaction method may include at least steps 1310-1320.

[0155] 1310: Displays a speech flow editing area containing multiple speech nodes, where each speech node has node attributes and speech content.

[0156] In this embodiment of the disclosure, the script editing area includes a visual graphical configuration interface, which displays multiple script nodes. The display format of these script nodes includes, but is not limited to, icons or geometric shapes. In a specific example, such as... Figure 14 As shown, it illustrates the main window interface for visually configuring the speech flow path in the AI ​​voice speech flow configuration system. The main window interface includes a rectangular speech flow editing area, in which multiple speech nodes (shown as circles) are placed and the corresponding node attribute display interface is displayed.

[0157] In the embodiments of this disclosure, the plurality of dialogue nodes have node attributes and dialogue content. The node attributes include, but are not limited to, start nodes, jump nodes, request nodes, reinforcement nodes, reply nodes, and termination nodes. Further descriptions of node attributes and dialogue content in the above embodiments of this disclosure can be found in the relevant descriptions in sections 110-120 of this disclosure, and will not be repeated here.

[0158] In this embodiment of the disclosure, for example in 1310 above, in each valid speech flow path, adjacent speech nodes conform to preset node association rules, wherein the adjacent speech nodes refer to nodes that are spatially close in the speech flow editing area. In one example, the preset node association rules include node attribute association rules and logical association rules for speech content.

[0159] In this embodiment of the disclosure, each valid speech flow path includes a starting speech node and an ending speech node, wherein the node attribute of the starting speech node is a starting node and the node attribute of the ending speech node is an ending node.

[0160] In this embodiment of the disclosure, adjacent nodes are connected by directed node lines in each valid speech flow path. In this embodiment of the disclosure, the directed node lines clearly display the connection direction in the interface, for example, represented by arrows. At the same time, the directed lines also represent the upstream and downstream relationships between the connected speech nodes.

[0161] In a specific example, such as Figure 14 As shown, in the script editing area, script node A2 and script node A1 are connected by a directed line. This directed line is displayed on the interface as an arrow from A2 to A1, indicating that in the script flow, script node A2 is the upstream node of script node A1, which is the transmitter of data or information; while script node A1 is the downstream node of script node A2, which is the receiver of data or information.

[0162] In a specific example disclosed herein, such as Figure 14As shown, it illustrates the main window interface for visually configuring the speech flow path in the AI ​​voice speech flow configuration system. It displays multiple speech nodes arranged in circles, along with their corresponding node attribute display interfaces (only a portion is shown). This includes two valid speech paths: Valid Speech Path 1: A2→A1→A→X→D→D1→D2, and Valid Speech Path 2: B2→B1→B→C→C1→C2. Adjacent speech nodes conform to node attribute association rules and the logical association rules of speech content, and are shown as directed arrows. The node attributes of the starting speech nodes A2 and B2 are the starting nodes, and the node attributes of the ending speech nodes D2 and C2 are the starting nodes.

[0163] Further descriptions of the preset node association rules and node attributes in the above embodiments of this disclosure can be found in the relevant descriptions in 110-120 of this disclosure, and will not be repeated here.

[0164] 1320: In response to the user selecting the target utterance node, display at least one valid utterance flow path containing the target utterance node.

[0165] In this embodiment of the disclosure, in response to a user's selection of a speech node, such as clicking or selecting by box, the speech node selected by the user is confirmed as the target speech node, and at least one valid speech flow path containing the target speech node will be displayed. In a specific example, such as Figure 14 As shown, when the user selects the dialogue node B2, the valid dialogue flow path B2→B1→B→C→C1→C2 (indicated by a solid black line) containing the dialogue node B2 is displayed, while other valid dialogue flow paths are not displayed, such as the valid dialogue flow path A2→A1→A→X→D→D1→D2 (indicated by a dashed line).

[0166] In yet another specific embodiment, reference is made to the foregoing. Figure 8 When the user selects a speech node X, the valid speech flow paths B2→B1→B→X→D→D1→D2 and A2→A1→A→X→C→C1→C2 containing speech node X are displayed (shown as solid black lines).

[0167] In some embodiments of this disclosure, the speech stream interaction method may further include 1330 (not shown):

[0168] 1330: In response to the user selecting a node connection in a valid speech flow path, highlight the valid speech flow path corresponding to the selected node connection.

[0169] In one example, the highlighting may include changing the color of the node connection, increasing the brightness of the node connection, or thickening the node connection, etc., and this application does not limit this. In a specific example, refer to... Figure 2 When the user selects the node connection A2→A1, the valid speech flow path A2→A1→A→X→C→C1→C2 corresponding to the node connection A2→A1 is highlighted (bold).

[0170] In this embodiment of the disclosure, the above-described 1330 can be combined with the methods described in the aforementioned 940 in a non-contradictory manner.

[0171] In some embodiments of this disclosure, the speech stream editing area includes a start area and an end area, and the speech stream interaction method may further include 1340 (not shown):

[0172] 1340: In response to a user dragging a script node into the starting area, define the node attribute of the script node as a starting node; and / or, in response to a user dragging a script node into the ending area, define the node attribute of the script node as an ending node.

[0173] For a detailed description of the above embodiment 1340 of this disclosure, please refer to the relevant descriptions in 110-120 of this disclosure, which will not be repeated here.

[0174] In some embodiments of this disclosure, the effective path includes a first effective path and a second effective path. After determining the effective path corresponding to the target speech node, the speech flow configuration method may further include steps 1350-1370 (not shown):

[0175] 1350: In response to a copy or cut operation on multiple consecutive speech nodes of the first valid path, determine the segment to be spliced ​​to the second valid path;

[0176] 1360: In response to the paste or move operation of the road segment to be spliced, determine the neighboring speech node in the second valid path that is adjacent to the endpoint speech node of the road segment to be spliced;

[0177] 1370: According to the preset node association rules, the endpoint speech nodes and the adjacent speech nodes that have an association relationship are spliced ​​together to form a third valid path.

[0178] For a detailed description of the embodiments 1350-1370 of this disclosure, please refer to the relevant descriptions in 1140-1160 of this disclosure, which will not be repeated here.

[0179] In some embodiments of this disclosure, the speech flow interaction method may further include:

[0180] Displays a node library including multiple candidate script nodes; in response to a user dragging at least one of the candidate script nodes from the node library into the script flow editing area, the at least one candidate script node is displayed in the script flow editing area.

[0181] In this embodiment of the disclosure, the alternative dialogue nodes also have node attributes and node content, and can be configured by the user. In a specific embodiment, refer to... Figure 14 It shows the main window interface for visually configuring the speech flow path. The main window interface includes a rectangular speech flow editing area 1401, in which multiple speech nodes (shown as circles) are placed and the corresponding speech node's node attribute display interface is displayed, as well as a node library 1402 including multiple candidate speech nodes. The speech node Y can be dragged out by the user from the node library 1402 and displayed in the speech flow editing area 1401.

[0182] In some embodiments of this disclosure, users can also drag and drop nodes from the script editing area into the node library to become candidate script nodes. In this embodiment, when a script node is dragged between the script editing area and the node library, its node attributes and node content may optionally remain unchanged.

[0183] This disclosure also provides a method for interactive speech flow, which displays a speech flow editing area including multiple speech nodes; in response to a user selecting a target speech node, it displays at least one valid speech flow path containing the target speech node. This method can display speech nodes and their valid paths in the speech flow configuration interface, allowing users to configure speech nodes and speech flow paths visually, providing intuitive, flexible, and efficient visual configuration, and improving user efficiency in configuring and managing speech flows.

[0184] The speech flow interaction method and its sub-steps in the above embodiments of this disclosure can be combined with the speech flow configuration method and its sub-steps of the above disclosure in a non-contradictory manner to form new embodiments, which fall within the scope of this application.

[0185] This disclosure also provides a speech flow configuration device. In such... Figure 15 In the illustrated embodiment, the script flow configuration device 1500 may include a generation module 1501, a traversal module 1502, and a determination module 1503. The window module 1501 may be configured to generate a script flow editing area including multiple script nodes. The traversal module 1502 may be configured to, according to preset node association rules, traverse upstream and / or downstream associated upstream and / or downstream script nodes in the script flow editing area to form a script flow path. The determination module 1503 may be configured to, when the script flow path has a starting script node and an ending script node, determine the script flow path as a valid path corresponding to the target script node, wherein the node attribute of the starting script node is a starting node, and the node attribute of the ending script node is an ending node.

[0186] In some embodiments of this disclosure, the speech flow configuration device 1500 may further include:

[0187] (1) The node completion module is configured to calculate and display the node completion of the utterance node.

[0188] (2) The node cross-path copy and paste module is configured to copy and paste across paths and splice the speech segments.

[0189] (3) The node cross-path retrieval module is configured to quickly retrieve the required speech nodes based on the retrieval formula.

[0190] (4) Node connection color module, which is configured to adjust the color of the connection between nodes in the speech flow.

[0191] (5) The node connection editing module is configured to display the connections between all speech nodes and edit the attributes of the connections, including connection name, connection type, connection direction, etc.

[0192] (6) The node optimization tracking module is configured to provide error correlation tracking.

[0193] (7) The node recommendation optimization module is configured to recommend script templates based on the user's historical script flow configuration data.

[0194] (8) Node data analysis module, which is configured to statistically analyze effective paths and application data indicators of each node.

[0195] This disclosure also provides a speech flow interaction device 1600, including: a display module 1601 and a response module 1602. Wherein:

[0196] The display module 1601 is configured to display a speech flow editing area including multiple speech nodes, wherein the speech nodes have node attributes and speech content.

[0197] The response module 1602 is configured to, in response to a user selecting a target speech node, display at least one valid speech flow path containing the target speech node in the display module 1601; wherein, in each valid speech flow path, adjacent nodes conform to a preset node association rule, and the adjacent nodes are connected by directed node lines; wherein, each valid speech flow path includes a starting speech node and an ending speech node, the node attribute of the starting speech node is a starting node, and the node attribute of the ending speech node is an ending node.

[0198] In some embodiments, the response module 1602 is configured to highlight the valid speech flow path corresponding to the selected node connection in response to the user selecting a node connection of a valid speech flow path.

[0199] In some embodiments, the speech stream editing area includes a start area and an end area, and the response module 1602 is configured to:

[0200] In response to a user dragging a script node into the starting area, the node attribute of the script node is defined as the starting node; and / or

[0201] In response to a user dragging a script node into the termination area, the node attribute of the script node is defined as a termination node.

[0202] In some embodiments, the valid path includes a first valid path and a second valid path, and the response module 1602 is configured to:

[0203] In response to a copy or cut operation on multiple consecutive speech nodes of the first valid path, determine the segment to be spliced ​​to the second valid path;

[0204] In response to the pasting or moving operation of the road segment to be spliced, determine the neighboring speech nodes in the second valid path that are adjacent to the endpoint speech node of the road segment to be spliced;

[0205] According to the preset node association rules, the endpoint speech nodes and the adjacent speech nodes that have a relationship are spliced ​​together to form a third effective path.

[0206] In some embodiments, the display module 1601 is configured to display a node library including a plurality of alternative script nodes; the response module 1602 is configured to display the at least one alternative script node in the script flow editing area in response to a user dragging at least one of the alternative script nodes from the node library into the script flow editing area.

[0207] In embodiments of this disclosure, an electronic device is provided, including: a processor and a memory storing a computer program, the processor being configured to perform the method of any of the embodiments of this disclosure when running the computer program.

[0208] Figure 17 The illustration shows a method or electronic device 1700 that can implement embodiments of the present disclosure. In some embodiments, it may include more or fewer electronic devices than illustrated. In some embodiments, it may be implemented using a single or multiple electronic devices. In some embodiments, it may be implemented using cloud-based or distributed electronic devices.

[0209] like Figure 17As shown, the electronic device 1700 includes a central processing unit (CPU) 1701, which can perform various appropriate operations and processes based on programs and / or data stored in read-only memory (ROM) 1702 or programs and / or data loaded from storage portion 1708 into random access memory (RAM) 1703. The CPU 1701 can be a multi-core processor or may contain multiple processors. In some embodiments, the CPU 1701 may include a general-purpose main processor and one or more special coprocessors, such as a graphics processing unit (GPU), a neural network processor (NPU), a digital signal processor (DSP), etc. Various programs and data required for the operation of the electronic device 1700 are also stored in RAM 1703. The CPU 1701, ROM 1702, and RAM 1703 are interconnected via bus 1704. An input / output (I / O) interface 1705 is also connected to bus 1704.

[0210] The processor and memory described above are used together to execute a program stored in the memory. When the program is executed by a computer, it can implement the steps or functions of the methods described in the above embodiments.

[0211] The following components are connected to I / O interface 1705: an input section 1706 including a keyboard, mouse, touchscreen, etc.; an output section 1707 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1708 including a hard disk, etc.; and a communication section 1709 including a network interface card such as a LAN card, modem, etc. The communication section 1709 performs communication processing via a network such as the Internet. Drive 1710 is also connected to I / O interface 1705 as needed. Removable media 1711, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 1710 as needed so that computer programs read from them can be installed into storage section 1708 as needed. Figure 17 The diagram only shows a portion of the components and does not imply that the computer system 1700 only includes... Figure 17 The components shown.

[0212] In some embodiments, the electronic device 1700 refers to a mobile terminal or computer, including mobile phones, vehicle terminals, smart TVs, etc. Taking a mobile phone as an example, the electronic device 1700 also includes a touch screen, external speaker, gyroscope, camera, 4G / 5G antenna, and other device modules.

[0213] The systems, devices, modules, or units described in the above embodiments can be implemented by a computer or its associated components. The computer may be, for example, a mobile terminal, smartphone, personal computer, laptop computer, in-vehicle human-machine interface device, personal digital assistant, media player, navigation device, game console, tablet computer, wearable device, smart TV, Internet of Things system, smart home, industrial computer, server, or a combination thereof.

[0214] Although not shown, in embodiments of this disclosure, a program product is provided, the program product comprising a computer program that, when executed by a processor, implements the methods of any embodiment of this disclosure.

[0215] Although not shown, in embodiments of this disclosure, a storage medium is provided that stores a computer program, which, when executed, implements the method of any embodiment of this disclosure.

[0216] The storage media in embodiments of this disclosure include articles that are permanent and non-permanent, removable and non-removable, capable of storing information by any method or technology. Examples of storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0217] The methods, programs, systems, apparatuses, etc., of the embodiments of this disclosure can be executed or implemented in a single or multiple networked computers, or practiced in a distributed computing environment. In the embodiments of this specification, in these distributed computing environments, tasks can be performed by remote processing devices connected via a communication network.

[0218] Those skilled in the art will understand that the embodiments described in this specification can be provided as methods, systems, or computer program products. Therefore, those skilled in the art will realize that the functional modules / units or controllers and related method steps described in the above embodiments can be implemented in software, hardware, or a combination of both.

[0219] Unless explicitly stated otherwise, the actions or steps of the methods or procedures described in the embodiments of this disclosure do not necessarily have to be performed in a specific order and can still achieve the desired results. In some implementations, multitasking and parallel processing are also possible or may be advantageous.

[0220] This document describes several embodiments of the present disclosure; however, for the sake of brevity, the descriptions of the embodiments are not exhaustive, and identical or similar features or portions between the embodiments may be omitted. In this document, "one embodiment," "some embodiments," "example," "specific example," or "some examples" refers to at least one embodiment or example applicable to the present disclosure, but not all embodiments. The above terms do not necessarily refer to the same embodiment or example. Without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of the different embodiments or examples.

[0221] The exemplary systems and methods of this disclosure have been specifically shown and described with reference to the foregoing embodiments, and are merely examples of the best mode for implementing the systems and methods. Those skilled in the art will understand that various changes can be made to the embodiments of the systems and methods described herein without departing from the spirit and scope of this disclosure as defined in the appended claims when implementing the systems and / or methods.

Claims

1. A method for configuring a speech flow, characterized in that, The speech stream configuration method includes: Generate a speech flow editing area that includes multiple speech nodes, wherein each speech node has node attributes and speech content; According to the preset node association rules, starting from the target speech node in the speech flow editing area, the associated upstream speech nodes and / or downstream speech nodes are traversed upstream and / or downstream to form a speech flow path. If the speech flow path has a starting speech node and an ending speech node, then the speech flow path is determined as the valid path corresponding to the target speech node, wherein the node attribute of the starting speech node is the starting node, and the node attribute of the ending speech node is the ending node.

2. The speech flow configuration method according to claim 1, characterized in that, The step of forming a speech flow path by traversing upstream and / or downstream associated upstream and / or downstream speech nodes from the target speech node in the speech flow editing area according to preset node association rules includes: Starting from the target speech node, the associated upstream speech nodes are determined sequentially according to the preset node association rules until there are no more associated upstream speech nodes. Each time the associated upstream speech node is determined, it serves as the starting point for the next determination. and / or Starting from the target speech node, related downstream speech nodes are sequentially determined according to the preset node association rules until there are no more related downstream speech nodes. Each determined related downstream speech node serves as the starting point for the next determination.

3. The speech flow configuration method according to claim 1, characterized in that, The step of forming a speech flow path by traversing upstream and / or downstream associated upstream and / or downstream speech nodes from the target speech node in the speech flow editing area according to preset node association rules includes: Obtain the relative position information of the target speech node and multiple adjacent speech nodes; Based on the node association rules and the relative position information, determine the first upstream speech node and / or the first downstream speech node associated with the target speech node.

4. The speech flow configuration method according to claim 3, characterized in that, The step of forming a speech flow path by traversing associated upstream and / or downstream speech nodes from the target speech node in the speech flow editing area according to preset node association rules also includes: Starting from the first upstream dialogue node, related second upstream dialogue nodes are sequentially determined according to the node association rules until there are no more related second upstream dialogue nodes. Each time a related second upstream dialogue node is determined, it serves as the starting point for the next determination. and / or Starting from the first downstream dialogue node, related second downstream dialogue nodes are sequentially determined according to the node association rules until there are no more related second downstream dialogue nodes. Each time a related second downstream dialogue node is determined, it serves as the starting point for the next determination.

5. The script flow configuration method according to any one of claims 1 to 4, characterized in that, The node association rules include node attribute association rules between nodes and logical association rules for the content of the speech.

6. The speech flow configuration method according to claim 1, characterized in that, The script editing area includes a start area and an end area, and the script configuration method further includes: In response to the introduction of a speech node in the starting region, the node attribute of the speech node is defined as the starting node; In response to the introduction of a speech node in the termination region, the node attribute of the speech node is defined as a termination node.

7. The speech flow configuration method according to claim 1, characterized in that, After determining the valid path corresponding to the target speech node, the speech flow configuration method further includes: displaying the valid path of the target speech node.

8. The speech flow configuration method according to claim 1, characterized in that, After determining the valid path corresponding to the target speech node, the speech flow configuration method further includes: Determine the completion rate of the speech content of all speech nodes in the effective path; Identify the completion rate of the speech content of at least a portion of the speech nodes in the valid path.

9. The speech flow configuration method according to claim 1, characterized in that, The valid path includes a first valid path and a second valid path. After determining the valid path corresponding to the target speech node, the speech flow configuration method further includes: In response to a copy or cut operation on multiple consecutive speech nodes of the first valid path, determine the segment to be spliced ​​to the second valid path; In response to the pasting or moving operation of the road segment to be spliced, determine the neighboring speech nodes in the second valid path that are adjacent to the endpoint speech node of the road segment to be spliced; According to the preset node association rules, the endpoint speech nodes and the adjacent speech nodes that have a relationship are spliced ​​together to form a third effective path.

10. A speech flow interaction method, characterized in that, include: Displays a speech flow editing area that includes multiple speech nodes, wherein each speech node has node attributes and speech content; In response to a user selecting a target speech node, at least one valid speech flow path containing the target speech node is displayed; In each valid speech flow path, adjacent nodes conform to preset node association rules, and the adjacent nodes are connected by directed node lines. Each valid speech flow path includes a starting speech node and an ending speech node. The starting speech node has the attribute of a starting node, and the ending speech node has the attribute of an ending node.

11. The speech flow interaction method according to claim 10, characterized in that, Also includes: In response to the user selecting a node connection in a valid speech flow path, the valid speech flow path corresponding to the selected node connection is highlighted.

12. The speech flow interaction method according to claim 10, characterized in that, The script editing area includes a start area and an end area. The method further includes: In response to a user dragging a script node into the starting area, the node attribute of the script node is defined as the starting node; and / or In response to a user dragging a script node into the termination area, the node attribute of the script node is defined as a termination node.

13. The speech flow interaction method according to claim 10, characterized in that, The valid path includes a first valid path and a second valid path. After determining the valid path corresponding to the target speech node, the speech flow configuration method further includes: In response to a copy or cut operation on multiple consecutive speech nodes of the first valid path, determine the segment to be spliced ​​to the second valid path; In response to the pasting or moving operation of the road segment to be spliced, determine the neighboring speech nodes in the second valid path that are adjacent to the endpoint speech node of the road segment to be spliced; According to the preset node association rules, the endpoint speech nodes and the adjacent speech nodes that have a relationship are spliced ​​together to form a third effective path.

14. The speech flow interaction method according to claim 9, characterized in that, Also includes: Displays a node library that includes multiple alternative dialogue nodes; In response to a user dragging at least one of the candidate script nodes from the node library into the script flow editing area, the at least one candidate script node is displayed in the script flow editing area.

15. A speech flow configuration device, characterized in that, include: The generation module is configured to generate a speech flow editing area that includes multiple speech nodes. The traversal module is configured to traverse the associated upstream and / or downstream dialogue nodes from the target dialogue node in the dialogue flow editing area according to the preset node association rules, thereby forming a dialogue flow path. The determination module is configured to determine the dialogue flow path as the valid path corresponding to the target dialogue node when there is a start dialogue node and an end dialogue node in the dialogue flow path, wherein the node attribute of the start dialogue node is a start node and the node attribute of the end dialogue node is an end node.

16. A speech flow interactive device, characterized in that, include: The display module is configured to display a script flow editing area that includes multiple script nodes, wherein the script nodes have node attributes and script content; The response module is configured to display at least one valid speech flow path containing the target speech node in response to a user selecting the target speech node; wherein, in each valid speech flow path, adjacent nodes conform to a preset node association rule, and the adjacent nodes are connected by directed node lines; wherein, each valid speech flow path includes a starting speech node and an ending speech node, the node attribute of the starting speech node is a starting node, and the node attribute of the ending speech node is an ending node.

17. An electronic device, characterized in that, include: A processor and a memory storing a computer program, the processor being configured to implement the method of any one of claims 1-14 when executing the computer program.

18. A program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-14.