Robot speech detection visualization method, device, electronic device and storage medium

By constructing a robot speech test flowchart and identifying difference nodes, the problem of non-intuitive robot speech detection is solved, and the visualization effect of detection is improved. It is suitable for financial automatic services and online medical self-service.

CN115525750BActive Publication Date: 2025-09-26ONE CONNECT SMART TECH CO LTD SHENZHEN
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211254123.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-13
Publication Date
2025-09-26
Estimated Expiration
2042-10-13

AI Technical Summary

Technical Problem

The existing robot speech detection method is black box detection. The testers cannot intuitively view the detection logic, resulting in the inability to effectively modify the speech system.

Method used

By extracting simulated user features, screening historical human-computer dialogue records as test reference samples, constructing a speech test flowchart, and identifying the different nodes from the reference flowchart, the differences are highlighted using preset rules.

Benefits of technology

It realizes the visualization of robot speech detection, improves the intuitiveness and modifiability of detection, and is suitable for financial automatic services and online medical self-service scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115525750B_ABST
    Figure CN115525750B_ABST
Patent Text Reader

Abstract

The present invention relates to artificial intelligence technology and discloses a method for visualizing robot speech detection, comprising: using historical human-computer conversation records corresponding to users whose simulated user characteristics meet a preset role distance threshold as test reference samples, using user input information in the test reference samples as test corpus, receiving simulated reply speech given by a preset robot for the test corpus, constructing a speech test flowchart for the preset robot based on the test corpus and the simulated reply speech, constructing a speech reference flowchart for the preset robot based on the test reference sample, identifying the difference nodes between the speech test flowchart and the speech reference flowchart, and highlighting the corresponding difference nodes according to preset rules. The present invention also proposes a device, equipment, and medium for visualizing robot speech detection. The present invention can solve the problem of unintuitive robot speech detection in scenarios such as automatic financial services and online medical self-service.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method, device, electronic device and computer-readable storage medium for detecting robot speech visualization. Background Art

[0002] Human-computer interaction is a common application in the current field of artificial intelligence. For example, in self-service financial services and online healthcare, conversational robots can quickly and accurately analyze user needs based on the content of the conversation and respond with pre-designed responses or guidance to meet user needs.

[0003] With the advancement of technology, the application scenarios of conversational robots are becoming increasingly complex, and the corresponding conversational robot scripts are becoming increasingly sophisticated and extensive. For example, as business application scenarios upgrade, conversational robot scripts also need to be upgraded. However, complex script systems may experience inconsistencies during the upgrade process. Therefore, pre-testing conversational robot scripts is becoming increasingly important.

[0004] Currently, it is more common to use deep learning-based language models to detect related robot speech. For the tester, only the input and output are visible in this detection method. The detection process is like a black box and cannot be viewed intuitively. Therefore, the tester cannot intuitively understand the detection logic and cannot make relevant modifications to the subsequent robot speech. Summary of the Invention

[0005] The present invention provides a method, device, electronic device and computer-readable storage medium for visualizing robot speech detection, the main purpose of which is to solve the problem of non-intuitive robot speech detection.

[0006] To achieve the above objectives, the present invention provides a method for visualizing robot speech detection, comprising:

[0007] Extracting a preset simulated user's simulated user characteristics, and selecting, from a preset human-computer dialogue library, historical human-computer dialogue records corresponding to users whose characteristics meet a preset role distance threshold with the simulated user characteristics as test reference samples;

[0008] Using the user input information in the test reference sample as test corpus, receiving a simulated response speech given by a preset robot for the test corpus, and constructing a speech test flow chart for the preset robot based on the test corpus and the simulated response speech;

[0009] Constructing a speech reference flow chart of the preset robot according to the test reference sample;

[0010] Identify the difference nodes between the speech test flowchart and the speech reference flowchart, and highlight the corresponding difference nodes according to preset rules.

[0011] Optionally, selecting, from a preset human-computer dialogue library, historical human-computer dialogue records corresponding to users whose characteristics with the simulated user meet a preset role distance threshold as test reference samples includes:

[0012] Obtaining a user portrait set in the preset human-computer dialogue library;

[0013] Calculating the role distance between the simulated user feature and each user portrait in the user portrait set;

[0014] The human-computer dialogue record corresponding to the user portrait whose role distance meets the preset role distance threshold is selected as the test reference sample.

[0015] Optionally, constructing a speech test flowchart for the preset robot based on the test corpus and the simulated reply speech includes:

[0016] Performing user intent recognition on the test corpus, and generating corresponding user intent nodes according to the recognized user intent;

[0017] Identifying the business intent of the simulated reply speech, and generating a corresponding business intent node according to the identified business intent;

[0018] Arrange all the user intent nodes and all the business intent nodes vertically according to the chronological order in which the test corpus and the simulated reply script appear in the question and answer of the preset robot to obtain an intent node queue;

[0019] Calculate the intention distance between every two adjacent intention nodes in the intention node queue, and generate a connecting line of preset proportional length according to the size of the intention distance, and use the connecting line to connect the adjacent intention nodes in series to obtain the speech test flowchart.

[0020] Optionally, the performing user intent recognition on the test corpus includes:

[0021] Generate a text vector matrix of the test corpus;

[0022] Extracting text features of the test corpus from the text vector matrix;

[0023] Calculate the probability value between the text feature and the preset user intention label using a pre-trained activation function;

[0024] The user intention label corresponding to the probability value greater than or equal to the preset probability threshold is selected as the user intention corresponding to the test corpus.

[0025] Optionally, generating a text vector matrix of the test corpus includes:

[0026] Performing word segmentation processing on the test corpus to obtain multiple text segmentations;

[0027] Selecting one of the text segmentations from the multiple text segmentations as a target segmentation, and counting the number of co-occurrences of the target segmentation and its adjacent text segmentations within a preset neighborhood of the target segmentation;

[0028] Construct a co-occurrence matrix using the co-occurrence counts of each text segmentation word;

[0029] Convert the multiple text segmentations into word vectors respectively, and concatenate the word vectors into a vector matrix;

[0030] The co-occurrence matrix and the vector matrix are multiplied to obtain a text vector matrix.

[0031] Optionally, the identifying the difference nodes between the speech technique test flowchart and the speech technique reference flowchart, and highlighting the corresponding difference nodes according to a preset rule, includes:

[0032] A user intention node in the speech test flow chart is sequentially used as a detection node, and the length of the connection line between the detection node and the business intention node directly connected to the detection node is used as the detection distance;

[0033] The user intention node consistent with the detection node in the speech reference flow chart is used as a reference node, and the length of the connection line between the reference node and the business intention node directly connected to the reference node is used as a reference distance;

[0034] The distance difference between the detection distance and the reference distance is calculated. When the distance difference is greater than a preset distance threshold, the business intention node directly connected to the detection node and the business intention node directly connected to the reference node are used as the difference nodes.

[0035] Optionally, the calculating the probability value between the text feature and the preset user intention label using a pre-trained activation function includes:

[0036] The following activation function is used to calculate the probability value between the text feature and the preset user intention label:

[0037]

[0038] Among them, p(a|x) is the probability value between the text feature x and the user intention label a, w ais the weight vector of user intention label a, T is the transposition operator, exp is the expectation operator, and a is the number of preset user intention labels.

[0039] In order to solve the above problems, the present invention further provides a device for visualizing robot speech detection, comprising:

[0040] Test sample acquisition module: used to extract simulated user features of a preset simulated user, and select historical human-computer dialogue records corresponding to users whose simulated user features meet a preset role distance threshold from a preset human-computer dialogue library as test reference samples;

[0041] A test flow chart generation module is configured to use the user input information in the test reference sample as test corpus, receive simulated response speech given by a preset robot based on the test corpus, and construct a speech test flow chart for the preset robot based on the test corpus and the simulated response speech;

[0042] Reference flow chart generation module: used to construct the speech reference flow chart of the preset robot according to the test reference sample;

[0043] Flowchart comparison module: used to identify the difference nodes between the speech test flowchart and the speech reference flowchart, and highlight the corresponding difference nodes according to preset rules.

[0044] In order to solve the above problem, the present invention further provides an electronic device, comprising:

[0045] a memory storing at least one computer program; and

[0046] The processor executes the program stored in the memory to implement the above-mentioned robot speech detection visualization method.

[0047] In order to solve the above problems, the present invention also provides a computer-readable storage medium, in which at least one computer program is stored. The at least one computer program is executed by a processor in an electronic device to implement the above-mentioned robot speech detection visualization method.

[0048] The embodiment of the present invention utilizes a test reference sample to construct a speech reference flowchart. The speech reference flowchart can intuitively reflect the historical performance of the preset robot before the test, and then utilizes the user input information in the test reference sample as the test corpus, and then constructs a speech test flowchart based on the test corpus and the simulated reply speech of the preset robot. The speech test flowchart can vividly reflect the test performance of the current preset robot. Finally, the speech test flowchart and the speech reference flowchart are compared to obtain the difference nodes in the two diagrams, and the corresponding difference nodes are highlighted, which can show the test effect more intuitively and three-dimensionally. Therefore, the embodiment of the present invention can solve the problem of non-intuitive robot speech detection in scenarios such as financial automatic services and online medical self-service, and improve the visualization effect of robot speech detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 A flowchart of a method for visualizing robot speech detection provided by one embodiment of the present invention;

[0050] Figure 2 A schematic diagram of a detailed implementation flow of one step in a method for visualizing robot speech detection provided by an embodiment of the present invention;

[0051] Figure 3 A schematic diagram of a detailed implementation flow of another step in the method for visualizing robot speech detection provided by one embodiment of the present invention;

[0052] Figure 4 A schematic diagram of a detailed implementation flow of another step in the method for visualizing robot speech detection provided by one embodiment of the present invention;

[0053] Figure 5 A functional module diagram of a device for visualizing robot speech detection provided by one embodiment of the present invention;

[0054] Figure 6 A schematic diagram of the structure of an electronic device for implementing the method for visualizing robot speech detection provided in one embodiment of the present invention.

[0055] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0056] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0057] The embodiment of the present application provides a method for visualizing the detection of robot speech. The execution subject of the method for visualizing the detection of robot speech includes but is not limited to at least one of the electronic devices such as a server and a terminal that can be configured to execute the method provided by the embodiment of the present application. In other words, the method for visualizing the detection of robot speech can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server can be an independent server, or it can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0058] Reference Figure 1 FIG. 1 is a flow chart of a method for visualizing robot speech detection according to an embodiment of the present invention. In this embodiment, the method for visualizing robot speech detection includes:

[0059] S1. Extracting simulated user features of a preset simulated user, and selecting historical human-computer dialogue records corresponding to users whose simulated user features meet a preset role distance threshold from a preset human-computer dialogue library as test reference samples;

[0060] In an embodiment of the present invention, the preset simulated user refers to the user role set in the robot speech detection scenario. In actual application, relevant user roles can be set according to the scenarios covered by the actual robot speech. For example, in the fund consultation scenario, user roles such as ordinary company employees, self-employed individuals, and retirees can be set. Correspondingly, the basic information of the simulated user refers to the basic user information related to the robot speech scenario. For example, in the medical outpatient registration self-service consultation scenario, the corresponding basic information of the simulated user includes age, name, occupation, historical medical history, allergic medications, current disease symptoms, and other information.

[0061] In this embodiment of the present invention, the basic information of the simulated user can be manually set through the interface input function provided by the simulated human-computer dialogue window. Alternatively, the basic information of the simulated user can be stored in a configuration file, and the basic information of the corresponding simulated user can be randomly read from the configuration file before the robot speech detection is performed.

[0062] In another optional embodiment of the present invention, the basic information of the preset simulated user can be initialized using the user information in the historical human-computer dialogue record.

[0063] In the embodiment of the present invention, the preset human-computer dialogue database refers to a database storing historical human-computer dialogue records. The human-computer dialogue database can retrieve relevant dialogue records according to conditions such as dialogue time, dialogue scene, and dialogue users.

[0064] For details, see Figure 2 As shown, the process of selecting historical human-computer dialogue records corresponding to users whose characteristics with the simulated user meet a preset role distance threshold from the preset human-computer dialogue library as test reference samples includes:

[0065] S11, obtaining a user portrait set in the preset human-computer dialogue library;

[0066] S12, calculating the role distance between the simulated user feature and each user portrait in the user portrait set;

[0067] S13. Selecting a human-computer dialogue record corresponding to a user portrait having a role distance that meets a preset role distance threshold as the test reference sample.

[0068] In the embodiment of the present invention, the key user characteristics may be occupation characteristics, or a combination of occupation and age, etc. In practical applications, the corresponding key user characteristics may be determined according to the robot speech scenario to be detected.

[0069] In the embodiment of the present invention, a preset key user feature entry may be matched with the basic information of the simulated user to obtain a corresponding key user feature.

[0070] It can be understood that the historical human-computer dialogue records stored in the preset human-computer dialogue library can extract corresponding user attributes for each group of human-computer dialogue records, and relevant user portraits can be constructed based on the user attributes. Usually, human-computer dialogue records corresponding to the same user portrait have the same or similar characteristics.

[0071] In the embodiment of the present invention, the role distance between the key user feature and each of the user portraits may be calculated using distance formulas such as Euclidean distance and Mahalanobis distance.

[0072] In the embodiment of the present invention, the role distance threshold may be configured according to actual conditions.

[0073] Furthermore, a time screening condition may be added to screen out the human-computer dialogue records with a relatively recent time as the test reference samples.

[0074] The embodiment of the present invention can narrow down the speech scenarios of the robot to be tested by matching historical human-computer dialogue records related to the basic information of the simulated user in a preset human-computer dialogue library as test reference sample operations. Correspondingly, the corresponding robot speech scenarios to be tested can be changed by changing the basic information of the simulated user, thereby achieving full coverage of the robot speech scenario test.

[0075] S2. Using the user input information in the test reference sample as test corpus, receiving a simulated response speech given by a preset robot for the test corpus, and constructing a speech test flow chart for the preset robot based on the test corpus and the simulated response speech;

[0076] In an embodiment of the present invention, the test corpus can be automatically entered into the human-computer dialogue window through a simulated human-computer dialogue window at a certain preset speed. After the preset robot receives the corresponding test corpus, it generates relevant speech according to the preset logic, that is, simulates the reply speech.

[0077] In an embodiment of the present invention, after the preset robot completes the response to the relevant test corpus, a speech test flowchart of the preset robot can be constructed by combining the test corpus and the actual situation of the simulated response speech.

[0078] In an embodiment of the present invention, the speech test flowchart refers to the process of conducting speech test on the preset robot, in which the question-answer logic between the test corpus and the simulated reply speech of the preset robot is connected and displayed in the form of a flowchart, which can realize the conversion of text information into graph structure information.

[0079] For details, see Figure 3 As shown, the process of constructing a speech test flow chart of the preset robot based on the test corpus and the simulated reply speech includes:

[0080] S21, performing user intent recognition on the test corpus, and generating a corresponding user intent node according to the recognized user intent;

[0081] S22, identifying the business intent of the simulated reply speech, and generating a corresponding business intent node according to the identified business intent;

[0082] S23. Arrange all the user intent nodes and all the business intent nodes vertically according to the chronological order in which the test corpus and the simulated response words appear in the question and answer of the preset robot to obtain an intent node queue;

[0083] S24. Calculate the intention distance between every two adjacent intention nodes in the intention node queue, and generate a connecting line of preset proportional length according to the size of the intention distance, and use the connecting line to connect the adjacent intention nodes in series to obtain the speech test flowchart.

[0084] In an embodiment of the present invention, a preset text intent recognition model based on deep learning can be used to identify the user intent or the business intent.

[0085] Specifically, identifying user intent on the test corpus includes:

[0086] Generate a text vector matrix of the test corpus;

[0087] Extracting text features of the test corpus from the text vector matrix;

[0088] Calculate the probability value between the text feature and the preset user intention label using a pre-trained activation function;

[0089] The user intention label corresponding to the probability value greater than or equal to the preset probability threshold is selected as the user intention corresponding to the test corpus.

[0090] In an embodiment of the present invention, since the test corpus is composed of natural language, if the test corpus is directly analyzed, a large amount of computing resources will be occupied, resulting in low analysis efficiency. Therefore, the test corpus can be converted into a text vector matrix, and then the text content expressed in natural language can be converted into a numerical form.

[0091] In an embodiment of the present invention, methods such as Glove (Global Vectors for Word Representation) and Embedding Layer may be used to convert the test corpus into a text vector matrix.

[0092] In one embodiment of the present invention, generating the text vector matrix of the test corpus includes:

[0093] Performing word segmentation processing on the test corpus to obtain multiple text segmentations;

[0094] Selecting one of the text segmentations from the multiple text segmentations as a target segmentation, and counting the number of co-occurrences of the target segmentation and its adjacent text segmentations within a preset neighborhood of the target segmentation;

[0095] Construct a co-occurrence matrix using the co-occurrence counts of each text segmentation word;

[0096] Convert the multiple text segmentations into word vectors respectively, and concatenate the word vectors into a vector matrix;

[0097] The co-occurrence matrix and the vector matrix are multiplied to obtain a text vector matrix.

[0098] In detail, a preset standard dictionary may be used to perform word segmentation processing on the test corpus to obtain a plurality of text segmentations, wherein the standard dictionary contains a plurality of standard segmentations.

[0099] For example, the test corpus is searched in the standard dictionary according to different lengths. If the same standard segmentation as the test corpus is retrieved, the retrieved standard segmentation can be determined to be the text segmentation of the test corpus.

[0100] For example, the co-occurrence count corresponding to each text segmentation word may be used to construct a co-occurrence matrix as shown below:

[0101]

[0102] Among them, X i,j is the number of co-occurrences of keyword i and its adjacent text segment j in the test corpus.

[0103] In an embodiment of the present invention, a model with a word vector conversion function such as a word2vec model and an NLP (Natural Language Processing) model can be used to convert the multiple test corpora into word vectors respectively, and then the word vectors are spliced ​​into a vector matrix of the test corpus, and the vector matrix is ​​multiplied with the co-occurrence matrix to obtain a text vector matrix.

[0104] In an embodiment of the present invention, a preset deep learning-based text intent recognition model can be used to extract text features of the test corpus based on the text vector matrix.

[0105] In this embodiment of the present invention, the activation function includes but is not limited to the softmax activation function, the sigmoid activation function, and the relu activation function. The preset user intent tags are determined based on the scenarios covered by the actual robot's speech. For example, taking insurance self-service consultation as an example, the preset user intent tags include but are not limited to insurance cancellation consultation, false reporting consultation, and policy benefit inquiry.

[0106] In one embodiment of the present invention, the probability value may be calculated using the following activation function:

[0107]

[0108] Among them, p(a|x) is the probability value between the text feature x and the user intention label a, w a is the weight vector of user intention label a, T is the transposition operator, exp is the expectation operator, and A is the number of preset user intention labels.

[0109] It should be noted that the method for identifying the business intent of the simulated reply speech can be the same as the method for identifying the user intent of the test corpus, which will not be repeated here.

[0110] In the embodiment of the present invention, the intent distance between two adjacent intent nodes may be calculated using Euclidean distance, Mahalanobis distance, Manhattan distance formula, and the like.

[0111] In detail, the step of generating a connecting line of a preset proportional length according to the intended distance includes:

[0112] Normalize all the intention distances;

[0113] The normalized intention distance is multiplied by the preset ratio to obtain a connecting line of preset proportional length.

[0114] The embodiment of the present invention can intuitively display the speech test process of the preset robot by constructing a speech test flow chart of the preset robot.

[0115] S3. Construct a speech reference flow chart for the preset robot based on the test reference sample;

[0116] In an embodiment of the present invention, the speech reference flowchart is constructed based on the user input information and the corresponding robot reply information in the test reference sample. Therefore, the speech reference flowchart can reflect the historical performance of the preset robot in the same business scenario.

[0117] It should be noted that the method of constructing the speech reference flowchart of the preset robot based on the test reference sample is the same as constructing the speech test flowchart of the preset robot based on the test corpus and the simulated reply speech, which will not be repeated here.

[0118] S4. Identify the difference nodes between the speech test flowchart and the speech reference flowchart, and highlight the corresponding difference nodes according to preset rules.

[0119] It is understood that the test reference samples are derived from historical human-machine dialogue records in the preset human-machine dialogue library and can serve as historical versions of the robot's dialogue before the upgrade. By comparing the current dialogue test flowchart with the dialogue reference flowchart reflecting the historical version, some different nodes can be identified. These different nodes may be adjusted nodes or nodes that are inconsistent with the preset robot dialogue. Business personnel can focus on the different nodes to analyze the relevant robot dialogue and further improve the robot dialogue.

[0120] For details, see Figure 4 As shown, the identification of the difference nodes between the speech test flowchart and the speech reference flowchart includes:

[0121] S41, sequentially using one user intention node in the speech test flow chart as a detection node, and using the length of the connection line between the detection node and the business intention node directly connected to the detection node as the detection distance;

[0122] S42: Using the user intention node in the speech technique reference flow chart that is consistent with the detection node as a reference node, and using the length of the connection line between the reference node and the business intention node directly connected to the reference node as a reference distance;

[0123] S43. Calculate the distance difference between the detection distance and the reference distance. When the distance difference is greater than a preset distance threshold, use the business intention node directly connected to the detection node and the business intention node directly connected to the reference node as the difference nodes.

[0124] In this embodiment of the present invention, the distance threshold can be set based on actual test conditions. When the distance difference is greater than the preset distance threshold, it indicates that the responses given by the preset robot in the historical Q&A and the current test Q&A for the same user intent are significantly different, and require further analysis.

[0125] In the embodiment of the present invention, the preset rule may be to perform rendering operations such as highlighting and amplifying the difference nodes, so as to make the detected contrast effect more prominent.

[0126] The embodiment of the present invention utilizes a test reference sample to construct a speech reference flowchart. The speech reference flowchart can intuitively reflect the historical performance of the preset robot before the test, and then utilizes the user input information in the test reference sample as the test corpus, and then constructs a speech test flowchart based on the test corpus and the simulated reply speech of the preset robot. The speech test flowchart can vividly reflect the test performance of the current preset robot. Finally, the speech test flowchart and the speech reference flowchart are compared to obtain the difference nodes in the two diagrams, and the corresponding difference nodes are highlighted, which can show the test effect more intuitively and three-dimensionally. Therefore, the embodiment of the present invention can solve the problem of non-intuitive robot speech detection in scenarios such as financial automatic services and online medical self-service, and improve the visualization effect of robot speech detection.

[0127] like Figure 5 , which is a functional module diagram of a robot speech detection visualization device provided by an embodiment of the present invention.

[0128] The robot speech detection visualization device 100 described in the present invention can be installed in an electronic device. Based on the functions implemented, the robot speech detection visualization device 100 includes: a test sample acquisition module 101, a test flowchart generation module 102, a reference flowchart generation module 103, and a flowchart comparison module 104. A module, also referred to as a unit, refers to a series of computer program segments that can be executed by an electronic device processor and perform a fixed function. These are stored in the electronic device's memory.

[0129] In this embodiment, the functions of each module / unit are as follows:

[0130] The test sample acquisition module 101 is used to extract simulated user features of a preset simulated user, and select historical human-computer dialogue records corresponding to users whose simulated user features meet a preset role distance threshold from a preset human-computer dialogue library as test reference samples;

[0131] The test flowchart generating module 102 is configured to use the user input information in the test reference sample as test corpus, receive simulated response speech given by a preset robot based on the test corpus, and construct a speech test flowchart for the preset robot based on the test corpus and the simulated response speech;

[0132] The reference flow chart generating module 103 is used to construct a speech reference flow chart of the preset robot according to the test reference sample;

[0133] The flowchart comparison module 104 is used to identify the difference nodes between the speech test flowchart and the speech reference flowchart, and highlight the corresponding difference nodes according to preset rules.

[0134] In detail, each module in the robot speech detection visualization device 100 in the embodiment of the present invention adopts the same Figures 1 to 4 The technical means are the same as the robot speech detection visualization method described in , and can produce the same technical effects, so I will not go into details here.

[0135] like Figure 6 , which is a schematic diagram of the structure of an electronic device for implementing a method for visualizing robot speech detection provided by an embodiment of the present invention.

[0136] The electronic device 1 may include a processor 10, a memory 11 and a bus, and may also include a computer program stored in the memory 11 and executable on the processor 10, such as a robot speech detection visualization program.

[0137] Among them, the memory 11 includes at least one type of readable storage medium, and the readable storage medium includes a flash memory, a mobile hard disk, a multimedia card, a card-type memory (for example, SD or DX memory, etc.), a magnetic memory, a disk, an optical disk, etc. In some embodiments, the memory 11 can be an internal storage unit of the electronic device 1, such as a mobile hard disk of the electronic device 1. In other embodiments, the memory 11 can also be an external storage device of the electronic device 1, such as a plug-in mobile hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the electronic device 1. Furthermore, the memory 11 can also include both an internal storage unit of the electronic device 1 and an external storage device. The memory 11 can not only be used to store application software and various types of data installed on the electronic device 1, such as the code of the robot speech detection visualization program, etc., but can also be used to temporarily store data that has been output or is to be output.

[0138] In some embodiments, the processor 10 may be composed of an integrated circuit, for example, a single packaged integrated circuit, or a plurality of packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and a combination of various control chips. The processor 10 is the control core (Control Unit) of the electronic device, connecting the various components of the entire electronic device using various interfaces and lines, and executing the programs or modules stored in the memory 11 (such as a robot speech detection visualization program, etc.), as well as calling the data stored in the memory 11, to perform various functions of the electronic device 1 and process data.

[0139] The bus may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. The bus is configured to enable connection and communication between the memory 11 and at least one processor 10, etc.

[0140] Figure 6 Only the electronic device with components is shown, and it can be understood by those skilled in the art that Figure 6 The structure shown does not constitute a limitation on the electronic device 1 , and may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.

[0141] For example, although not shown, the electronic device 1 may further include a power source (such as a battery) for powering the various components. Preferably, the power source may be logically connected to the at least one processor 10 via a power management device, thereby implementing functions such as charging management, discharging management, and power consumption management through the power management device. The power source may further include any components such as one or more DC or AC power sources, a recharging device, a power failure detection circuit, a power converter or inverter, a power status indicator, etc. The electronic device 1 may further include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be described in detail here.

[0142] Furthermore, the electronic device 1 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device 1 and other electronic devices.

[0143] Optionally, the electronic device 1 may further include a user interface, which may be a display or an input unit (such as a keyboard). Optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touch device. The display may also be appropriately referred to as a display screen or a display unit, which is used to display information processed in the electronic device 1 and to display a visual user interface.

[0144] It should be understood that the embodiment is for illustration only and the scope of the patent application is not limited to this structure.

[0145] The robot speech detection visualization program stored in the memory 11 of the electronic device 1 is a combination of multiple instructions. When running in the processor 10, it can achieve the following:

[0146] Extracting a preset simulated user's simulated user characteristics, and selecting, from a preset human-computer dialogue library, historical human-computer dialogue records corresponding to users whose characteristics meet a preset role distance threshold with the simulated user characteristics as test reference samples;

[0147] Using the user input information in the test reference sample as test corpus, receiving a simulated response speech given by a preset robot for the test corpus, and constructing a speech test flow chart for the preset robot based on the test corpus and the simulated response speech;

[0148] Constructing a speech reference flow chart of the preset robot according to the test reference sample;

[0149] Identify the difference nodes between the speech test flowchart and the speech reference flowchart, and highlight the corresponding difference nodes according to preset rules.

[0150] Furthermore, if the modules / units integrated into the electronic device 1 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard drive, a magnetic disk, an optical disk, a computer memory, or a read-only memory (ROM).

[0151] The present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program. When the computer program is executed by a processor of an electronic device, the computer program can implement:

[0152] Extracting a preset simulated user's simulated user characteristics, and selecting, from a preset human-computer dialogue library, historical human-computer dialogue records corresponding to users whose characteristics meet a preset role distance threshold with the simulated user characteristics as test reference samples;

[0153] Using the user input information in the test reference sample as test corpus, receiving a simulated response speech given by a preset robot for the test corpus, and constructing a speech test flow chart for the preset robot based on the test corpus and the simulated response speech;

[0154] Constructing a speech reference flow chart of the preset robot according to the test reference sample;

[0155] Identify the difference nodes between the speech test flowchart and the speech reference flowchart, and highlight the corresponding difference nodes according to preset rules.

[0156] In addition, the functional modules in various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or hardware plus software functional modules.

[0157] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0158] Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims rather than the foregoing description, and all changes that come within the meaning and range of equivalents of the claims are intended to be embraced therein. Any reference to a figure in a claim should not be construed as limiting the claim to which it relates.

[0159] Blockchain, as used in this article, refers to a novel application model for computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Blockchain is essentially a decentralized database, a series of data blocks generated using cryptographic methods. Each block contains information about a batch of online transactions, used to verify the validity of this information (to prevent counterfeiting) and generate the next block. Blockchain can include the underlying blockchain platform, the platform product service layer, and the application service layer.

[0160] The embodiments of the present application can acquire and process relevant data based on holographic projection technology. Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to achieve the best results.

[0161] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a system claim may also be implemented by a single unit or device through software or hardware. Second-order terms are used to indicate names and do not imply any particular order.

[0162] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A visualization method for detecting robot speech, characterized in that: The method comprises: Extracting a preset simulated user's simulated user characteristics, and selecting, from a preset human-computer dialogue library, historical human-computer dialogue records corresponding to users whose characteristics meet a preset role distance threshold with the simulated user characteristics as test reference samples; Using the user input information in the test reference sample as test corpus, receiving a simulated response speech given by a preset robot for the test corpus, and constructing a speech test flow chart for the preset robot based on the test corpus and the simulated response speech; Constructing a speech reference flow chart of the preset robot according to the test reference sample; Identify the difference nodes between the speech test flowchart and the speech reference flowchart, and highlight the corresponding difference nodes according to preset rules; Among them, the construction of the speech test flowchart of the preset robot based on the test corpus and the simulated reply speech includes: performing user intent recognition on the test corpus, and generating corresponding user intent nodes according to the recognized user intent; performing business intent recognition on the simulated reply speech, and generating corresponding business intent nodes according to the recognized business intent; arranging all the user intent nodes and all the business intent nodes vertically according to the chronological order in which the test corpus and the simulated reply speech appear in the questions and answers of the preset robot to obtain an intention node queue; calculating the intention distance between every two adjacent intention nodes in the intention node queue, and generating a connecting line of a preset proportional length according to the size of the intention distance, and using the connecting line to connect adjacent intention nodes in series to obtain the speech test flowchart; The method of identifying the difference nodes between the speech test flowchart and the speech reference flowchart, and highlighting the corresponding difference nodes according to preset rules, includes: taking a user intention node in the speech test flowchart as a detection node in turn, and taking the length of the connection line between the detection node and the business intention node directly connected to the detection node as the detection distance; taking the user intention node consistent with the detection node in the speech reference flowchart as a reference node, and taking the length of the connection line between the reference node and the business intention node directly connected to the reference node as the reference distance; calculating the distance difference between the detection distance and the reference distance, and when the distance difference is greater than a preset distance threshold, taking the business intention node directly connected to the detection node and the business intention node directly connected to the reference node as the difference nodes.

2. The method for visualizing robot speech detection according to claim 1, wherein: The method of selecting historical human-computer dialogue records corresponding to users whose characteristics with the simulated user meet a preset role distance threshold from the preset human-computer dialogue library as test reference samples includes: Obtaining a user portrait set in the preset human-computer dialogue library; Calculating the role distance between the simulated user feature and each user portrait in the user portrait set; The human-computer dialogue record corresponding to the user portrait whose role distance meets the preset role distance threshold is selected as the test reference sample.

3. The method for visualizing robot speech detection according to claim 1, wherein: The performing user intent identification on the test corpus includes: Generate a text vector matrix of the test corpus; Extracting text features of the test corpus from the text vector matrix; Calculate the probability value between the text feature and the preset user intention label using a pre-trained activation function; The user intention label corresponding to the probability value greater than or equal to the preset probability threshold is selected as the user intention corresponding to the test corpus.

4. The method for visualizing robot speech detection according to claim 3, wherein: Generating the text vector matrix of the test corpus includes: Performing word segmentation processing on the test corpus to obtain multiple text segmentations; Selecting one of the text segmentations from the multiple text segmentations as a target segmentation, and counting the number of co-occurrences of the target segmentation and its adjacent text segmentations within a preset neighborhood of the target segmentation; Construct a co-occurrence matrix using the co-occurrence counts of each text segmentation word; Convert the multiple text segmentations into word vectors respectively, and concatenate the word vectors into a vector matrix; The co-occurrence matrix and the vector matrix are multiplied to obtain a text vector matrix.

5. The method for visualizing robot speech detection according to claim 3, wherein: The calculating of the probability value between the text feature and the preset user intention label using the pre-trained activation function includes: The following activation function is used to calculate the probability value between the text feature and the preset user intention label: in, Text features and user intent tags The probability value between Labeling user intent The weight vector of To find the transposition operator, To find the expected operator, The number of preset user intent labels.

6. A device for visualizing robot speech detection, used to implement the method for visualizing robot speech detection according to any one of claims 1 to 5, characterized in that: The device comprises: Test sample acquisition module: used to extract simulated user features of a preset simulated user, and select historical human-computer dialogue records corresponding to users whose simulated user features meet a preset role distance threshold from a preset human-computer dialogue library as test reference samples; A test flow chart generation module is configured to use the user input information in the test reference sample as test corpus, receive simulated response speech given by a preset robot based on the test corpus, and construct a speech test flow chart for the preset robot based on the test corpus and the simulated response speech; Reference flow chart generation module: used to construct the speech reference flow chart of the preset robot according to the test reference sample; Flowchart comparison module: used to identify the difference nodes between the speech test flowchart and the speech reference flowchart, and highlight the corresponding difference nodes according to preset rules.

7. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the robot speech detection visualization method according to any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for visualizing robot speech detection according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Intelligent question-and-answer multi-round interactive method and system based on visual flowchart

    CN106776649A

  • Robot verbal skill resource input method and device, electronic equipment and storage medium

    CN112148845A