Method, apparatus, electronic device, and storage medium for judging integrity of dialogue information

By extracting feature and vectorized representation of multiple rounds of dialogue information, and combining the keyword information of the theoretical text for weighting and quantitative calculation, the problem of low accuracy in dialogue information integrity judgment is solved, and high accuracy in dialogue information integrity judgment is achieved.

CN114840654BActive Publication Date: 2025-06-24PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210512467.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-12
Publication Date
2025-06-24
Estimated Expiration
2042-05-12

AI Technical Summary

Technical Problem

The prior art is difficult to effectively judge the integrity of dialogue information, resulting in low accuracy of information acquisition and the system cannot determine whether the dialogue information is complete in a timely manner.

Method used

By obtaining multiple rounds of dialogue information, extracting word sequences and performing vectorized representations, obtaining theoretical texts corresponding to dialogue information, extracting keyword information and performing weight quantization calculations, collecting keyword information with weights greater than preset weights as a set of strongly associated contents, calculating the matching values ​​of the main complaint representation and strongly associated contents set, and determining whether the dialogue information is complete based on the matching value and association threshold.

Benefits of technology

It improves the accuracy of judging the integrity of dialogue information, can effectively judge the integrity of dialogue information, and ensures the comprehensiveness and accuracy of information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114840654B_ABST
    Figure CN114840654B_ABST
Patent Text Reader

Abstract

The present invention relates to artificial intelligence technology, and discloses a method for judging the integrity of dialogue information, including: obtaining multi-round dialogue information, and extracting the main complaint representation of the dialogue information; obtaining a theoretical text corresponding to the dialogue information, and extracting keyword information of the theoretical text; performing weighted quantization calculation on the keyword information to obtain the weight of each keyword information, and collecting the keyword information with weights greater than a preset weight into a strongly associated content set; calculating a matching value between the main complaint representation and the strongly associated content set, and judging whether the dialogue information is complete according to the matching value and an association threshold. In addition, the present invention also relates to blockchain technology, and the data list can be stored in the nodes of the blockchain. The present invention also proposes a device for judging the integrity of dialogue information, an electronic device, and a storage medium. The present invention can solve the problem of low accuracy in judging the integrity of dialogue information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, apparatus, electronic device, and computer-readable storage medium for judging the integrity of dialogue information. Background Art

[0002] With the advent of the information age, especially the Internet age, all aspects of people's lives increasingly rely on the Internet, and any information is data. The Internet has facilitated our lives. For example, we can make an appointment with a doctor online instead of going to the hospital to register and consult a doctor specifically; we can use online shopping to meet our needs for purchasing items. Only when complete information is obtained and integrated and analyzed can we fully understand the needs of users and better provide services according to the wishes of users.

[0003] The massive text data emerging on the Internet has brought rich corpus resources, but at the same time has also posed great challenges to text perception, analysis, and processing. The redundancy of information in the big data era has led to low accuracy in people's access to information, and it is difficult to control the integrity of information. In the application systems of most enterprises, the dialogue between users and enterprises takes the form of one question and one answer, and cannot automatically enter the next program. Moreover, the obtained dialogue information is either missing or repeated, and the system cannot timely judge whether the dialogue information is complete. Therefore, improving the accuracy of judging the integrity of dialogue information has become an urgent problem to be solved. Summary of the Invention

[0004] The present invention provides a method, apparatus, and computer-readable storage medium for judging the integrity of dialogue information, and its main purpose is to solve the problem of low accuracy in judging the integrity of dialogue information.

[0005] To achieve the above object, a method for judging the integrity of dialogue information provided by the present invention includes:

[0006] Obtain multi-round dialogue information and extract the word sequence of the dialogue information;

[0007] Perform vectorization representation on the word sequence to obtain a main complaint representation;

[0008] Obtain the theoretical text corresponding to the dialogue information and extract the keyword information of the theoretical text;

[0009] Perform weighted quantization calculation on the keyword information to obtain the weight of each keyword information, and collect the keyword information with weights greater than the preset weight into a strongly associated content set;

[0010] Calculate the matching value between the main complaint representation and the strongly associated content set, and judge whether the dialogue information is complete according to the matching value and the association threshold.

[0011] Optionally, the word sequence for extracting the conversation information includes:

[0012] Generating a text set of the conversation information;

[0013] Using a preset stop word list to filter out the stop words in the text set;

[0014] Performing low-frequency word removal processing on the filtered text set;

[0015] Performing word segmentation on the text set obtained after low-frequency word removal processing to obtain a word sequence.

[0016] Optionally, the vectorized representation of the word sequence to obtain the chief complaint representation includes:

[0017] Using a pre-trained corpus model to represent each word sequence as an n-dimensional word vector;

[0018] Performing weighted calculation on the word vectors to obtain the weight values of the word vectors;

[0019] Selecting the top N items with the highest weight values according to the preset vector dimensionality reduction setting as the chief complaint representation.

[0020] Optionally, the extraction of the keyword information of the theoretical text includes:

[0021] Randomly selecting a part of the theoretical text to generate a theoretical data set of the selected part of the theoretical text, and using stratified sampling to divide the theoretical data set into a training set and a test set;

[0022] Performing stop word and word segmentation processing on the training set to obtain a training set corpus;

[0023] Performing stop word and word segmentation processing on the test set to obtain a test set corpus;

[0024] Constructing a keyword extraction model according to the training set corpus and the test set corpus, and using the keyword extraction model to extract the keyword information of the theoretical text.

[0025] Optionally, the constructing of the keyword extraction model according to the training set corpus and the test set corpus includes:

[0026] Constructing a corpus matrix of the training set corpus, and using the corpus matrix to train a preset keyword extraction model;

[0027] Constructing a corpus matrix of the test set corpus, and using the test set corpus to verify the accuracy rate of the keyword extraction model until the accuracy rate is greater than a preset accuracy rate threshold to obtain a trained keyword extraction model.

[0028] Optionally, the weighted quantization calculation of the keyword information to obtain the weight of each piece of keyword information includes:

[0029] Select one of the keyword information one by one as the target keyword information;

[0030] Perform unique ID coding on each keyword to obtain a word ID;

[0031] Count the occurrence frequency of each word ID in the target keyword information to obtain an ID word frequency;

[0032] Assign the ID word frequency to a blank matrix to obtain a word frequency statistical matrix;

[0033] Use a preset weight algorithm to calculate the word frequency statistical matrix to obtain the weight of each piece of keyword information in the keyword information.

[0034] Optionally, the calculation of the matching value between the main complaint representation and the strongly associated content set includes:

[0035]

[0036] where A = (a1, a2,..., a i ,..., a n ), B = (b1, b2,..., b i ,..., b n ), cos x is the matching value, A is the main complaint representation, B is the strongly associated content set, a i is the i-th word sequence in the main complaint representation, b i is the i-th keyword in the strongly associated content set, and n represents the number of keywords in the strongly associated content set.

[0037] To solve the above problems, the present invention also provides a device for judging the integrity of dialogue information, and the device includes:

[0038] A feature extraction module, configured to obtain multi-round dialogue information and extract the word sequence of the dialogue information;

[0039] A main complaint representation module, configured to perform vectorization representation on the word sequence to obtain a main complaint representation;

[0040] A theoretical text module, configured to obtain a theoretical text corresponding to the dialogue information and extract the keyword information of the theoretical text;

[0041] A weight calculation module, configured to perform weighted quantization calculation on the keyword information to obtain the weight of each piece of keyword information, and collect the keyword information with weights greater than a preset weight as a strongly associated content set;

[0042] A matching value calculation module, configured to calculate a matching value between the main complaint representation and the strongly associated content set, and determine whether the conversation information is complete according to the matching value and the association threshold.

[0043] To solve the above problems, the present invention also provides an electronic device, which includes:

[0044] At least one processor; and,

[0045] A memory communicatively connected to the at least one processor; wherein,

[0046] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the above-mentioned conversation information integrity judgment method.

[0047] To solve the above problems, the present invention also provides a computer-readable storage medium, in which at least one computer program is stored, and the at least one computer program is executed by a processor in an electronic device to implement the above-mentioned conversation information integrity judgment method.

[0048] In an embodiment of the present invention, feature extraction is performed on the obtained multi-round conversation information, and text vectorization is performed to generate a main complaint representation of the conversation information, which helps to improve the retrieval query of the conversation information. Furthermore, it is easy to calculate the matching value between the main complaint representation and the strongly associated content set, obtain a theoretical text corresponding to the conversation information, calculate the weight of the theoretical text to generate a strongly associated content set, ensure the comprehensiveness of the corpus information related to the theoretical text, and enable the judgment of whether the conversation information is complete. Therefore, the present invention proposes a conversation information integrity judgment method, device, electronic device, and computer-readable storage medium, which can solve the problem of low accuracy in judging the integrity of conversation information. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 A flowchart of the conversation information integrity judgment method provided by an embodiment of the present invention;

[0050] Figure 2 A flowchart of the feature extraction provided by an embodiment of the present invention;

[0051] Figure 3 A flowchart of the main complaint representation provided by an embodiment of the present invention;

[0052] Figure 4 A functional module diagram of the conversation information integrity judgment device provided by an embodiment of the present invention;

[0053] Figure 5 A schematic structural diagram of an electronic device for implementing the method for judging the integrity of conversation information provided by an embodiment of the present invention.

[0054] The implementation, functional features and advantages of the object of the present invention will be further described in conjunction with the embodiments with reference to the accompanying drawings. Specific embodiments

[0055] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0056] The embodiments of the present application provide a method for judging the integrity of conversation information. The execution subject of the method for judging the integrity of conversation information includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiments of the present application. In other words, the method for judging the integrity of conversation information can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0057] Refer to Figure 1 As shown, it is a flowchart of a method for judging the integrity of conversation information provided by an embodiment of the present invention. In this embodiment, the method for judging the integrity of conversation information includes:

[0058] S1. Obtain multi-round conversation information and extract the word sequence of the conversation information.

[0059] In the embodiments of the present invention, the conversation information can be a description of the patient's own symptom information entered in the dialog box during an online medical consultation; it can be an inquiry about a commodity by a Taobao user when having a purchase intention; it can be an inquiry about some time and location information by a Dianping user when wanting to go to a certain merchant.

[0060] Specifically, some common questions can be displayed according to the results of big data surveys. For example, on the medical consultation interface, it can display: "What help do you need?", "Do you have a history of allergies?", "Do you have a family genetic history?", "What symptoms do you have now?"; It can also set automatic replies according to customer questions to obtain multi-round conversation information.

[0061] Furthermore, the conversation information contains a large amount of text data. If the conversation information is directly analyzed, it will consume a large amount of computing resources and lead to low analysis efficiency. Therefore, sub-word sequence extraction is performed on the conversation information to obtain the word sequence corresponding to the conversation information, where the word sequence is used to represent each word itself and its frequency.

[0062] In the embodiment of the present invention, the extraction of the word sequence of the conversation information includes:

[0063] S21. Generate a text set of the conversation information;

[0064] S22. Use a preset stop word list to filter out the stop words in the text set;

[0065] S23. Perform low-frequency word removal processing on the filtered text set;

[0066] S24. Perform word segmentation processing on the text set obtained after low-frequency word removal processing to obtain a word sequence.

[0067] Specifically, the so-called stop words refer to words that, although having a high frequency of occurrence in the text set, contribute nothing to classification and only increase the dimensionality of the feature space and the complexity of the classification operation. Such useless words include function words such as modal particles, adverbs, conjunctions, and prepositions. Before performing word segmentation processing, it is necessary to introduce a stop word list to filter out stop words, which can achieve the effect of reducing noise; low-frequency words, as the name implies, are words that appear less frequently in the data. Such data actually has a certain amount of information, but when low-frequency words are put into the model for operation, they often maintain their random initial state, adding noise to the model. For the processing of low-frequency words, a simple method is to remove them.

[0068] Specifically, the establishment method of the preset stop word list can be divided into manual establishment and automatic establishment of the stop word list based on probability statistics. The manual establishment of the stop word list is to select certain word sets according to the subjective judgment of linguistics experts or select specific words for a specific application field to form a stop word list: for the English stop word list, the more famous ones are the stop word list published by VanRijsbergen and the Brown Corpus stop word list. The automatic establishment based on probability statistics is to construct a stop word list based on word frequency information, or obtain some stop words from the preliminary word segmentation results, and then continuously update and verify according to the segmentation results during the subsequent word segmentation process. The automatic establishment based on probability statistics mainly automatically obtains the stop word list by adopting techniques such as entropy, joint entropy, and resampling based on the KL distribution of TF / IDF words.

[0069] Specifically, the training text set for generating the conversation information refers to importing the conversation information into a pre-stored database, using numpy to read out all the required data, designing identifiers, concatenating all the data, generating corresponding identifiers, and storing them in a classified and structured manner.

[0070] Specifically, the processing of removing low-frequency words from the filtered training text set can be carried out according to the word frequency processing method. When the number of occurrences of a word is lower than the minimum threshold, it is removed; or the words are sorted according to the number of occurrences, and a certain proportion of low-frequency words are eliminated. There are two ways to calculate word frequency. One is directly the number of times a word appears in the text, and the other calculation method is the proportion of the total number of times a word appears in the text to the total number of words in the text.

[0071] S2. Perform vectorization representation on the word sequence to obtain the main complaint representation.

[0072] In the embodiment of the present invention, the performing vectorization representation on the word sequence to obtain the main complaint representation includes:

[0073] S31. Use a pre-trained corpus model to represent each word sequence as an n-dimensional word vector;

[0074] S32. Perform weighted calculation on the word vectors to obtain the weight values of the word vectors;

[0075] S33. Select the top N items with the highest weight values as the main complaint representation according to the preset vector dimensionality reduction setting.

[0076] Specifically, the corpus model may include, but is not limited to, the One-Hot Representation model, the BOW model, the word set model, etc. When using the One-Hot Representation model to represent text as a vector, each word is represented as a long vector, the dimension of the vector is the size of the vocabulary, the current position of the word is represented by 1, and other positions are represented by 0, without considering the frequency of word occurrence.

[0077] Specifically, the vector dimensionality reduction setting can be understood as when training a model, features will be selected to control the dimensionality range of each type of feature. For example, age is represented in the One-Hot way and divided into segments such as 0 - 10, 10 - 20,..., >100, that is, age is mapped into 11 dimensions. Compared with representing each age in one dimension, this is also a method of dimensionality reduction. Because in the original high-dimensional space, there is redundant information and noise information, which will introduce errors in actual applications and affect the accuracy; while dimensionality reduction can extract the essential structure inside the data, reduce the errors caused by redundant information and noise information, and improve the accuracy in applications.

[0078] In an embodiment of the present invention, the weighted calculation of the word vectors may be to generate word vectors using the Skip-Gram model of Word2Vec, and then calculate the weights of the word vectors using TF-IDF values to weight the word vectors with TF-IDF values. The weighted vectors can better represent the text. The TF-IDF values of each word in different documents are pre-calculated, and the Word2Vec word vectors of each word are also pre-trained. When selecting the word vector matrix data of each input text, the TF-IDF value of each word in this text is obtained by using the hashing method, and the weighted word vector matrix is obtained through matrix operations.

[0079] Specifically, the selection of the top N items with the highest weight values as the main complaint representation according to the preset vector dimensionality reduction setting can be understood as follows: when we set the size of the word vector matrix of each text to be the same, both are 150×100, where 150 is the number of text words and 100 is the dimensionality of the word vector. When the text lengths are different and the number of words is more than 150, the first 150 words are taken according to the word occurrence frequency; when the number of words is less than 150, the matrix is filled with 0s. We can also set the size of the word vector matrix of each text to be the same, both are 50×10, where 50 is the number of text words and 10 is the dimensionality of the word vector. When the text lengths are different and the number of words is more than 50, the first 50 words are taken according to the word occurrence frequency; when the number of words is less than 50, the matrix is filled with 0s.

[0080] S3. Obtain the theoretical text corresponding to the conversation information, and extract the keyword information of the theoretical text.

[0081] In an embodiment of the present invention, the theoretical text corresponding to the conversation information may be, in the medical field, a targeted solution given by an expert to a patient's condition, or detailed theoretical knowledge of a certain disease by a doctor. The keyword information may be "cold", "fever", "inflammation", "infection", etc.; it may also be, in the financial field, a professional analysis by a securities broker of the recent market. The keyword information may be "bottom fishing", "annualized rate of return", "LOF fund", "stock index futures", "offshore finance", and so on.

[0082] In an embodiment of the present invention, the extraction of the keyword information of the theoretical text includes: randomly selecting a part of the theoretical text to generate a theoretical data set of the selected part of the theoretical text, and using stratified sampling to divide the theoretical data set into a training set and a test set; performing stop word and word segmentation processing on the training set to obtain the training set corpus; performing stop word and word segmentation processing on the test set to obtain the test set corpus; constructing a keyword extraction model according to the training set corpus and the test set corpus, and using the keyword extraction model to extract the keyword information of the theoretical text.

[0083] Specifically, the stratified sampling is to group the overall survey objects according to different characteristics, and then use a random method to extract a certain number of samples from each layer to form a new set; the word segmentation process can be understood as processing a sentence into several words. For example, for a sentence: "Xiaoming came to Liwan District", the result of word segmentation after statistical analysis in our corpus is: "Xiaoming / came / Liwan / District".

[0084] Specifically, constructing the keyword extraction model according to the training set corpus and the test set corpus includes: constructing the corpus matrix of the training set corpus, and using the corpus matrix to train a preset keyword extraction model; constructing the corpus matrix of the test set corpus, and using the test set corpus to verify the accuracy rate of the keyword extraction model until the accuracy rate is greater than the preset accuracy rate threshold, so as to obtain the trained keyword extraction model.

[0085] Specifically, the corpus matrix of the training set corpus can be obtained by using the Gensim module. The process can be to import relevant data packets, such as the jieba and gensim packages, and then load and clean the corpus. The processed corpus is used to train a word vector model through Gensim, and this model is saved for future use. Use the word vector model to obtain the word vector of a specified vocabulary, and judge the similarity between the word vector of the specified vocabulary and the vocabulary in the word vector model to obtain the corpus matrix.

[0086] S4. Perform weighted quantization calculation on the keyword information to obtain the weight of each piece of keyword information, and collect the keyword information with weights greater than the preset weight as the strong association content set.

[0087] In the embodiment of the present invention, performing weighted quantization calculation on the keyword information to obtain the weight of each piece of keyword information includes: selecting one of the keyword information one by one as the target keyword information; performing unique ID coding on each keyword to obtain a word ID; counting the occurrence frequency of each word ID in the target keyword information to obtain the ID word frequency; assigning the ID word frequency to a blank matrix to obtain a word frequency statistical matrix; using a preset weight algorithm to calculate the word frequency statistical matrix to obtain the weight of each piece of keyword information in the keyword information.

[0088] Specifically, the target keyword information can be a certain keyword information among many keyword information. For example, the keyword information can be "I have a cold today", "When I got up this morning, I felt my face was a bit swollen", etc. When a certain keyword information is "I have a cold today", the keywords related to the keyword information "I have a cold today" include: "cough", "dizziness", "fever", "fatigue", "nasal congestion", etc.; the ID is unique, can represent the attributes of each keyword, and has a one-to-one correspondence. For example, Taobao user ID, each member has a unique user ID, which represents identity and qualification. Another example is everyone's ID number, which can be used to distinguish citizenship.

[0089] Specifically, the statistics of the occurrence frequency of each word ID in the target keyword information includes: reading a text from a file, reading an article from a file. Considering that some words need to be capitalized at the beginning of a sentence, in fact, uppercase and lowercase should be regarded as one word, but they will be separated during statistics. Therefore, all uppercase letters are converted to lowercase. Then there will be various symbols in a text, and these symbols will affect the ranking of word frequencies. Therefore, these symbols need to be removed. Here, the replace method of the string can be used to replace other symbols with space characters. Then every time a word appears, the count needs to be incremented by 1. Treat each word and its occurrence count as a key-value pair for processing, and output each word and its appearance frequency; the assignment of the ID word frequency to the blank matrix can be performed through the zeros function. Enter a = zeros(2,3), which generates a matrix of all zeros, and use Matlab to assign values to the blank matrix.

[0090] S5. Calculate the matching value between the main complaint representation and the strongly associated content set, and judge whether the dialogue information is complete according to the matching value and the association threshold.

[0091] In the embodiment of the present invention, the calculation of the matching value between the main complaint representation and the strongly associated content set includes:

[0092]

[0093] Among them, A = (a1, a2,..., a i ,..., a n ), B = (b1, b2,..., b i ,..., b n ), cos x is the matching value, A is the main complaint representation, B is the strongly associated content set, a i is the i-th word sequence in the main complaint representation, b i is the i-th keyword in the strongly associated content set, and n represents the number of keywords in the strongly associated content set.

[0094] Specifically, when the matching value ≥ the correlation threshold, the main complaint characterization has a high degree of matching with the strongly associated content set, and it can be determined that the conversation information is complete; when the matching value ≤ the correlation threshold, the main complaint characterization has a low degree of matching with the strongly associated content set, and it can be determined that the conversation information is incomplete.

[0095] In the embodiments of the present invention, feature extraction is performed on the obtained multi-round conversation information, and text vectorization is performed to generate a main complaint characterization of the conversation information, which helps to improve the retrieval and query of the conversation information. Furthermore, it is easy to calculate the matching value between the main complaint characterization and the strongly associated content set, obtain the theoretical text corresponding to the conversation information, perform weight calculation on the theoretical text to generate a strongly associated content set, ensure the comprehensiveness of the corpus information related to the theoretical text, and enable the determination of whether the conversation information is complete. Therefore, the present invention proposes a method for judging the integrity of conversation information, which can solve the problem of low accuracy in judging the integrity of conversation information.

[0096] As Figure 4 shown, it is a functional module diagram of a device for judging the integrity of conversation information provided by an embodiment of the present invention.

[0097] The device 100 for judging the integrity of conversation information according to the present invention can be installed in an electronic device. According to the functions achieved, the device 100 for judging the integrity of conversation information can include a feature extraction module 101, a main complaint characterization module 102, a theoretical text module 103, a weight calculation module 104, and a matching value calculation module 105. The modules in the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.

[0098] In this embodiment, the functions of each module / unit are as follows:

[0099] The feature extraction module is used to obtain multi-round conversation information and extract the word sequence of the conversation information;

[0100] The main complaint characterization module is used to perform vectorized characterization on the word sequence to obtain a main complaint characterization;

[0101] The theoretical text module is used to obtain the theoretical text corresponding to the conversation information and extract the keyword information of the theoretical text;

[0102] The weight calculation module is used to perform weighted quantization calculation on the keyword information to obtain the weight of each keyword information, and collect the keyword information with weights greater than the preset weight as the strongly associated content set;

[0103] A matching value calculation module is used to calculate the matching value between the chief complaint characterization and the strongly associated content set, and determine whether the dialogue information is complete according to the matching value and the association threshold.

[0104] As Figure 5 shown, it is a schematic structural diagram of an electronic device for implementing a method for judging the integrity of dialogue information provided by an embodiment of the present invention.

[0105] The electronic device 1 may include a processor 10, a memory 11, a communication bus 12, and a communication interface 13, and may further include a computer program stored in the memory 11 and operable on the processor 10, such as a dialogue information integrity judgment program.

[0106] Among them, the processor 10 may be composed of integrated circuits in some embodiments. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and lines, and by running or executing programs or modules stored in the memory 11 (such as executing a dialogue information integrity judgment program, etc.), and calling data stored in the memory 11, to perform various functions of the electronic device and process data.

[0107] The memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), magnetic memories, magnetic disks, optical disks, etc. The memory 11 may be an internal storage unit of the electronic device in some embodiments, such as the mobile hard disk of the electronic device. The memory 11 may also be an external storage device of the electronic device in other embodiments, such as a plug-in mobile hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. Further, the memory 11 may also include both an internal storage unit and an external storage device of the electronic device. The memory 11 can be used not only to store application software installed on the electronic device and various types of data, such as the code of a dialogue information integrity judgment program, etc., but also to temporarily store data that has been output or will be output.

[0108] The communication bus 12 can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is configured to enable connection communication between the memory 11 and at least one processor 10, etc.

[0109] The communication interface 13 is used for communication between the above-mentioned electronic device and other devices, including a network interface and a user interface. Optionally, the network interface can include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between this electronic device and other electronic devices. The user interface can be a display, an input unit (such as a keyboard), and optionally, the user interface can also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display can be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display can also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the electronic device and to display a visual user interface.

[0110] Only the electronic device with components is shown in the figure. Those skilled in the art can understand that the structure shown in the figure does not constitute a limitation on the electronic device, and it can include fewer or more components than shown in the figure, or combine certain components, or have different component arrangements.

[0111] For example, although not shown, the electronic device may further include a power source (such as a battery) for supplying power to each component. Preferably, the power source can be logically connected to the at least one processor 10 through a power management device, so as to implement functions such as charge management, discharge management, and power consumption management through the power management device. The power source can also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or an inverter, and a power status indicator. The electronic device may further include a variety of sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.

[0112] It should be understood that the above embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.

[0113] The conversation information integrity judgment program stored in the memory 11 of the electronic device 1 is a combination of multiple instructions, and when running in the processor 10, it can achieve:

[0114] Obtain multi-round conversation information and extract the word sequence of the conversation information;

[0115] Perform vectorized representation on the word sequence to obtain the main complaint representation;

[0116] Obtain the theoretical text corresponding to the conversation information and extract the keyword information of the theoretical text;

[0117] Perform weighted quantization calculation on the keyword information to obtain the weight of each keyword information, and collect the keyword information with weights greater than the preset weight as the strongly associated content set;

[0118] Calculate the matching value between the main complaint representation and the strongly associated content set, and judge whether the conversation information is complete according to the matching value and the association threshold. Specifically, the specific implementation method of the processor 10 for the above instructions can refer to the description of the relevant steps in the corresponding embodiments of the attached drawings, which will not be elaborated here.

[0119] Furthermore, if the module / unit integrated in the electronic device 1 is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory).

[0120] The present invention also provides a computer-readable storage medium, and the readable storage medium stores a computer program, and when the computer program is executed by the processor of the electronic device, it can achieve:

[0121] Obtain multi-round conversation information and extract the word sequence of the conversation information;

[0122] Perform vectorized representation on the word sequence to obtain the main complaint representation;

[0123] Obtain the theoretical text corresponding to the conversation information and extract the keyword information of the theoretical text;

[0124] Perform weighted quantization calculation on the keyword information to obtain the weight of each keyword information, and collect the keyword information with weights greater than the preset weight as the strongly associated content set;

[0125] Calculate the matching value between the main complaint representation and the strongly associated content set, and determine whether the dialogue information is complete according to the matching value and the association threshold.

[0126] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.

[0127] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0128] In addition, the functional modules in various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of hardware plus software functional modules.

[0129] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention.

[0130] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any associated drawing marks in the claims should not be regarded as limiting the claims involved.

[0131] The blockchain referred to in the present invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain, essentially a decentralized database, is a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity (anti-counterfeiting) of the information and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, an application service layer, etc.

[0132] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use the knowledge to obtain the best results.

[0133] In addition, it is obvious that the word "including" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices stated in the system claims can also be implemented by one unit or device through software or hardware. The terms such as first and second are used to represent names and do not represent any specific order.

[0134] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for judging the integrity of dialogue information, characterized in that, The method includes: Obtaining multi-round conversation information, extracting word sequences obtained by segmenting the conversation information, where the word sequences represent each word in the conversation information and the corresponding word frequencies; Using a pre-trained corpus model to represent each word sequence as an n-dimensional word vector to obtain multiple word vectors, performing weighted calculation on each word vector to obtain the weight value of each word vector, and selecting the top N word vectors with the highest weight values as the main complaint representation according to a preset vector dimensionality reduction setting; Obtaining a theoretical text corresponding to the conversation information, and extracting keyword information of the theoretical text; Selecting one of them as the target keyword information one by one from the keyword information, performing unique ID encoding on each keyword to obtain a word ID, counting the occurrence frequency of each word ID in the target keyword information to obtain an ID word frequency, assigning the ID word frequency to a blank matrix to obtain a word frequency statistical matrix, calculating the word frequency statistical matrix using a preset weight algorithm to obtain the weight of each keyword information in the keyword information, and collecting the keyword information with weights greater than the preset weight as a strongly associated content set; Calculating the matching value between the main complaint representation and the strongly associated content set, and judging whether the conversation information is complete according to the matching value and the association threshold.

2. The method for judging the integrity of conversation information according to claim 1, characterized in that, The extracting word sequences obtained by segmenting the conversation information includes: Generating a text set of the conversation information; Filtering stop words in the text set using a preset stop word list; Performing low-frequency word removal processing on the filtered text set; Performing word segmentation processing on the text set obtained after low-frequency word removal processing to obtain word sequences.

3. The method for judging the integrity of conversation information according to claim 1, wherein, The extracting keyword information of the theoretical text includes: Randomly selecting a part of the theoretical text to generate a theoretical data set of the selected part of the theoretical text, and using stratified sampling to divide the theoretical data set into a training set and a test set; Performing stop word and word segmentation processing on the training set to obtain a training set corpus; Performing stop word and word segmentation processing on the test set to obtain a test set corpus; Constructing a keyword extraction model according to the training set corpus and the test set corpus, and using the keyword extraction model to extract keyword information of the theoretical text.

4. The method for judging the integrity of conversation information according to claim 3, wherein The constructing a keyword extraction model according to the training set corpus and the test set corpus includes: Constructing a corpus matrix of the training set corpus, and training a preset keyword extraction model using the corpus matrix; Constructing a corpus matrix of the test set corpus, and verifying the accuracy rate of the keyword extraction model using the test set corpus until the accuracy rate is greater than a preset accuracy rate threshold to obtain a trained keyword extraction model.

5. The method for judging the integrity of dialogue information according to any one of claims 1 to 4, characterized in that The calculating the matching value between the main complaint representation and the strongly associated content set includes: Among them, A = ( , ,..., ,..., ), B = ( , ,..., ,..., ), is the matching value, A is the main complaint characterization, B is the set of strongly associated content, is the th word sequence in the main complaint characterization, is the th keyword in the set of strongly associated content, and n represents the number of keywords in the set of strongly associated content.

6. A device for judging the integrity of dialogue information, characterized in that, The device includes: A feature extraction module, configured to obtain multi-round conversation information, extract word sequences obtained by segmenting the conversation information, where the word sequences represent each word in the conversation information and the corresponding word frequencies; The chief complaint representation module is used to represent each word sequence as an n-dimensional word vector by using a pre-trained corpus model to obtain multiple word vectors, perform weighted calculation on each word vector to obtain the weight value of each word vector, and select the top N word vectors with the highest weight values as the chief complaint representation according to the preset vector dimensionality reduction setting; The theoretical text module is used to obtain the theoretical text corresponding to the dialogue information and extract the keyword information of the theoretical text; The weight calculation module is used to select one of them as the target keyword information from the keyword information one by one, perform unique ID encoding on each keyword to obtain the word ID, count the occurrence frequency of each word ID in the target keyword information to obtain the ID word frequency, assign the ID word frequency to the blank matrix to obtain the word frequency statistical matrix, and calculate the word frequency statistical matrix by using the preset weight algorithm to obtain the weight of each keyword information in the keyword information, and collect the keyword information with the weight greater than the preset weight as the strongly associated content set; The matching value calculation module is used to calculate the matching value between the chief complaint representation and the strongly associated content set, and judge whether the dialogue information is complete according to the matching value and the association threshold.

7. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the dialogue information integrity judgment method according to any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the dialogue information integrity judgment method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Oral retelling marking method and system

    CN108428382A

  • Automatic question answering method and device

    CN109522395A