A data processing method, device and electronic device for mining user requirements

By locating the user's negative intentions in online promotion activities, expanding the dialogue content, performing cluster analysis, and determining the weight of the core words, the problem that the existing technology cannot systematically explore user needs, and achieving more efficient and accurate user needs identification.

CN112541792BActive Publication Date: 2025-06-27ZUOYEBANG EDUCATION TECH (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011522401.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-22
Publication Date
2025-06-27
Estimated Expiration
2040-12-22

AI Technical Summary

Technical Problem

The existing online promotion activities cannot systematically tap into user needs, resulting in low efficiency and inability to promote on a large scale.

Method used

By positioning the statements expressing negative intentions in the user's dialogue content, expanding the dialogue content, performing cluster analysis, determining the weight of the core words, and determining the theme of user needs based on the weight.

Benefits of technology

It improves the efficiency and effectiveness of user demand mining, makes it more systematic and can more accurately identify users' real needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112541792B_ABST
    Figure CN112541792B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of data information processing, and provides a data processing method, device, electronic device and recording medium for mining user requirements. The method includes: locating the sentences expressing negative willingness in the user's conversation content, and expanding the conversation content with this sentence as the center; performing clustering analysis on the expanded conversation content to determine the weights of the core words in each category; determining the themes of each category, that is, user requirements, based on the weights of the core words. By locating the approximate position of user requirements, improving the quality and concentration of alternative sentences through clustering analysis, and sorting within each category using the weights of core words, the present invention improves the efficiency and effect of subsequent induction work, making the mining of user requirements more systematic.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data information processing, and is particularly applicable to data information processing in online services. More specifically, it relates to a data processing method, device, and electronic device for mining user needs. Background Art

[0002] In traditional commercial promotion activities, publicity and promotion are often carried out offline to users. Now, the customer service center is gradually used to replace the offline promotion method to publicize to users.

[0003] When the salesperson in the customer service center promotes and publicizes to users, many users will refuse for various reasons. Some users really don't need it, but some users refuse because the product does not meet their own needs. In order to make more users accept and use the product, it is necessary to mine the needs of users, improve in subsequent activities, and solve the needs of users.

[0004] The existing method for mining user needs is to summarize and generalize by the salesperson himself. This method is inefficient and cannot be widely promoted. Currently, a systematic method that can conveniently and accurately mine user needs is required. Summary of the Invention

[0005] (1) Technical Problems to be Solved

[0006] The present invention aims to solve the problem that user needs cannot be systematically mined in existing online promotion activities.

[0007] (2) Technical Solutions

[0008] To solve the above technical problems, one aspect of the present invention proposes a data processing method for mining user needs, including:

[0009] Locate the statements expressing negative willingness in the user's conversation content, and expand the conversation content centered on this statement;

[0010] Perform cluster analysis on the expanded conversation content, and determine the weights of the core words in each category;

[0011] Determine the theme of each category based on the weights of the core words, that is, the user needs.

[0012] According to the preferred embodiment of the present invention, locating the statements expressing negative willingness in the user's conversation content is specifically:

[0013] Adopt the method of manual indexing and / or indexing by an intention recognition model.

[0014] According to the preferred embodiment of the present invention, expanding the conversation content centered on this statement is specifically:

[0015] Centered on the indexed sentence expressing negative intention, the dialogue is extended forward by N rounds and backward by M rounds, where N and M are natural numbers.

[0016] According to a preferred embodiment of the present invention, cluster analysis is performed on the extended conversation content, specifically:

[0017] Split the extended dialogue into short sentences according to punctuation marks;

[0018] A cluster analysis model is used to perform cluster analysis on short sentences, and the cluster analysis model adopts the Biterm TopicModel algorithm.

[0019] According to a preferred embodiment of the present invention, the weight of the core words in each category is determined as follows:

[0020] The Biterm Topic Model algorithm is used to determine the weight of words in each category, and the first i words are selected as core words, where i is a natural number.

[0021] According to a preferred embodiment of the present invention, determining the topics of each category based on the weight of the core words further includes:

[0022] The weight of the core words in each category is used to determine the weight of the short sentence in each category.

[0023] According to a preferred embodiment of the present invention, the short sentences are sorted in each category according to their weights, and the theme of each category is determined according to the top three short sentences in the category.

[0024] A second aspect of the present invention provides a data processing device for mining user needs, comprising:

[0025] The dialogue content expansion module is used to locate the sentences expressing negative intentions in the user's dialogue content and expand the dialogue content based on the sentences;

[0026] The cluster analysis module is used to perform cluster analysis on the expanded conversation content and determine the weight of the core words in each category;

[0027] The topic determination module is used to determine the topic of each category, that is, user needs, based on the weight of the core words.

[0028] According to a preferred embodiment of the present invention, the sentences expressing negative intention in the user's conversation content are specifically located as follows:

[0029] Use manual indexing and / or intent recognition model indexing.

[0030] According to a preferred embodiment of the present invention, the expansion of the conversation content centered on the sentence is specifically as follows:

[0031] Centered on the indexed sentence expressing negative intention, the dialogue is extended forward by N rounds and backward by M rounds, where N and M are natural numbers.

[0032] According to a preferred embodiment of the present invention, cluster analysis is performed on the extended conversation content, specifically:

[0033] Split the extended dialogue into short sentences according to punctuation marks;

[0034] A cluster analysis model is used to perform cluster analysis on short sentences, and the cluster analysis model adopts the Biterm TopicModel algorithm.

[0035] According to a preferred embodiment of the present invention, the weight of the core words in each category is determined as follows:

[0036] The Biterm Topic Model algorithm is used to determine the weight of words in each category, and the first i words are selected as core words, where i is a natural number.

[0037] According to a preferred embodiment of the present invention, determining the topics of each category based on the weight of the core words further includes:

[0038] The weight of the core words in each category is used to determine the weight of the short sentence in each category.

[0039] According to a preferred embodiment of the present invention, the short sentences are sorted in each category according to their weights, and the theme of each category is determined according to the top three short sentences in the category.

[0040] A third aspect of the present invention provides an electronic device, comprising a processor and a memory, wherein the memory is used to store a computer executable program, and when the computer program is executed by the processor, the processor executes the described method.

[0041] A fourth aspect of the present invention further proposes a computer-readable medium storing a computer-executable program, wherein when the computer-executable program is executed, the method described is implemented.

[0042] (III) Beneficial effects

[0043] The present invention locates the approximate position of user needs, improves the quality and concentration of alternative sentences through cluster analysis, and uses core word weights to sort within each category, thereby improving the efficiency and effect of subsequent summarization work and making the mining of user needs more systematic. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 It is a schematic flowchart of a data processing method for mining user requirements according to an embodiment of the present invention;

[0045] Figure 2 It is a schematic structural diagram of a data processing device for mining user requirements according to an embodiment of the present invention;

[0046] Figure 3 It is a schematic structural diagram of an electronic device according to an embodiment of the present invention;

[0047] Figure 4 It is a schematic diagram of a computer-readable recording medium according to an embodiment of the present invention. Detailed implementation manners

[0048] During the introduction of specific embodiments, the detailed description of structures, performances, effects or other features is to enable those skilled in the art to fully understand the embodiments. However, it does not exclude that those skilled in the art can implement the present invention with technical solutions that do not contain the above-mentioned structures, performances, effects or other features under specific circumstances.

[0049] The flowcharts in the accompanying drawings are only exemplary flow demonstrations, and do not represent that all the contents, operations and steps in the flowcharts must be included in the solutions of the present invention, nor does it represent that they must be executed in the order shown in the figures. For example, some operations / steps in the flowchart can be decomposed, some operations / steps can be combined or partially combined, etc. Without departing from the gist of the present invention, the execution order shown in the flowchart can be changed according to the actual situation.

[0050] The blocks in the accompanying drawings Figure 1 Generally represent functional entities, and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processing unit devices and / or microcontroller devices.

[0051] The same reference numerals in the drawings represent the same or similar elements, components or parts. Therefore, the repeated description of the same or similar elements, components or parts may be omitted hereinafter. It should also be understood that although the first, second, third, etc. attributives indicating numbers may be used in this article to describe various devices, elements, components or parts, these devices, elements, components or parts should not be limited by these attributives. That is, these attributives are only used to distinguish one from another. For example, the first device can also be called the second device without departing from the essence of the technical solution of the present invention. In addition, the terms "and / or", "or / and" mean all combinations including any one or more of the listed items.

[0052] To solve the above technical problems, the present invention proposes a data processing method for mining user needs. The method flow chart is as follows Figure 1 shown, including:

[0053] S101. Locate the statement expressing negative willingness in the user's conversation content, and expand the conversation content centered on this statement.

[0054] In this embodiment, the online promotion activity is the promotion of online education services. With the development of Internet technology, online education services have also flourished. Users can freely choose learning content, learning time, and learning location, making the acquisition of knowledge more flexible and diverse. In other embodiments, it can be other online promotion activities, and only examples are given here.

[0055] Based on the above technical solutions, further, the specific method for locating the statement expressing negative willingness in the user's conversation content is:

[0056] Adopt the method of manual indexing and / or indexing by an intention recognition model.

[0057] In this embodiment, the business personnel in the customer service center usually communicate with users by phone or voice. During the communication process, the business personnel can mark the words that the user clearly refuses. For example, when the user answers "No, no, we won't sign up for now", the business personnel can mark this sentence in the system to facilitate the subsequent expansion of the conversation content.

[0058] In other embodiments, it is also possible to identify the conversation content through an intention recognition model after the business personnel communicate with the user, and find out the statements expressing negative willingness of the user. First, convert the communication content between the business personnel and the user into text content, and then input the converted text content into the intention recognition model for judgment by the intention recognition model. The intention recognition model is a TextCNN convolutional neural network model based on deep learning. The intention recognition model can be trained using the communication history records between the business personnel and the user.

[0059] In other embodiments, it is also possible to adopt a combination of manual indexing and indexing by an intention recognition model. First, the intention recognition model preliminarily indexes the statements expressing negative willingness of the user, and then the business personnel manually perform secondary indexing on the results of the preliminary indexing to improve the accuracy and avoid errors in indexing by the intention recognition model.

[0060] Based on the above technical solutions, further, the specific method for expanding the conversation content centered on this statement is:

[0061] Expand N rounds of conversations forward and M rounds of conversations backward centered on the indexed statement expressing negative willingness, where N and M are natural numbers.

[0062] During the communication between business personnel and users, users usually state their reasons in the early part of the conversation. Therefore, in this embodiment, N > M, and usually it is not too far from the statement expressing the negative intention. So the value range of N is from 2 rounds to 6 rounds, usually set to 3, and the value range of M is from 1 round to 5 rounds, usually set to 2.

[0063] In this embodiment, a time interval threshold is also set. If the time interval between the statement of the extended conversation content and the statement expressing the negative intention exceeds the time interval threshold, the possibility of an association between the two is very small. Therefore, this statement can be discarded. The time interval threshold and the number of rounds of conversation extension ensure the degree of association between the extended conversation content and the user's negative intention.

[0064] S102. Perform clustering analysis on the extended conversation content to determine the weights of the core words in each category.

[0065] On the basis of the above technical solution, further, the clustering analysis of the extended conversation content is specifically as follows:

[0066] Segment the extended conversation content into short sentences according to punctuation marks;

[0067] Use a clustering analysis model to perform clustering analysis on the short sentences. The clustering analysis model adopts the Biterm TopicModel algorithm.

[0068] In this embodiment, after the conversation content is extended, the conversation content is segmented into short sentences based on punctuation marks. Compared with a complete long sentence, the semantic content of each short sentence is more pure and more representative. Before performing clustering analysis on the short sentences, word segmentation processing is also performed on the short sentences. After dividing the short sentences into individual words, vectorization processing is performed on the words. The vectorization of word text means using digital features to represent text because computers cannot directly understand the language and characters created by humans. To enable computers to understand text, we need to map the text information into a numerical semantic space, which we can call the word vector space. There are various algorithms for converting text into vectors, such as TF-IDF, BOW, One-Hot, word2vec, etc. In this embodiment, the vectorization of text adopts the word2vec algorithm. The word2vec model is an unsupervised learning model, and the mapping of text information to the semantic space can be achieved through the training of an unlabeled corpus.

[0069] In this embodiment, the text is split into individual Chinese characters and converted into vectors according to the word2vec model. In other embodiments, a semantic vector library can be preset, all Chinese characters are converted into vectors in advance for storage, and when in use, the vectors corresponding to the Chinese characters are directly selected from the semantic vector library.

[0070] In this embodiment, when the amount of content to be mined in the communication between business personnel and users accumulates to a certain quantity, the vector values of the words contained in the short sentences included in this content are input into the clustering analysis model. The clustering analysis model classifies these short sentences, such as into category A, category B, category C, category D, and so on. In this embodiment, it is also sorted according to the number of short sentences included in each classification, and the top 2 to 5 classifications are selected and retained, and the remaining classifications are discarded. For example, the top 3 classifications are selected, namely category A, category B, and category C. Or a short sentence quantity threshold is set, and those classifications with the number of short sentences less than the short sentence quantity threshold are discarded.

[0071] Based on the above technical solution, further, the specific method for determining the weights of the core words in each classification is as follows:

[0072] The Biterm Topic Model algorithm is used to determine the weights of the words in each classification, and the top i words are selected as the core words, where i is a natural number.

[0073] In this embodiment, during the process of the clustering analysis model clustering the short sentences, the Biterm Topic Model algorithm calculates the weights of the words included in each classification under that classification. Similarly, the top i words in each classification are selected as the core words. Usually, the value range of i is 2 to 5. For example, if i is 3, then the core words in category A are a, b, and c, and the weights are 0.24, 0.11, and 0.04 respectively; the core words in category B are x, y, and z, and the weights are 0.31, 0.19, and 0.03 respectively; the core words in category C are α, β, and γ, and the weights are 0.15, 0.11, and 0.01 respectively.

[0074] S103. Determine the theme of each classification based on the weights of the core words, that is, the user requirements.

[0075] Based on the above technical solution, further, determining the theme of each classification based on the weights of the core words further includes:

[0076] Use the weights of the core words in each classification to determine the weights of the short sentences in each classification.

[0077] Based on the above technical solution, further, sort each classification according to the weights of the short sentences in each classification, and determine the theme of the classification according to the top three short sentences in each classification.

[0078] In this embodiment, the weights of short sentences in each category are determined according to the core words in each category. For example, if a sentence contains three core words, a, c, and y, after calculation, the weight of this sentence in category A is 0.28, the weight in category B is 0.19, and the weight in category C is 0. Then, in each category, the short sentences are sorted according to their weights, and the theme of the category is determined based on the top three short sentences. Since the meaning of short sentences is relatively pure, the top three short sentences can summarize the general meaning of the category. In this embodiment, the first-ranked short sentence is directly selected as the theme of the category. For example, if the first-ranked short sentence in category A is "There is a conflict in class time", then this sentence is directly selected as the theme of category A, that is, the user's demand is to solve the conflict in class time, and subsequently, the user's demand can be solved by adjusting the class time or providing courses at other times for the user. In other embodiments, the business personnel can also summarize and determine the theme of the category based on the top three short sentences.

[0079] Figure 2 This is a data processing device 200 for mining user needs in an embodiment of the present invention, including:

[0080] A conversation content expansion module 201, configured to locate the sentences expressing negative intentions in the user's conversation content, and expand the conversation content centered on this sentence.

[0081] In this embodiment, the online promotion activity is the promotion of online education services. With the development of Internet technology, online education services have also flourished. Users can freely choose learning content, learning time, and learning location, making the acquisition of knowledge more flexible and diverse. In other embodiments, it can be other online promotion activities, and only examples are given here.

[0082] Based on the above technical solution, further, the specific method for locating the sentences expressing negative intentions in the user's conversation content is:

[0083] Adopt the method of manual indexing and / or intention recognition model indexing.

[0084] In this embodiment, the business personnel in the customer service center usually communicate with users by phone or voice. During the communication process, the business personnel can mark the words that the user clearly refuses. For example, if the user answers "No, no, we won't sign up for now", the business personnel can mark this sentence in the system to facilitate the subsequent expansion of the conversation content.

[0085] In other implementations, the intent recognition model can also be used to recognize the content of the conversation after the business personnel communicate with the user to find out the sentences in which the user clearly expresses the negative intention. First, the content of the communication between the business personnel and the user is converted into text content, and then the converted text content is input into the intent recognition model, and the intent recognition model makes a judgment. The intent recognition model is a TextCNN convolutional neural network model based on deep learning. The intent recognition model can be trained using the communication history between the business personnel and the user.

[0086] In other implementations, a combination of manual indexing and intent recognition model indexing may be used. The intent recognition model may first preliminarily index the user's sentences expressing negative intentions, and then business personnel may manually perform a secondary indexing on the results of the preliminary indexing to improve accuracy and avoid errors in intent recognition model indexing.

[0087] On the basis of the above technical solution, further, the conversation content is expanded with the sentence as the center as follows:

[0088] Centered on the indexed sentence expressing negative intention, the dialogue is extended forward by N rounds and backward by M rounds, where N and M are natural numbers.

[0089] In the process of communication between business personnel and users, users usually put the reasons in the previous dialogue, so in this implementation, N>M, and it is usually not too far away from the statement expressing negative intention, so the value range of N is 2 to 6 rounds, usually set to 3, and the value range of M is 1 to 5 rounds, usually set to 2.

[0090] In this embodiment, a time interval threshold is also set. If the time interval between the sentence of the extended conversation content and the sentence expressing the negative intention exceeds the time interval threshold, the possibility of the association between the two is very small, so the sentence can be discarded. The degree of association between the extended conversation content and the user's negative intention is guaranteed by the two aspects of the time interval threshold and the conversation extension round.

[0091] The cluster analysis module 202 is used to perform cluster analysis on the expanded conversation content to determine the weight of the core words in each category.

[0092] On the basis of the above technical solution, further, cluster analysis is performed on the extended conversation content, specifically:

[0093] Split the extended dialogue into short sentences according to punctuation marks;

[0094] A cluster analysis model is used to perform cluster analysis on short sentences, and the cluster analysis model adopts the Biterm TopicModel algorithm.

[0095] In this embodiment, after the dialogue content is expanded, the dialogue content is separated into short sentences based on punctuation marks. Compared with complete long sentences, the semantic content of each short sentence is more pure and more representative. Before performing clustering analysis on the short sentences, word segmentation processing is also performed on the short sentences, that is, the short sentences are divided into individual words, and then the words are vectorized. The vectorization of word text, that is, using digital features to represent text, because computers cannot directly understand the languages and characters created by humans. In order to enable computers to understand text, we need to map the text information to a numerical semantic space, which we can call the word vector space. There are various algorithms for converting text into vectors, such as TF-IDF, BOW, One-Hot, word2vec, etc. In this embodiment, the vectorization of text adopts the word2vec algorithm. The word2vec model is an unsupervised learning model, and the mapping of text information to the semantic space can be achieved by training with an unindexed corpus.

[0096] In this embodiment, the text is split into individual Chinese characters and converted into vectors according to the word2vec model. In other embodiments, a semantic vector library can be preset, and all Chinese characters are converted into vectors in advance for storage. When in use, the vectors corresponding to the Chinese characters are directly selected from the semantic vector library.

[0097] In this embodiment, when the content to be mined from the communication between business personnel and users accumulates to a certain quantity, the vector values of the words in the short sentences included in this content are input into the clustering analysis model. The clustering analysis model classifies these short sentences, such as into category A, category B, category C, category D, etc. In this embodiment, sorting is also performed according to the number of short sentences included in each classification, and the top 2 to 5 classifications are selected and retained, and the remaining classifications are discarded. For example, the top 3 categories are selected, namely category A, category B, and category C. Or a short sentence quantity threshold is set, and the classifications with the number of short sentences less than the short sentence quantity threshold are discarded.

[0098] On the basis of the above technical solutions, further, the specific method for determining the weights of the core words in each classification is as follows:

[0099] The Biterm Topic Model algorithm is used to determine the weights of the words in each classification, and the first i words are selected as the core words, where i is a natural number.

[0100] In this embodiment, during the process of clustering short sentences by the clustering analysis model, the Biterm Topic Model algorithm calculates the weights of the words included in each classification under that classification. Similarly, the first i words are selected as the core words for each classification, and generally, the value range of i is from 2 to 5. For example, when i is 3, the core words in class A are a, b, and c, with weights of 0.24, 0.11, and 0.04 respectively; the core words in class B are x, y, and z, with weights of 0.31, 0.19, and 0.03 respectively; the core words in class C are α, β, and γ, with weights of 0.15, 0.11, and 0.01 respectively.

[0101] The theme determination module 203 is used to determine the theme of each classification, that is, the user requirement, based on the weights of the core words.

[0102] On the basis of the above technical solution, further, determining the theme of each classification based on the weights of the core words further includes:

[0103] Use the weights of the core words in each classification to determine the weights of the short sentences in each classification.

[0104] On the basis of the above technical solution, further, sort the short sentences in each classification according to the weights of the short sentences in each classification, and determine the theme of the classification according to the top three short sentences in each classification.

[0105] In this embodiment, the weights of the short sentences in each classification are determined according to the core words in each classification. For example, a certain sentence contains three core words, a, c, and y. After calculation, the weight of this sentence in class A is 0.28, the weight in class B is 0.19, and the weight in class C is 0. Then, the short sentences are sorted according to their weights in each classification, and the theme of the classification is determined according to the top three short sentences. Since the meaning of the short sentences is relatively pure, the top three short sentences can summarize the general meaning of the classification. In this embodiment, the first-ranked short sentence is directly selected as the theme of the classification. For example, if the first-ranked short sentence in class A is "There is a conflict in class time", then this sentence is directly selected as the theme of class A, that is, the user requirement is to solve the conflict in class time. Subsequently, the user requirement can be solved by adjusting the class time or providing courses at other times for the user. In other embodiments, the business personnel can also summarize and determine the theme of the classification based on the top three short sentences.

[0106] The process of mining user requirements is described below through Example 1.

[0107] Example 1:

[0108] The business personnel communicate and contact with the user through the customer service center. During the communication process, there is a round of conversation content between the business personnel and the user as follows.

[0109] Business staff: The summer and autumn courses will start soon. It's really a great deal if you sign up now. Won't you sign up?

[0110] User: Let's discuss it further.

[0111] The business staff marks the user's sentence "Let's discuss it further." in the system as a statement indicating a negative intention.

[0112] Based on the content of this conversation, expand it. The expansion range is to expand 2 rounds of conversations forward and 1 round of conversation backward.

[0113] Business staff: Hello, is this Ms. Wang?

[0114] User: This is me.

[0115] Business staff: We have an online math course for the summer and autumn now. There are a total of XX classes. Now it only costs XXX yuan after the discount. Would you like to sign up one for your child?

[0116] User: Is it an online course? I'm not considering signing up for an online course for the time being.

[0117] Business staff: The summer and autumn courses will start soon. It's really a great deal if you sign up now. Won't you sign up?

[0118] User: Let's discuss it further. (Mark: Negative intention)

[0119] Business staff: Okay, sorry to trouble you.

[0120] User: It's okay. Goodbye.

[0121] Segment the user's conversation content in the above conversation by punctuation marks. The expanded content is split into short sentences such as "This is me.", "Is it an online course?", "I'm not considering signing up for an online course for the time being.", "Let's discuss it further.", "It's okay." and "Goodbye".

[0122] Then vectorize these short sentences. During the vectorization process, use the internal word list to segment these short sentences into individual words, and finally use the word2vec algorithm to convert them into vectors.

[0123] When a certain amount of conversation data from different users has been accumulated, convert all these conversation data into short sentences and perform vectorization processing. Input these vectorized short sentences into the clustering analysis model for clustering analysis.

[0124] After clustering analysis, the short sentences are divided into multiple categories such as category 0, category 1, category 2, category 3, etc. Among them, the number of sentences in these three categories, namely category 0, category 1, and category 2, is relatively large and they are retained, while other categories are discarded. The clustering analysis model also calculates the weights of the keywords in each category using the Biterm Topic Model algorithm.

[0125] Among them, in category 0, the weight of the keyword "temporarily" is 0.122, the weight of "not" is 0.087, the weight of "sign up" is 0.074, and the weight of "online" is "0.213".

[0126] The top 3 sentences in category 0 are respectively:

[0127] "Temporarily do not consider signing up for online courses for now", with a weight of 0.496;

[0128] "Do not consider signing up online", with a weight of 0.374;

[0129] "Do not sign up for the summer and autumn courses temporarily for now", with a weight of 0.283.

[0130] Therefore, the theme of category 0 can be summarized as "temporarily do not sign up for online courses".

[0131] In category 1, the weight of the keyword "temporarily" is 0, the weight of "not" is 0.014, the weight of "sign up" is 0.007, and the weight of "online" is 0. The weight is 0 because "temporarily" and "online" do not appear within the top 40 of the topic word weights. Therefore, in category 1, the weight of "temporarily do not consider signing up for online courses for now" is 0.021.

[0132] In category 2, the weights of "temporarily", "not", "sign up", and "online" are all 0. Therefore, in category 2, the weight of "temporarily do not consider signing up for online courses for now" is 0.

[0133] Therefore, it can be determined that "temporarily do not consider signing up for online courses for now" belongs to category 0, and the corresponding user's demand for this sentence is to temporarily not sign up for online courses. Business personnel can adjust their strategies based on this result and recommend offline courses to users.

[0134] Figure 3 It is a schematic structural diagram of an electronic device according to an embodiment of the present invention. The electronic device includes a processor and a memory. The memory is used to store computer-executable programs. When the computer program is executed by the processor, the processor executes a vehicle intelligent assisted pushing method based on rotation angle monitoring.

[0135] Such as Figure 3As shown, the electronic device is presented in the form of a general computing device. The processor can be one or multiple and work collaboratively. The present invention does not exclude distributed processing, that is, the processors can be dispersed in different physical devices. The electronic device of the present invention is not limited to a single entity and can also be the sum of multiple physical devices.

[0136] The memory stores computer-executable programs, usually machine-readable codes. The computer-readable programs can be executed by the processor so that the electronic device can execute the method of the present invention or at least some of the steps in the method.

[0137] The memory includes volatile memory, such as random access storage units (RAM) and / or cache storage units, and can also be non-volatile memory, such as read-only storage units (ROM).

[0138] Optionally, in this embodiment, the electronic device further includes an I / O interface, which is used for the electronic device to exchange data with external devices. The I / O interface can represent one or more of several bus structures, including a memory unit bus or a memory unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the multiple bus structures.

[0139] It should be understood that Figure 3 The shown electronic device is only an example of the present invention. The electronic device of the present invention may also include elements or components not shown in the above example. For example, some electronic devices also include a display unit such as a display screen, and some electronic devices also include human-computer interaction elements, such as buttons, keyboards, etc. As long as the electronic device can execute the computer-readable program in the memory to implement the method of the present invention or at least some of the steps in the method, it can be considered as the electronic device covered by the present invention.

[0140] Figure 4 is a schematic diagram of a computer-readable recording medium of an embodiment of the present invention. As Figure 4As shown, a computer-readable recording medium stores a computer-executable program. When the computer-executable program is executed, the above-mentioned vehicle intelligent assisted propulsion method based on rotation angle monitoring of the present invention is implemented. The computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium may also be any readable medium other than the readable storage medium, and this readable medium may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the above.

[0141] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, executed as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).

[0142] From the above description of the embodiments, those skilled in the art can easily understand that the present invention can be implemented by hardware capable of executing specific computer programs, such as the system of the present invention, as well as the electronic processing units, servers, clients, mobile phones, control units, processors, etc. included in the system. The present invention can also be implemented by a vehicle including at least a part of the above system or components. The present invention can also be implemented by computer software for executing the method of the present invention, such as control software executed by a microprocessor or an electronic control unit at the locomotive end, a client, a server end, etc. However, it should be noted that the computer software for executing the method of the present invention is not limited to being executed in one or a specific hardware entity. It can also be implemented in a distributed manner by unspecified specific hardware. For example, some method steps executed by a computer program can be executed at the locomotive end, and another part can be executed in a mobile terminal or a smart helmet, etc. For computer software, the software product can be stored in a computer-readable storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), or can be distributed and stored on a network, as long as it can enable an electronic device to execute the method according to the present invention.

[0143] The specific embodiments described above further elaborate on the object, technical solution, and beneficial effects of the present invention. It should be understood that the present invention is not inherently related to any specific computer, virtual device, or electronic device, and various general-purpose devices can also implement the present invention. The above description is only for the specific embodiments of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A data processing method for mining user needs, characterized in that, The method includes: Locating the sentences expressing negative willingness in the user's conversation content by means of manual indexing and / or indexing by an intention recognition model, expanding the conversation content centered on the indexed sentences expressing negative willingness, and ensuring the degree of association between the expanded conversation content and the user's negative willingness by setting a time interval threshold and the number of conversation expansion rounds, including: expanding N rounds of conversation forward and M rounds of conversation backward, where N and M are natural numbers and N > M, the value range of N is from 2 rounds to 6 rounds, and the value range of M is from 1 round to 5 rounds; and setting a time interval threshold, and discarding the sentences expressing negative willingness if the time interval between the sentences of the expanded conversation content and the sentences expressing negative willingness exceeds the time interval threshold; Segmenting the expanded conversation content into short sentences based on punctuation marks, performing clustering analysis on the short sentences, and determining the weights of the core words in each category, including: dividing the short sentences into individual words and vectorizing each word; inputting the vector values of the words contained in the short sentences into a clustering analysis model and classifying these short sentences using the clustering analysis model; sorting according to the number of short sentences contained in each category, and retaining the categories with the top rankings; calculating the weights of the words contained in each retained category under the corresponding category and selecting the core words; where before calculating the weights of the words contained in each category under the corresponding category, it includes: sorting according to the number of short sentences contained in each category and selecting and retaining the categories with the top rankings; Determining the theme of each category, that is, the user's needs, based on the weights of the core words, including: using the weights of the core words selected by the words contained in each category to determine the weights of the short sentences in each category, and sorting in each category according to the determined weights of the short sentences in each category to determine the theme of each category.

2. The method according to claim 1, characterized in that, The specific method of locating the sentences expressing negative willingness in the user's conversation content by means of manual indexing and / or indexing by an intention recognition model is: Marking the words clearly refused by the user in the user's conversation content by means of manual indexing.

3. The method according to claim 1, wherein The specific method of locating the sentences expressing negative willingness in the user's conversation content by means of manual indexing and / or indexing by an intention recognition model is: Identifying the sentences in which the user clearly expresses negative willingness in the user's conversation content by using an intention recognition model and marking them.

4. The method according to claim 1, characterized in that The specific method of locating the sentences expressing negative willingness in the user's conversation content by means of manual indexing and / or indexing by an intention recognition model is: First, identifying the sentences in which the user clearly expresses negative willingness in the user's conversation content by using an intention recognition model for preliminary indexing, and then performing secondary indexing on the results of the preliminary indexing by means of manual indexing.

5. The method according to any one of claims 1-4, wherein the clustering analysis model adopts the Biterm Topic Model algorithm; and / or, determining the weights of the core words in each category includes: using the Biterm Topic Model algorithm to determine the weights of the words in each category, and selecting the first i words as the core words, where i is a natural number.

6. The method according to claim 1, characterized in that, Before calculating the weights of the words included in each category under the corresponding category, it also includes: Discarding the category according to whether the number of short sentences included in each category is less than the set short sentence number threshold.

7. The method according to claim 1, wherein Determining the theme of each category by sorting according to the weights of the short sentences in each category respectively in each category also includes: Determining the theme of the category according to the top three short sentences in each category.

8. A data processing device for mining user requirements, characterized in that, The device includes: A dialogue content expansion module, which is used to locate the sentences expressing negative willingness in the user's dialogue content by means of manual indexing and / or indexing by an intention recognition model, and expand the dialogue content centered on the indexed sentences expressing negative willingness. The dialogue content expansion is ensured by setting a time interval threshold and the number of dialogue expansion rounds. It includes: expanding the dialogue forward by N rounds and backward by M rounds. The number of dialogue expansion rounds N and M are natural numbers and N>M. The value range of N is from 2 rounds to 6 rounds, and the value range of M is from 1 round to 5 rounds; and, setting a time interval threshold, and discarding the sentences expressing negative willingness if the time interval between the sentences of the expanded dialogue content and the sentences expressing negative willingness exceeds the time interval threshold; A clustering analysis module, which is used to split the expanded dialogue content into short sentences based on punctuation marks, perform clustering analysis on the short sentences, and determine the weights of the core words in each category, including: dividing the short sentences into individual words and vectorizing each word; inputting the vector values of the words contained in the short sentences into a clustering analysis model and classifying these short sentences with the clustering analysis model; sorting according to the number of short sentences included in each category, and selecting the top-ranked categories to retain; calculating the weights of the words included in each retained category under the corresponding category and selecting the core words; among them, before calculating the weights of the words included in each category under the corresponding category, it includes: sorting according to the number of short sentences included in each category and selecting and retaining the top-ranked categories; A theme determination module, which is used to determine the theme of each category, that is, the user's needs, based on the weights of the core words, including: using the weights of the core words selected from the words included in each category to determine the weights of the short sentences in each category, and determining the theme of each category by sorting according to the determined weights of the short sentences in each category respectively in each category.

9. The device according to claim 8, characterized in that, Locating the sentences expressing negative willingness in the user's dialogue content by means of manual indexing and / or indexing by an intention recognition model specifically is: Marking the words clearly refused by the user in the user's dialogue content by means of manual indexing.

10. The device according to claim 8, characterized in that, Locating the sentences expressing negative willingness in the user's dialogue content by means of manual indexing and / or indexing by an intention recognition model specifically is: Identifying the sentences clearly expressing negative willingness in the user's dialogue content by using an intention recognition model and marking them.

11. The device according to claim 8, wherein, Locating the sentences expressing negative willingness in the user's dialogue content by means of manual indexing and / or indexing by an intention recognition model specifically is: First, the intention recognition model is used to identify the user's conversation content to find out the sentences in which the user clearly expresses negative intentions and perform preliminary indexing. Then, manual indexing is used to perform secondary indexing on the preliminary indexing results.

12. The device according to any one of claims 8-11, characterized in that, The cluster analysis model adopts the BitermTopic Model algorithm, and / or determining the weight of the core word in each category includes: using the Biterm Topic Model algorithm to determine the weight of the words in each category, selecting the first i words as the core words, where i is a natural number.

13. The device according to claim 8, characterized in that, Before calculating the weight of the words included in each category under the corresponding category, the method further includes: discarding the category according to whether the number of short sentences included in each category is less than a set short sentence number threshold.

14. The device according to claim 8, characterized in that According to the weight of the determined short sentences in each category, the topics of each category are sorted in each category and determined, and the following are also included: The top three sentences in each category were used to determine the theme of that category.

15. An electronic device comprising a processor and a memory, wherein the memory is used to store a computer executable program, wherein: When the computer executable program is executed by the processor, the processor performs the method according to any one of claims 1 to 7.

16. A computer-readable medium stores a computer-executable program, characterized in that, When the computer executable program is executed, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Dynamic user attribute extraction method based on social media

    CN106354818A

  • Cl Simple sentence intention recognition method, device and system based on classification

    CN111191030A