A method, apparatus, storage medium and electronic device for determining a conversation topic
By clustering and topic analysis of the conversation text of internal communication tools in the enterprise, the problem that senior decision makers cannot accurately determine the conversation topic is solved, and a more efficient workflow is achieved.
Patent Information
- Application Number
- CN202110458206.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-27
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2041-04-27
AI Technical Summary
In the internal communication tools of the enterprise, high-level decision makers are unable to accurately, comprehensively and effectively determine the conversation topic, resulting in less efficient work.
The preset algorithm is used to calculate the conversation text, and clustering results are generated and the conversation topic is determined by dividing time periods, determining keyword frequency, combining similarity and continuity analysis.
Accurate and comprehensive identification of conversation topics improves work efficiency.
Smart Images

Figure CN113177418B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a method, device, storage medium and electronic device for determining a conversation topic. Background Art
[0002] As the functions of instant messaging tools (such as applications developed for communication, etc.) become increasingly complete and their operations become more convenient, they appear more and more in people's lives. At the same time, instant messaging tools are gradually replacing traditional telephone and email communication methods with their advantages of being cheaper, more convenient and more efficient.
[0003] Currently, many companies rely on internal communication tools to streamline management and facilitate more efficient and convenient internal communication among employees. These tools can be used to allow multiple employees to communicate simultaneously within the same communication environment. However, this generates a large volume of text, making it difficult for senior decision-makers to accurately, comprehensively, and effectively determine the topic of conversations, resulting in low work efficiency. Summary of the Invention
[0004] In view of this, the embodiments of the present application propose a method, device, storage medium and electronic device for determining a conversation topic to solve the problem of low work efficiency in the prior art.
[0005] In a first aspect, an embodiment of the present application provides a method for determining a conversation topic, which includes:
[0006] Get multiple conversation texts;
[0007] Calculating the plurality of conversation texts using a preset algorithm to obtain at least one clustering result;
[0008] A conversation topic is determined based on each clustering result.
[0009] In a possible implementation, the using a preset algorithm to calculate the plurality of conversation texts to obtain at least one clustering result includes:
[0010] Dividing the time periods corresponding to the plurality of conversation texts into a plurality of sub-time periods in chronological order;
[0011] determining the frequency of occurrence of the keyword in each of the sub-time periods;
[0012] Determine at least one conversation interval based on each of the occurrence frequencies, a first preset frequency, and a second preset frequency; wherein the first preset frequency is greater than the second preset frequency;
[0013] A clustering result is determined for each session interval.
[0014] In a possible implementation, determining at least one conversation interval based on each of the occurrence frequencies, the first preset frequency, and the second preset frequency includes:
[0015] Determine, in chronological order, a first occurrence frequency greater than or equal to the first preset frequency and a second occurrence frequency less than or equal to the second preset frequency;
[0016] A conversation interval is determined based on conversation texts in a plurality of sub-time periods determined based on the first occurrence frequency of the adjacent first occurrence and the second occurrence frequency of the adjacent last occurrence.
[0017] In a possible implementation, the method for determining a conversation topic further includes:
[0018] Parsing N end-of-session texts in a first session interval and M start-of-session texts in a second session interval of two adjacent session intervals;
[0019] When it is determined that the continuity between the N end conversation texts in the previous conversation interval and the M start conversation texts in the next conversation interval meets a preset condition, the two adjacent conversation intervals are merged into one conversation interval.
[0020] In a possible implementation, the method for determining a conversation topic further includes:
[0021] Calculating the similarity between the N end-of-conversation texts in the first of two adjacent conversation intervals and the M start-of-conversation texts in the second of two adjacent conversation intervals;
[0022] When the similarity is greater than a preset threshold, the two adjacent conversation intervals are merged into one conversation interval.
[0023] In a possible implementation, determining a clustering result for each session interval includes:
[0024] For each of the conversation intervals, a clustering algorithm is used to calculate the conversation texts in the conversation interval to obtain a clustering result corresponding to the conversation interval.
[0025] In a possible implementation, determining a conversation topic based on each clustering result includes:
[0026] A conversation topic of a conversation section corresponding to the clustering result is determined based on a plurality of text words included in the clustering result.
[0027] In a second aspect, an embodiment of the present application further provides a device for determining a conversation topic, comprising:
[0028] An acquisition module configured to acquire a plurality of conversation texts;
[0029] a calculation module configured to calculate the plurality of conversation texts using a preset algorithm to obtain at least one clustering result;
[0030] A determination module is configured to determine a conversation topic based on each clustering result.
[0031] In a third aspect, the present disclosure further provides a storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are performed:
[0032] Get multiple conversation texts;
[0033] Calculating the plurality of conversation texts using a preset algorithm to obtain at least one clustering result;
[0034] A conversation topic is determined based on each clustering result.
[0035] In a fourth aspect, the present disclosure further provides an electronic device, comprising: a processor and a memory, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate via a bus, and when the machine-readable instructions are executed by the processor, the following steps are performed:
[0036] Get multiple conversation texts;
[0037] Calculating the plurality of conversation texts using a preset algorithm to obtain at least one clustering result;
[0038] A conversation topic is determined based on each clustering result.
[0039] The embodiment of the present application uses a preset algorithm to calculate multiple conversation texts to obtain at least one clustering result, and then modifies each clustering result based on keywords, semantic analysis and similarity calculation. Then, a conversation topic is determined based on each clustering result. It can accurately, comprehensively and effectively determine at least one conversation topic corresponding to multiple conversation texts, thereby improving work efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0041] Figure 1 A flowchart of a method for determining a conversation topic provided by the present application is shown;
[0042] Figure 2 A flowchart showing at least one clustering result obtained in a method for determining a conversation topic provided by the present application is shown;
[0043] Figure 3 A flowchart of determining a conversation interval in a conversation topic determination method provided by the present application is shown;
[0044] Figure 4 A flowchart of further determining a conversation interval in determining a conversation topic provided by the present application is shown;
[0045] Figure 5 A flowchart showing another method for determining a conversation topic provided by the present application, further determining a conversation interval;
[0046] Figure 6 A schematic diagram showing the structure of a device for determining a conversation topic provided by the present application is shown;
[0047] Figure 7 A schematic structural diagram of an electronic device provided in this application is shown. DETAILED DESCRIPTION
[0048] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions of the embodiments of the present application will be clearly and completely described below in conjunction with the drawings of the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the described embodiments of the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0049] Unless otherwise defined, the technical or scientific terms used in this application should have the usual meanings understood by people with ordinary skills in the field to which this application belongs. The "first", "second" and similar words used in this application do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative position relationships. When the absolute position of the object being described changes, the relative position relationship may also change accordingly.
[0050] In order to keep the following description of the embodiments of the present application clear and concise, detailed descriptions of known functions and known components are omitted in this application.
[0051] like Figure 1 The flowchart shown is a method for determining a conversation topic according to the first aspect of the present application. The method utilizes deep learning techniques and natural language processing techniques, such as semantic structure analysis and similarity calculation, to accurately determine the conversation topic contained in a conversation text, thereby improving work efficiency. Specifically, steps S101-S103 are described.
[0052] S101, obtaining multiple conversation texts.
[0053] Here, the conversation text is obtained from the current conversation tool, and the conversation text includes an ID that identifies the user's identity information, the time when the user sends the conversation text, and the text content of the conversation text sent by the user.
[0054] Specifically, the rules for obtaining conversation texts can be set according to actual needs. For example, if there are many conversation texts, some conversation texts can be obtained based on time, such as obtaining conversation texts from 10:00 am to 12:00 pm on the same day. If there are multiple members participating in a conversation, conversation texts can be obtained based on member, such as if the members participating in the conversation are member A, member B, member C, member D, and member E, and only the conversation texts sent by member B and member E are extracted. Of course, all conversation texts in the current conversation tool can also be obtained.
[0055] The backend server connected to the current conversation tool generates an acquisition instruction and sends the acquisition instruction to the processor corresponding to the current conversation tool, so that the processor returns the conversation text corresponding to the acquisition instruction to the server. The acquisition instruction instructs the processor corresponding to the current conversation tool to acquire multiple conversation texts according to the acquisition rule.
[0056] S102: Calculate the plurality of conversation texts using a preset algorithm to obtain at least one clustering result.
[0057] In a specific implementation, the preset algorithm can be a pre-trained clustering model or a pre-set calculation rule, which is not specifically limited in the present embodiment. The present embodiment takes the preset algorithm as a pre-set calculation rule as an example for detailed description.
[0058] Specifically, Figure 2 The specific method steps of calculating multiple conversation texts using a preset algorithm to obtain at least one clustering result are shown, including S201-S204.
[0059] S201 , dividing a time period corresponding to a plurality of conversation texts into a plurality of sub-time periods in chronological order.
[0060] S202: Determine the occurrence frequency of the keyword in each sub-time period.
[0061] S203: Determine at least one conversation interval based on each occurrence frequency, a first preset frequency, and a second preset frequency; wherein the first preset frequency is greater than the second preset frequency.
[0062] After acquiring multiple conversation texts, the time periods corresponding to the multiple conversation texts are determined. Specifically, the time interval between the time point corresponding to the first conversation text and the time point corresponding to the last conversation text is used as the time period corresponding to the multiple conversation texts. Further, the time period is divided into multiple sub-time periods according to the chronological order.
[0063] For example, determine the first and last conversation texts among multiple conversation texts, find that the time point when the first conversation text was sent is 9:00, and the time point when the last conversation text was sent is 12:00, then use 9:00-12:00 as the time period corresponding to the multiple conversation texts. Afterwards, you can divide the sub-time period from 9:00 to 12:00 with a length of five minutes as a sub-time period to obtain multiple sub-time periods.
[0064] Of course, the duration of each sub-time period is not limited to five minutes, and can be determined according to the amount of all conversation texts. For example, when there are more conversation texts, the duration of the sub-time period can be set to be shorter; when there are fewer conversation texts, the duration of the sub-time period can be set to be longer, etc.
[0065] In a specific implementation, the keywords may be pre-set or determined based on all conversation texts. Specifically, one or more words that appear frequently in all conversation texts are searched as keywords.
[0066] After determining the keyword, the occurrence frequency of the keyword is calculated for each sub-time period; then, at least one conversation interval is determined based on each occurrence frequency, a first preset frequency, and a second preset frequency; wherein the first preset frequency is greater than the second preset frequency.
[0067] Specifically, you can follow Figure 3 The method steps shown in the figure are for determining at least one session interval, specifically including S301 and S302.
[0068] S301 : determining, in chronological order, a first occurrence frequency that is greater than or equal to a first preset frequency and a second occurrence frequency that is less than or equal to a second preset frequency.
[0069] S302 : Determine a conversation interval based on conversation texts in a plurality of sub-time periods determined by adjacent first-appearance first frequency and last-appearance second frequency.
[0070] In the specific implementation, after calculating the frequency of occurrence of the keyword in each sub-time period, each occurrence frequency is compared with the first preset frequency and the second preset frequency respectively, and the frequency greater than or equal to the first preset frequency is determined to be the first occurrence frequency, and the frequency less than or equal to the second preset frequency is determined to be the second occurrence frequency, and all the first occurrence frequencies and second occurrence frequencies are arranged in chronological order.
[0071] The conversation texts within multiple sub-time periods determined by the adjacent first-occurrence first frequency and the last-occurrence second frequency are determined as a conversation interval. For example, the sub-time periods are A, B, C, D, E, F, G, H, and I, and A, B, C, D, E, F, G, H, and I are arranged in chronological order, where A, B, F, G correspond to the first-occurrence frequency, and D, E, and I correspond to the second-occurrence frequency. A is the first-occurrence first frequency, and the sub-time period corresponding to the last-occurrence second frequency adjacent to it is E. Therefore, a conversation interval is determined based on the conversation texts within A, B, C, D, and E. Similarly, a conversation interval is determined based on the conversation texts within F, G, H, and I.
[0072] Furthermore, considering that the conversation interval determined based on keywords may have certain errors, the embodiment of the present application also provides Figure 4 and Figure 5 The determination method shown is used to improve the accuracy of the determined session interval.
[0073] Figure 4 S401 and S402 in FIG. 4 show a determination method, specifically:
[0074] S401 , parsing N conversation ending texts in a former conversation interval and M conversation starting texts in a latter conversation interval of two adjacent conversation intervals.
[0075] S402: When it is determined that the continuity between the N end conversation texts in the previous conversation interval and the M start conversation texts in the next conversation interval meets a preset condition, the two adjacent conversation intervals are merged into one conversation interval.
[0076] After identifying multiple conversation segments based on keywords, the system extracts N end-of-session text from the first segment and M start-of-session text from the second segment for each adjacent segment and parses them. Specifically, the system analyzes the semantic structure of each conversation segment, including its part of speech, brevity, and fragmentation. N and M can be the same or different.
[0077] Specifically, when the parts of speech of the text content are demonstrative pronouns ("you", "I", "he"), conjunctions ("but", "and", etc.), and symbols ("@" and other information), it represents a high degree of correlation, that is, there is a certain degree of continuity; when the text content is relatively short, such as "OK", "received", etc., it can represent the end of the topic, that is, the continuity is low; the judgment rule for text content fragmentation is: if the text content does not have a predicate structure, it can represent the end of the current topic; if the text content has a predicate structure, it can represent the beginning of the current topic.
[0078] Based on the above parsing rules, if it is determined that the continuity between the N end-of-session texts in the previous session and the M start-of-session texts in the next session meets a preset condition, the two adjacent session intervals are merged into one session interval. The preset condition can be set as, for example, that at least half of the N end-of-session texts and the M start-of-session texts have a high degree of continuity.
[0079] Figure 5 S501 and S502 in FIG. 5 show another determination method, specifically:
[0080] S501 , calculating the similarity between N ending conversation texts in a former conversation interval and N starting conversation texts in a latter conversation interval in two adjacent conversation intervals.
[0081] S502: When the similarity is greater than a preset threshold, the two adjacent conversation intervals are merged into one conversation interval.
[0082] After determining multiple conversation intervals based on keywords or determining multiple conversation intervals by parsing the conversation text, the similarity between the N ending conversation texts in the first conversation interval and the N starting conversation texts in the second conversation interval of two adjacent conversation intervals can also be calculated using calculation methods such as cosine similarity semanticSim(A, B), Jaccard similarity coefficient jacSim(A, B) and the summed average of the edit distance levSim(A, B).
[0083] When the similarity is greater than a preset threshold, that is, it indicates that the conversation topics of the two adjacent conversation intervals are the same, the two adjacent conversation intervals are merged into one conversation interval.
[0084] Among them, you can only use Figure 4 To improve the accuracy of the determined session interval, you can also use only Figure 5 The determination method in can improve the accuracy of the determined session interval, and can also be used at the same time Figure 4 and Figure 5 The determination method is used to improve the accuracy of the determined session interval.
[0085] S204: Determine a clustering result for each conversation interval.
[0086] After determining multiple conversation intervals, a clustering algorithm is used to calculate the conversation text within each conversation interval to obtain a clustering result corresponding to the conversation interval. Preferably, a traditional K-means clustering algorithm is used to calculate the conversation text within each conversation interval. Of course, the traditional K-means clustering algorithm is only one embodiment, and any method that can achieve the clustering purpose will be sufficient.
[0087] S103: Determine a conversation topic based on each clustering result.
[0088] Here, each clustering result includes at least one or more keywords, and then the conversation topic of the conversation section corresponding to the clustering result is determined based on the multiple text words included in the clustering result.
[0089] The conversation topic can be pre-set, for example, a mapping relationship between keywords and conversation topics is pre-set. After the clustering results are determined, the conversation topic corresponding to the keyword in the clustering results is searched, and the conversation topic corresponding to the keyword is used as the conversation topic corresponding to the clustering result; the conversation can also be generated according to preset rules using the keywords in the clustering results, and the embodiments of the present application do not make specific limitations on this.
[0090] The embodiment of the present application uses a preset algorithm to calculate multiple conversation texts to obtain at least one clustering result, and then modifies each clustering result based on keywords, semantic analysis and similarity calculation. Then, a conversation topic is determined based on each clustering result. It can accurately, comprehensively and effectively determine at least one conversation topic corresponding to multiple conversation texts, thereby improving work efficiency.
[0091] Based on the same inventive concept, the second aspect of this application also provides a conversation topic determination device corresponding to the conversation topic determination method. Since the principle of solving the problem by the conversation topic determination device in this application is similar to the above-mentioned conversation topic determination method of this application, the implementation of the conversation topic determination device can refer to the implementation of the method, and the repeated parts will not be repeated.
[0092] Figure 6A schematic diagram of an electronic device provided in an embodiment of the present application is shown, specifically including:
[0093] An acquisition module 601 is configured to acquire a plurality of conversation texts;
[0094] A calculation module 602 is configured to calculate the plurality of conversation texts using a preset algorithm to obtain at least one clustering result;
[0095] The determination module 603 is configured to determine a conversation topic based on each clustering result.
[0096] In another embodiment, the calculation module 602 is specifically configured as follows:
[0097] Dividing the time periods corresponding to the plurality of conversation texts into a plurality of sub-time periods in chronological order;
[0098] determining the frequency of occurrence of the keyword in each of the sub-time periods;
[0099] Determine at least one conversation interval based on each of the occurrence frequencies, a first preset frequency, and a second preset frequency; wherein the first preset frequency is greater than the second preset frequency;
[0100] A clustering result is determined for each session interval.
[0101] In another embodiment, when the calculation module 602 determines at least one conversation interval based on each of the occurrence frequencies, the first preset frequency, and the second preset frequency, the specific configuration is as follows:
[0102] Determine, in chronological order, a first occurrence frequency greater than or equal to the first preset frequency and a second occurrence frequency less than or equal to the second preset frequency;
[0103] A conversation interval is determined based on conversation texts in a plurality of sub-time periods determined based on the first occurrence frequency of the adjacent first occurrence and the second occurrence frequency of the adjacent last occurrence.
[0104] In another embodiment, the conversation topic determination apparatus further includes a first merging module 604, which is configured to:
[0105] Parsing N end-of-session texts in a first session interval and M start-of-session texts in a second session interval of two adjacent session intervals;
[0106] When it is determined that the continuity between the N end conversation texts in the previous conversation interval and the M start conversation texts in the next conversation interval meets a preset condition, the two adjacent conversation intervals are merged into one conversation interval.
[0107] In another embodiment, the conversation topic determination apparatus further includes a second merging module 605, which is configured to:
[0108] Calculating the similarity between the N end-of-conversation texts in the first of two adjacent conversation intervals and the M start-of-conversation texts in the second of two adjacent conversation intervals;
[0109] When the similarity is greater than a preset threshold, the two adjacent conversation intervals are merged into one conversation interval.
[0110] In another embodiment, when the calculation module 602 determines a clustering result for each session interval, the specific configuration is as follows:
[0111] For each of the conversation intervals, a clustering algorithm is used to calculate the conversation texts in the conversation interval to obtain a clustering result corresponding to the conversation interval.
[0112] In yet another embodiment, the determining module 603 is specifically configured to:
[0113] A conversation topic of a conversation section corresponding to the clustering result is determined based on a plurality of text words included in the clustering result.
[0114] The embodiment of the present application uses a preset algorithm to calculate multiple conversation texts to obtain at least one clustering result, and then modifies each clustering result based on keywords, semantic analysis and similarity calculation. Then, a conversation topic is determined based on each clustering result. It can accurately, comprehensively and effectively determine at least one conversation topic corresponding to multiple conversation texts, thereby improving work efficiency.
[0115] An embodiment of the present application provides a storage medium, which is a computer-readable medium and stores a computer program. When the computer program is executed by a processor, the method provided in any embodiment of the present application is implemented, including the following steps S11 to S13:
[0116] S11, obtaining multiple conversation texts;
[0117] S12, calculating the plurality of conversation texts using a preset algorithm to obtain at least one clustering result;
[0118] S13, determining a conversation topic based on each clustering result.
[0119] When the computer program is executed by the processor to calculate the multiple conversation texts using a preset algorithm and obtain at least one clustering result, the processor specifically performs the following steps: dividing the time periods corresponding to the multiple conversation texts into multiple sub-time periods in chronological order; determining the frequency of occurrence of keywords in each of the sub-time periods; determining at least one conversation interval based on each of the occurrence frequencies, a first preset frequency, and a second preset frequency; wherein the first preset frequency is greater than the second preset frequency; and determining a clustering result for each of the conversation intervals.
[0120] When the computer program is executed by the processor to determine at least one conversation interval based on each of the occurrence frequencies, the first preset frequency, and the second preset frequency, the processor specifically performs the following steps: determining, in chronological order, a first occurrence frequency greater than or equal to the first preset frequency and a second occurrence frequency less than or equal to the second preset frequency; and determining a conversation interval based on conversation texts within multiple sub-time periods determined based on adjacent first occurrences of the first occurrence frequency and last occurrence of the second occurrence frequency.
[0121] When the computer program is executed by the processor to execute the conversation topic determination method, the processor also executes the following steps: parsing N end conversation texts in the first conversation interval and M start conversation texts in the second conversation interval of two adjacent conversation intervals; when it is determined that the continuity between the N end conversation texts in the first conversation interval and the M start conversation texts in the second conversation interval meets a preset condition, merging the two adjacent conversation intervals into one conversation interval.
[0122] When the computer program is executed by the processor to execute the conversation topic determination method, the processor also executes the following steps: calculating the similarity between the N end conversation texts in the first conversation interval and the M start conversation texts in the second conversation interval of two adjacent conversation intervals; when the similarity is greater than a preset threshold, merging the two adjacent conversation intervals into one conversation interval.
[0123] When the computer program is executed by the processor to determine a clustering result for each conversation interval, the processor further executes the following steps: for each conversation interval, using a clustering algorithm to calculate the conversation text in the conversation interval to obtain a clustering result corresponding to the conversation interval.
[0124] When the computer program is executed by the processor to determine a conversation topic based on each clustering result, the processor also executes the following step: determining a conversation topic of a conversation interval corresponding to the clustering result based on multiple text words included in the clustering result.
[0125] The embodiment of the present application uses a preset algorithm to calculate multiple conversation texts to obtain at least one clustering result, and then modifies each clustering result based on keywords, semantic analysis and similarity calculation. Then, a conversation topic is determined based on each clustering result. It can accurately, comprehensively and effectively determine at least one conversation topic corresponding to multiple conversation texts, thereby improving work efficiency.
[0126] An embodiment of the present application provides an electronic device. The structural diagram of the electronic device can be as follows: Figure 7 As shown, the electronic device comprises at least a memory 701 and a processor 702. The memory 701 stores a computer program. The processor 702 implements the method provided by any embodiment of the present disclosure when executing the computer program on the memory 701. Exemplarily, the steps of the electronic device computer program are as follows S21 to S23:
[0127] S21, obtaining multiple conversation texts;
[0128] S22, calculating the plurality of conversation texts using a preset algorithm to obtain at least one clustering result;
[0129] S23, determining a conversation topic based on each clustering result.
[0130] When the processor calculates the multiple conversation texts using a preset algorithm stored in the execution memory to obtain at least one clustering result, it also executes the following computer program: dividing the time periods corresponding to the multiple conversation texts into multiple sub-time periods in chronological order; determining the frequency of occurrence of keywords in each of the sub-time periods; determining at least one conversation interval based on each of the occurrence frequencies, a first preset frequency, and a second preset frequency; wherein the first preset frequency is greater than the second preset frequency; and determining a clustering result for each of the conversation intervals.
[0131] When the processor determines at least one conversation interval based on each of the occurrence frequencies, the first preset frequency, and the second preset frequency stored in the execution memory, it also executes the following computer program: determining, in chronological order, a first occurrence frequency greater than or equal to the first preset frequency and a second occurrence frequency less than or equal to the second preset frequency; and determining a conversation interval based on conversation texts within multiple sub-time periods determined based on adjacent first occurrences of the first occurrence frequency and last occurrence of the second occurrence frequency.
[0132] When executing the conversation topic determination method stored in the memory, the processor also executes the following computer program: parsing N end-conversation texts in the first of two adjacent conversation intervals and M start-conversation texts in the second of two adjacent conversation intervals; and merging the two adjacent conversation intervals into one conversation interval when it is determined that the continuity between the N end-conversation texts in the first of the conversation intervals and the M start-conversation texts in the second of the conversation intervals meets a preset condition.
[0133] When executing the conversation topic determination method stored in the memory, the processor also executes the following computer program: calculating the similarity between the N end conversation texts in the first conversation interval and the M start conversation texts in the second conversation interval of two adjacent conversation intervals; when the similarity is greater than a preset threshold, merging the two adjacent conversation intervals into one conversation interval.
[0134] When determining a clustering result for each conversation interval stored in the execution memory, the processor further executes the following computer program: for each conversation interval, using a clustering algorithm to calculate the conversation text in the conversation interval to obtain a clustering result corresponding to the conversation interval.
[0135] When determining a conversation topic based on each clustering result stored in the memory, the processor further executes the following computer program: determining a conversation topic of a conversation interval corresponding to the clustering result based on a plurality of text words included in the clustering result.
[0136] The embodiment of the present application uses a preset algorithm to calculate multiple conversation texts to obtain at least one clustering result, and then modifies each clustering result based on keywords, semantic analysis and similarity calculation. Then, a conversation topic is determined based on each clustering result. It can accurately, comprehensively and effectively determine at least one conversation topic corresponding to multiple conversation texts, thereby improving work efficiency.
[0137] Optionally, in this embodiment, the above-mentioned storage medium may include, but is not limited to, various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk. Optionally, in this embodiment, the processor executes the method steps recorded in the above-mentioned embodiment according to the program code stored in the storage medium. Optionally, the specific examples in this embodiment can refer to the examples described in the above-mentioned embodiment and the optional implementation mode, and this embodiment will not be repeated here. Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present application can be implemented using a general-purpose computing device, which can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented using program codes executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be performed in a different order than here, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. In this way, the present application is not limited to any specific combination of hardware and software.
[0138] In addition, although exemplary embodiments have been described herein, the scope includes any and all embodiments based on the present application with equivalent elements, modifications, omissions, combinations (e.g., solutions that intersect various embodiments), adaptations, or changes. The elements in the claims are to be interpreted broadly based on the language employed in the claims and are not limited to the examples described in this specification or during the prosecution of this application, which examples are to be interpreted as non-exclusive. Therefore, this specification and examples are intended to be considered as examples only, with the true scope and spirit being indicated by the following claims and the full scope of their equivalents.
[0139] The above description is intended to be illustrative rather than restrictive. For example, the above examples (or one or more of their solutions) can be used in combination with each other. For example, a person of ordinary skill in the art may use other embodiments when reading the above description. In addition, in the above-mentioned specific embodiments, various features can be grouped together to simplify the application. This should not be interpreted as an intention that a disclosed feature that is not required to be protected is necessary for any claim. On the contrary, the subject matter of the present application may be less than all the features of a specific disclosed embodiment. Thus, the following claims are incorporated into the specific embodiments as examples or embodiments, wherein each claim is independently a separate embodiment, and it is considered that these embodiments can be combined with each other in various combinations or arrangements. The scope of this application should be determined with reference to the appended claims and the full scope of equivalents to which these claims are entitled.
[0140] The above describes in detail several embodiments of the present application, but the present application is not limited to these specific embodiments. Those skilled in the art can make various variations and modifications to the embodiments based on the concept of the present application, and these variations and modifications should all fall within the scope of protection claimed by the present application.
Claims
1. A method for determining a conversation topic, characterized in that: include: Get multiple conversation texts; Calculating the plurality of conversation texts using a preset algorithm to obtain at least one clustering result includes: dividing the time period corresponding to the plurality of conversation texts into a plurality of sub-time periods in chronological order, wherein the duration of each sub-time period is dynamically adjusted according to the number of all conversation texts: when there are more conversation texts, the sub-time period is shorter; when there are fewer conversation texts, the sub-time period is longer; determining the frequency of occurrence of a keyword in each sub-time period, wherein the keyword is dynamically extracted by searching for words with higher frequency of occurrence in all conversation texts; determining at least one conversation interval based on each of the occurrence frequencies, a first preset frequency, and a second preset frequency, wherein the first preset frequency is greater than the second preset frequency; determining a clustering result for each conversation interval, wherein each clustering result includes at least one keyword, and each keyword corresponds to a conversation topic; determining at least one conversation interval based on each of the occurrence frequencies, the first preset frequency, and the second preset frequency includes: sequentially determining a first occurrence frequency greater than or equal to the first preset frequency and a second occurrence frequency less than or equal to the second preset frequency in chronological order; determining a conversation interval based on the conversation texts in the plurality of sub-time periods determined by the first occurrence of the first occurrence frequency and the last occurrence of the second occurrence frequency; Determining a conversation topic based on each clustering result includes: searching for a conversation topic corresponding to a keyword in each clustering result according to a preset mapping relationship between the keyword and the conversation topic, and using the conversation topic as the conversation topic corresponding to each clustering result.
2. The method for determining a conversation topic according to claim 1, wherein: Also includes: Parsing N end-of-session texts in a first session interval and M start-of-session texts in a second session interval of two adjacent session intervals; When it is determined that the continuity between the N end conversation texts in the previous conversation interval and the M start conversation texts in the next conversation interval meets a preset condition, the two adjacent conversation intervals are merged into one conversation interval.
3. The method for determining a conversation topic according to claim 1 or 2, wherein: Also includes: Calculating the similarity between the N end-of-conversation texts in the first of two adjacent conversation intervals and the M start-of-conversation texts in the second of two adjacent conversation intervals; When the similarity is greater than a preset threshold, the two adjacent conversation intervals are merged into one conversation interval.
4. The method for determining a conversation topic according to claim 1, wherein: Determining a clustering result for each session interval includes: For each of the conversation intervals, a clustering algorithm is used to calculate the conversation texts in the conversation interval to obtain a clustering result corresponding to the conversation interval.
5. The method for determining a conversation topic according to claim 1, wherein: Determining a conversation topic based on each clustering result includes: A conversation topic of a conversation section corresponding to the clustering result is determined based on a plurality of text words included in the clustering result.
6. A device for determining a conversation topic, characterized in that: include: An acquisition module configured to acquire a plurality of conversation texts; The calculation module is configured to calculate the plurality of conversation texts using a preset algorithm to obtain at least one clustering result, including: dividing the time period corresponding to the plurality of conversation texts into a plurality of sub-time periods in chronological order, wherein the duration of each sub-time period is dynamically adjusted according to the number of all conversation texts: when there are more conversation texts, the sub-time period is shorter; when there are fewer conversation texts, the sub-time period is longer; determining the frequency of occurrence of a keyword in each sub-time period, wherein the keyword is dynamically extracted by searching for words with higher frequency of occurrence in all conversation texts; determining at least one conversation interval based on each of the occurrence frequencies, a first preset frequency, and a second preset frequency, wherein the first preset frequency is greater than the second preset frequency; determining a clustering result for each conversation interval, wherein each clustering result includes at least one keyword, and each keyword corresponds to a conversation topic; determining at least one conversation interval based on each of the occurrence frequencies, the first preset frequency, and the second preset frequency, including: sequentially determining a first occurrence frequency greater than or equal to the first preset frequency and a second occurrence frequency less than or equal to the second preset frequency in chronological order; determining a conversation interval based on the conversation texts within the plurality of sub-time periods determined by the first occurrence of the first occurrence frequency and the last occurrence of the second occurrence frequency; A determination module is configured to determine a conversation topic based on each clustering result, including: searching for a conversation topic corresponding to a keyword in each clustering result according to a pre-set mapping relationship between the keyword and the conversation topic, and using the conversation topic as the conversation topic corresponding to each clustering result.
7. A storage medium, characterized in that: The storage medium stores a computer program, which, when executed by a processor, performs the following steps: Get multiple conversation texts; Calculating the plurality of conversation texts using a preset algorithm to obtain at least one clustering result includes: dividing the time period corresponding to the plurality of conversation texts into a plurality of sub-time periods in chronological order, wherein the duration of each sub-time period is dynamically adjusted according to the number of all conversation texts: when there are more conversation texts, the sub-time period is shorter; when there are fewer conversation texts, the sub-time period is longer; determining the frequency of occurrence of a keyword in each sub-time period, wherein the keyword is dynamically extracted by searching for words with higher frequency of occurrence in all conversation texts; determining at least one conversation interval based on each of the occurrence frequencies, a first preset frequency, and a second preset frequency, wherein the first preset frequency is greater than the second preset frequency; determining a clustering result for each conversation interval, wherein each clustering result includes at least one keyword, and each keyword corresponds to a conversation topic; determining at least one conversation interval based on each of the occurrence frequencies, the first preset frequency, and the second preset frequency includes: sequentially determining a first occurrence frequency greater than or equal to the first preset frequency and a second occurrence frequency less than or equal to the second preset frequency in chronological order; determining a conversation interval based on the conversation texts in the plurality of sub-time periods determined by the first occurrence of the first occurrence frequency and the last occurrence of the second occurrence frequency; Determining a conversation topic based on each clustering result includes: searching for a conversation topic corresponding to a keyword in each clustering result according to a preset mapping relationship between the keyword and the conversation topic, and using the conversation topic as the conversation topic corresponding to each clustering result.
8. An electronic device, characterized in that: include: A processor and a memory, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the memory communicate via a bus. When the machine-readable instructions are executed by the processor, the following steps are performed: Get multiple conversation texts; Calculating the plurality of conversation texts using a preset algorithm to obtain at least one clustering result includes: dividing the time period corresponding to the plurality of conversation texts into a plurality of sub-time periods in chronological order, wherein the duration of each sub-time period is dynamically adjusted according to the number of all conversation texts: when there are more conversation texts, the sub-time period is shorter; when there are fewer conversation texts, the sub-time period is longer; determining the frequency of occurrence of a keyword in each sub-time period, wherein the keyword is dynamically extracted by searching for words with higher frequency of occurrence in all conversation texts; determining at least one conversation interval based on each of the occurrence frequencies, a first preset frequency, and a second preset frequency, wherein the first preset frequency is greater than the second preset frequency; determining a clustering result for each conversation interval, wherein each clustering result includes at least one keyword, and each keyword corresponds to a conversation topic; determining at least one conversation interval based on each of the occurrence frequencies, the first preset frequency, and the second preset frequency includes: sequentially determining a first occurrence frequency greater than or equal to the first preset frequency and a second occurrence frequency less than or equal to the second preset frequency in chronological order; determining a conversation interval based on the conversation texts in the plurality of sub-time periods determined by the first occurrence of the first occurrence frequency and the last occurrence of the second occurrence frequency; Determining a conversation topic based on each clustering result includes: searching for a conversation topic corresponding to a keyword in each clustering result according to a preset mapping relationship between the keyword and the conversation topic, and using the conversation topic as the conversation topic corresponding to each clustering result.
Citation Information
Patent Citations
Automatic extraction method of conversation text topic
CN101599071A
Device and method for classifying and sectioning information, device and method for retrieving and extracting information, recording medium, and information retrieval system
JP2002169592A