Event Information Processing Method, System, Device, and Storage Medium

By obtaining the target text information in the current period in the e-commerce platform, a new word discovery method that combines solidification and the longest word chain is adopted, hot trend word mining is combined with popularity and trend information, and a diversified part-of-speech template is used to extract event information, solving the problem that existing technology is difficult to capture sudden hot spot information, and achieving efficient and accurate hot spot information capture.

CN114943208BActive Publication Date: 2025-05-27ALIBABA (CHINA) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210442488.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-25
Publication Date
2025-05-27
Estimated Expiration
2042-04-25

AI Technical Summary

Technical Problem

It is difficult for the prior art to capture effective hot information in a timely and accurate manner in scenarios where hot information is relatively sudden and unpredictable.

Method used

By obtaining the target text information generated by the target application scenario in the current period, new word discovery is performed by combining solidification and the longest word chain, hot trend word mining is performed by combining the popularity information and popularity trend information of new words, and finally using a diversified part-of-speech template for event information extraction.

Benefits of technology

It realizes the timely and accurate capture of effective hot-spot event information in the unexpected and unpredictable hot-spot information scenarios, and improves the efficiency and accuracy of hot-spot information capture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114943208B_ABST
    Figure CN114943208B_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide an event information processing method, system, device, and storage medium. In the embodiments of the present application, for the target text information generated in the current period in the target application scenario, new words are first discovered by using the coagulation degree and the longest word chain to obtain the current new word set, and words with ambiguous meanings can be excluded. Then, combined with the heat information and heat trend information of each new word in the new word set, hot trend word mining is performed to obtain hot trend words with a certain heat and heat trend from the current new word set, thereby ensuring the utilization rate of hot trend words. Further, using a diverse part-of-speech template, event information extraction is performed on multiple text information where the hot trend words are located to obtain representative information as the event description information corresponding to the corresponding hot trend words, thus achieving the purpose of timely and accurately capturing effective hot event information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet technologies, and in particular, to an event information processing method, system, device, and storage medium. Background Art

[0002] In the e-commerce field, as a medium between merchants and consumer users, an e-commerce platform, on the one hand, needs to perceive the needs of consumer users, select merchants and products according to the needs of consumer users, and on the other hand, needs to conduct operational communication with merchants, and finally publish the products of merchants on the platform for consumer users to purchase. In the process of operational communication with merchants, if relevant public opinion information can be discovered in a timely and accurate manner, it is beneficial for the platform to make more reasonable decisions.

[0003] In the prior art, there are various public opinion discovery schemes, but these schemes all need to define public opinion intentions. For example, a user specifies keywords, selects text information related to the keywords from a text library, and further uses methods such as sentiment analysis, word cloud analysis, or channel source analysis to mine hot information that meets the pre-defined intentions from the selected text information. These public opinion discovery schemes are more suitable for scenarios where information changes smoothly and has a long cycle, and are not applicable to scenarios where hot information is relatively sudden and unpredictable, and cannot capture effective hot information in a timely and accurate manner. Summary of the Invention

[0004] Multiple aspects of this application provide an event information processing method, system, device, and storage medium to solve the problem that in scenarios where hot information is relatively sudden and unpredictable, effective hot information cannot be captured in a timely and accurate manner.

[0005] An embodiment of this application provides an event information processing method, including: obtaining target text information generated in a current time period for a target application scenario, where the target text information includes original text information and / or text information converted from voice information; performing new word discovery on the target text information by using a method that combines coagulation degree and the longest word chain to obtain a current new word set; mining hot trend words based on the heat information and heat trend information of each new word in the current new word set to obtain hot trend words in the current new word set; and performing event information extraction on the text information where the hot trend words are located based on a diversified part-of-speech template to obtain event description information corresponding to the hot trend words.

[0006] The embodiments of the present application further provide an event information processing system, including: a text acquisition module, configured to acquire target text information generated in a current period of a target application scenario, where the target text information includes original text information and / or text information converted from voice information; a new word discovery module, configured to discover new words from the target text information by integrating the cohesion degree and the longest word chain, so as to obtain a current new word set; a hot trend mining module, configured to perform hot trend word mining based on the heat information and heat trend information of each new word in the current new word set, so as to obtain hot trend words in the current new word set; and an event extraction module, configured to perform event information extraction on the text information where the hot trend words are located based on a diversified part-of-speech template, so as to obtain event description information corresponding to the hot trend words.

[0007] The embodiments of the present application further provide an electronic device, including: a memory and a processor, where the memory is configured to store a computer program, and the processor is coupled to the memory and configured to execute the computer program to implement the steps in the above method.

[0008] The embodiments of the present application further provide a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to implement the steps in the above method.

[0009] The technical solutions provided in the embodiments of the present application are directed to the target text information generated in a current period of a target application scenario. First, new words are discovered by integrating the cohesion degree and the longest word chain to obtain a current new word set, which can ensure the tightness between individual characters in each new word and exclude words with ambiguous meanings, thereby improving the accuracy of new words. Then, hot trend word mining is performed by combining the heat information and heat trend information of each new word in the new word set, taking into account both the heat and heat trend of the new words, to obtain hot trend words with a certain heat and heat trend from the current new word set, thereby ensuring the utilization rate of the hot trend words. Further, using a diversified part-of-speech template, event information extraction is performed on multiple text information where the hot trend words are located to obtain representative information as the event description information corresponding to the corresponding hot trend words, thereby achieving the purpose of timely and accurately capturing effective hot event information. Description of the Drawings

[0010] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0011] Figure 1a It is a schematic flowchart of an event information processing method provided by an exemplary embodiment of the present application;

[0012] Figure 1bA trend chart of the change in the popularity information of new words provided by an exemplary embodiment of the present application;

[0013] Figure 2a A schematic structural diagram of an event information processing system provided by an exemplary embodiment of the present application;

[0014] Figure 2b A schematic internal structure diagram of each module in an information processing system provided by an exemplary embodiment of the present application;

[0015] Figure 3 A schematic structural diagram of an electronic device provided by an exemplary embodiment of the present application. Detailed implementation manners

[0016] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0017] In the prior art, there are various public opinion discovery platforms for discovering relevant public opinion information. These public opinion discovery platforms need to first define the public opinion intention. For example, the user specifies keywords, selects text information related to the keywords from the text library, and further uses a certain analysis method to mine the hot information that meets the pre-defined intention from the selected text information. These public opinion discovery solutions rely on the existing text library and the definition of public opinion intention, and are more suitable for scenarios where the information changes smoothly and has a long cycle. They are not applicable to scenarios where hot information is relatively sudden and unpredictable, and cannot capture effective hot information in a timely and accurate manner.

[0018] To solve the above technical problems, the embodiments of the present application provide the following technical solutions. For a target application scenario, based on the target text information generated within the current period, a method of fusing solidification degree and the longest word chain is used to discover new words, obtaining a new word set with relatively high accuracy. Furthermore, by combining the popularity information and popularity trend information of each new word in the new word set, hot trend words are mined. Finally, using diverse part-of-speech templates, event information is extracted from the text information where the hot trend words are located, obtaining representative hot event description information. It no longer relies on the text library and does not need to pre-perceive the public opinion intention in advance. Instead, it can mine new words from the text information in the current period, capture effective hot event information in a timely and accurate manner based on the new words, and real-time perceive the problems, dynamics, etc. of the target application scenario, which are hot information with a certain trend, facilitating timely solution or giving coping strategies to meet the application requirements in this scenario.

[0019] The following will, in conjunction with the accompanying drawings, elaborate in detail on the technical solutions provided by various embodiments of the present application.

[0020] Figure 1a It is a schematic flowchart of an event information processing method provided for an exemplary embodiment of the present application. As Figure 1a shown, the method includes:

[0021] 101. Obtain target text information generated by a target application scenario during the current period, where the target text information includes original text information and / or text information converted from voice information;

[0022] 102. Discover new words from the target text information by fusing solidification degree and the longest word chain to obtain the current new word set;

[0023] 103. Mine hot trend words based on the popularity information and popularity trend information of each new word in the current new word set to obtain the hot trend words in the current new word set;

[0024] 104. Extract event information from the text information where the hot trend words are located based on diverse part-of-speech templates to obtain event description information corresponding to the hot trend words.

[0025] In this embodiment, an application scenario with the need for public opinion information analysis is referred to as a target application scenario. In this embodiment, the continuous process of the target application scenario is divided into different periods, and the duration of each period can be the same or different. The target application scenario will generate corresponding text information in each period. For the convenience of description and distinction, the different periods corresponding to the target application scenario are divided into the current period and the historical period. Of course, as time goes by, the current period will become the historical period.

[0026] In this embodiment, taking the text information generated by the target application scenario during the current period as the information basis, real-time public opinion information discovery is carried out for the target application scenario. In this embodiment, the current period can be a relatively short current period or a relatively long current period. For example, it can be the current week, the current few days, or the current few hours, etc., or it can be the current month, the current few months, or the current year, etc. This embodiment does not limit the length of the current period, and it can be specifically determined according to the application scenario and the requirement for the timeliness of hot event information discovery in the specific application scenario.

[0027] In this embodiment, public opinion information specifically refers to the description information of hot events with a certain trend that occur in the target application scenario in the current period, and the hot events are instructive or representative. For the convenience of description and distinction, the text information generated by the target application scenario in the current period is referred to as target text information. It should be noted that different target application scenarios have different target text information generated in the current period. This embodiment does not specifically limit the target application scenario. For example, the target application scenario can be various technical fields with public opinion information analysis needs, such as the e-commerce field, the financial field, the gaming field, the catering field, the automotive field, and so on.

[0028] Furthermore, for the e-commerce field, it can also be an e-commerce production site operation scenario. The e-commerce production site operation scenario refers to the operation stage for merchants in which sales personnel on the e-commerce production site use various communication methods to communicate with various merchants, explore merchants and products that cooperate with the e-commerce platform, and face the merchant. In the e-commerce production site operation scenario, the method provided in the embodiment of the present application can be used to discover hot events with certain trends in the communication between sales personnel and merchants in real time, and capture dynamic information of merchants or markets in real time, so that the e-commerce platform can quickly discover and solve problems, and achieve the purpose of timely adjusting the platform strategy.

[0029] In this embodiment, first, obtain the target text information generated by the target application scenario during the current period. The ways of generating target text information in different target application scenarios may vary. Regardless of the application scenario, the sources of the target text information include the following two categories. One is the original text information generated by the target application scenario during the current period, and the other is the text information converted from the voice information generated by the target application scenario during the current period. For example, taking the e-commerce origin operation scenario as an example, marketers on the e-commerce origin side can communicate with merchants face-to-face in language and record the communication process in text and upload it to the e-commerce platform. The communication process recorded in text can be called operation process information, and this recording method can be called the note-taking method, but it is not limited to this. In addition, marketers on the e-commerce origin side can use telephone or real-time communication software to remotely conduct online audio and video communication with merchants, and voice information can be generated during this process. Whether it is the operation process information recorded in the above note-taking method or the voice information generated by the online voice communication between e-commerce origin marketers and merchants, there may be information that can reflect the real-time needs of merchants, the changing trends of the market, and the real-time changes and impacts of relevant environmental information, etc. To a certain extent, these information can reflect the trends that the application scenario may have during the current period and for some time to come. Based on this, the operation process information recorded by e-commerce origin marketers and merchants can be obtained, and the ASR technology can be used to convert the voice information generated by the online audio and video communication between e-commerce origin marketers and merchants into text information, and at least one of these two types of text information can be obtained as the target text information generated by the e-commerce origin operation scenario during the current period.

[0030] Further, after obtaining the target text information generated by the target application scenario during the current period, new word discovery processing can be performed on the target text to obtain accurate and meaningful new words. In this embodiment, considering that the target text information is real-time data during the current period and may have the long-tail characteristic, the long-tail characteristic has the characteristic of skewed distribution, which means that a large amount of effective information may mainly exist in a small amount of text information (referred to as tail class information), while most of the text information (referred to as head class information) contains a small amount of effective information. That is to say, the hot event information to be mined may be sudden and unpredictable. Therefore, conventional text classification or clustering methods cannot be used to achieve the classification and analysis of hot information, and the recognition accuracy of the conventional method in tail class information will be significantly reduced, which will further reduce the accuracy of capturing hot event information. In the embodiment of this application, fully considering that the target text information is limited text information during the current period and may have the long-tail characteristic, the method of fusing solidification degree and the longest word chain is used to perform new word discovery on the target text information to obtain the current new word set, which can ensure the tightness between individual characters in each new word, exclude words with ambiguous meanings, so as to improve the accuracy of new words and reduce the influence of the long-tail characteristic.

[0031] Among them, a specific implementation manner of discovering new words from the target text information by fusing the solidification degree and the longest word chain is as follows: The first step is to use the n-gram language model to identify words in the target text information and perform solidification degree filtering on the identified words to obtain an initial word set. Here, n is an integer greater than or equal to 2, which can not only ensure the tightness of the semantic connection between individual characters forming each new word in the initial word set, but also filter out words with inaccurate or ambiguous meanings, improving the accuracy of the initial words. The second step is to use the longest word chain algorithm to segment the target text information based on the initial word set to obtain candidate segmentations. Using the longest word chain algorithm can ensure obtaining as rich segmentations as possible from the target text information. The third step is to perform backtracking filtering on the candidate segmentations based on the initial word set, filtering out invalid segmentations (such as wrongly segmented ones) to obtain valid segmentations. The valid segmentations form the current new word set, improving the accuracy of new words, reducing the number of new words, and improving the subsequent calculation efficiency based on segmentation.

[0032] In this embodiment, the n-gram language model is used to identify words in the target text information and perform solidification degree filtering on the identified words to obtain an initial word set. The specific implementation manner is as follows: Select a fixed value n (n≥2), count the words composed of 2-grams, 3-grams,..., n-grams, calculate the internal solidification degree of each word respectively, and only retain the words with a solidification degree higher than the set solidification degree threshold to form the initial word set G. Among them, different solidification degree thresholds can be set for 2-grams, 3-grams,..., n-grams, or the same solidification degree threshold can be set. However, generally, the larger the value of n, the less sufficient the statistics of the formed words are, that is, the larger the value of n, the solidification degree threshold needs to be increased adaptively. Taking a text fragment in the target text information as "new customer" as an example, and selecting n equal to 3 to identify words in the text information, and setting the thresholds corresponding to 2-grams and 3-grams as A1 and A2 respectively, then it is necessary to count the words composed of n = 2 and n = 3, calculate the internal solidification degree of each word composed of 2-grams and 3-grams, retain the words with a solidification degree greater than A1 among the words composed of 2-grams, and retain the words with a solidification degree greater than A2 among the words composed of 3-grams to form the initial word set.

[0033] In this embodiment, the solidification degree can be calculated by the following formula: Among them, a, b, and c represent the smallest units that make up a word. ab and bc represent the word segments obtained by segmentation when n equals 2, and abc represents the word segment obtained by segmentation when n equals 3. P(a), P(c), P(ab), P(bc), and P(abc) respectively represent the probabilities of a, c, ab, bc, and abc appearing in the initial word set. Based on this, the initial word set formed by the above recognition method can be {new customer, customer, new customer}. The above values of n and the setting of the coagulation degree threshold are only for illustrative purposes, but are not limited thereto.

[0034] After obtaining the initial word set, the longest word chain algorithm is adopted to segment the target text information based on the initial word set to obtain candidate word segments. The specific implementation method is as follows: A dictionary tree is constructed based on the initial word segment set, and the depth-first search algorithm is used to search the dictionary tree for the target text information to obtain the longest word chain, and the longest word chain is used as the candidate word segment. It should be noted that the algorithm for searching the dictionary tree is not limited to the depth-first search algorithm.

[0035] Among them, when constructing a dictionary tree based on the initial word segment set, the dictionary tree can be constructed based on the priorities of each word. The word with the lowest priority is used as the top of the dictionary tree, and the word with the highest priority is used as the end of the dictionary tree. The remaining words are constructed from the top to the end of the dictionary tree in ascending order of their priorities; or, the word with the lowest priority can also be used as the end of the dictionary tree, and the word with the highest priority is used as the top of the dictionary tree. The remaining words are constructed from the top to the end of the dictionary tree in ascending order of their priorities. Among them, there are various types of dictionary trees. For example, the Ngram-Trie tree can be used to quickly search whether the words in the initial word segment set appear in a string. Still taking the initial word set that can be {new customer, customer, new customer} as an example, the priority of "new customer" is higher than the priority of "customer". The construction of the dictionary tree based on the initial word segment set is only for illustrative purposes, but is not limited thereto.

[0036] After the trie tree constructed based on the initial word segmentation set is completed, for the target text information, the depth-first search algorithm is used to search the trie tree to obtain the longest word chain. The specific implementation method is as follows: Based on the trie tree, the words in the same branch of the trie tree are combined pairwise in ascending order of priority to obtain multiple combined new words until the combined new words of the longest word chain are obtained, and the word frequencies of the multiple combined new words are calculated. The combined new words that meet the word frequency threshold and the combined new words with the longest word chain are used as candidate word segmentations. More specifically, the above-mentioned longest word chain algorithm is used to roughly segment the target text information, and the frequency of each word segmentation is counted. The segmentation rule is that if only one word segmentation appears in the initial word set obtained in the first step, this word segmentation is not segmented. For example, for "new customer", as long as "new guest" and "customer" are in the initial word set, even if "new customer" is not in the initial word set, then "new customer" is not segmented and used as a candidate word segmentation; and the word frequencies of "new guest" and "customer" are calculated, and the word segmentations that meet the word frequency threshold are also used as candidate word segmentations.

[0037] In this embodiment, after obtaining the candidate word segmentations, in order to ensure the accuracy of the candidate word segmentations, the candidate word segmentations can be backtracked and filtered to filter out the incorrect candidate word segmentations obtained in the longest word chain algorithm to obtain more accurate word segmentations. In this embodiment, the candidate word segmentations can be backtracked and filtered based on the initial word set to obtain valid word segmentations. Among them, backtracking and filtering is a word filtering method used to filter out invalid word segmentations. The specific implementation method is as follows: For each candidate word segmentation, if the length of the candidate word segmentation is less than or equal to n and the candidate word segmentation appears in the initial word set, then the candidate word segmentation is used as a valid word segmentation; if the length of the candidate word segmentation is greater than n, then multiple word segments of length n are sequentially segmented from the candidate word segmentation, and the multiple word segments are respectively used as valid word segmentations. Thus, whether the length of the candidate word segmentation is greater than n or less than or equal to n, it can be determined whether it is a valid word segmentation. These valid word segmentations are the new word set mined from the text information generated in the current period of this application embodiment, simply referred to as the current new word set. Among them, n is a set value. For example, n = 3 or 4, and this set value can be determined according to specific requirements, and this embodiment does not limit this.

[0038] Furthermore, after obtaining the current new word set, hotspot trend words can be mined based on the popularity information and popularity trend information of each new word in the current new word set to obtain the hotspot trend words in the current new word set. A hotspot trend word refers to a new hotspot word that can represent the development trend of things to a certain extent. "Can represent the development trend of things to a certain extent" means having a certain trend. Correspondingly, a new hotspot word with a certain trend is called a hotspot trend word. One specific implementation method for mining hotspot trend words is as follows: Combine the historical new word sets in historical time periods to determine the popularity information and popularity trend information of each new word in the current new word set; Calculate the comprehensive hotspot trend score corresponding to each new word based on the popularity information and popularity trend information of each new word in the current new word set; Select the hotspot trend words from each new word according to the comprehensive hotspot trend score corresponding to each new word. Among them, the popularity information is used to reflect the popularity of the corresponding new word in the current time period. Optionally, it can be represented by the popularity ratio of each new word in the current time period or historical time periods, but not limited to this; The popularity trend information is used to reflect the change in the popularity of the corresponding new word from the historical time period to the current time period, such as whether the popularity has increased or decreased. Optionally, it can be represented by the change ratio of the popularity of each new word in the current time period and historical time periods, but not limited to this.

[0039] In the embodiments of the present application, the historical time period used is not limited. For example, it can be multiple historical time periods or one historical time period. Further optionally, it can be one or more historical time periods closest to the current time period, or one or more historical time periods not adjacent to the current time period selected according to a set strategy. One optional implementation method for combining the historical new word sets in historical time periods to determine the popularity information and popularity trend information of each new word in the current new word set is as follows: For each new word in the current new word set, determine the first popularity information of the new word in the current time period according to the word frequency of the new word in the current time period and the total word frequency of the current new word set in the current time period; And determine the second popularity information of the new word in the historical time period according to the word frequency of the new word in the historical time period and the total word frequency of the historical new word set in the historical time period; Then generate the popularity information of the new word according to the larger value of the first popularity information and the second popularity information, and the smaller value of the minimum popularity information in the current time period and the minimum popularity information in the historical time period; Generate the popularity trend information of the new word according to the difference in popularity between the first popularity information and the second popularity information, and the larger value of the first popularity information and the second popularity information.

[0040] In an optional embodiment, taking one historical time period as an example, each new word in the current new word set can be represented by a, and the current time period is represented by T 1 denote, and the word frequency of the new word a in the current time period is represented by freq(a, T 1)It is represented that the total word frequency of the current new word set in the current time period is count(T 1 ) and refers to the sum of the word frequencies of each new word in the current new word set in the current time period; the first heat information is represented by rate(a, T 1 ), and the minimum heat information in the previous time period is min rate (T 1 ) and refers to the minimum heat information among the heat information of each new word in the current new word set. The historical time period is represented by T 2 , the word frequency of new word a in the historical time period is represented by freq(a, T 2 ), and the total word frequency of the historical new word set in the historical time period is count(T 2 ) and refers to the sum of the word frequencies of each new word in the historical new word set in the historical time period. The second heat information is represented by rate(a, T 2 ), and the minimum heat information in the historical time period is min rate (T 2 ). The larger value between the first heat information and the second heat information is represented by max(rate(a, T 1 ), rate(a, T 2 )), and the smaller value between the minimum heat information in the current time period and the minimum heat information in the historical time period is represented by min(min rate (T 1 ), min rate (T 2 ))). The difference between the first heat information and the second heat information is represented by |rate(a, T 1 ) - rate(a, T 2 )|. Based on the above definitions, the following is an exemplary description of the determination or generation methods of the first heat information, the second heat information, the heat information and heat trend information of new words, and the comprehensive score of the hot trend.

[0041] In an alternative embodiment, the first heat information can be determined by the ratio of freq(a, T 1 ) to count(T 1 ), for example, it can be represented by the following formula: However, it is not limited thereto. For example, the first heat information can also be determined by weighted summation, weighted product, or other numerical calculation methods of freq(a, T 1 ) and count(T 1 ). Correspondingly, the second heat information can be determined by the ratio of freq(a, T 2 ) to count(T 2 ), for example, it can be represented by the following formula: However, it is not limited thereto. For example, the second heat information can also be determined by weighted summation of freq(a, T 2 ) and count(T 2 ), weighted product and other numerical calculation methods. In this embodiment, the determination methods of the first heat information and the second heat information are not limited. It should be noted that the determination methods of the first heat information and the second heat information are preferably the same.

[0042] Continuing with the above optional embodiment, after the first heat information and the second heat information are determined, when the ratio of the larger of the first heat information and the second heat information to the smaller of the minimum heat information in the current time period and the minimum heat information in the historical time period is determined as the heat information of the new word, the heat information of the new word can be generated by the following formula: In this formula, the logarithmic function log() is taken as an example to determine the heat information of the new word, but it is not limited thereto.

[0043] Similarly, continuing with the above optional embodiment, when the heat difference between the first heat information and the second heat information is determined as the heat trend information of the new word by the ratio to the larger of the first heat information and the second heat information, for example, it can be expressed by the following formula: When rate(a, T 1 ) ≥ rate(a, T 2 ), When rate(a, T 1 ) < rate(a, T 2 ), In this formula, the square root function sqrt() is taken as an example to determine the heat information of the new word, but it is not limited thereto, and other functions such as cube root can also be used.

[0044] That is to say, the determination methods of the new word heat information and the heat trend information are related to the magnitudes of rate(a, T1) and rate(a, T2). It should be noted that the above determination methods of the heat information and the heat trend information are only exemplary descriptions, but not limited thereto.

[0045] Furthermore, continuing with the above optional embodiment, after the new word heat information and the heat trend information are determined, the product of the new word heat information and the heat trend information can be used as the comprehensive score of the hot trend. In this embodiment, the comprehensive score of the hot trend can be calculated by the following formula:

[0046] When rate(a, T 1 ) ≥ rate(a, T 2 ),

[0047]

[0048] When rate(a, T1 ) < rate(a, T 2 ) When

[0049]

[0050] It should be noted that in the above formula, score(a) represents that the method for determining the comprehensive score of the above hot trend is only an exemplary illustration, but not limited thereto. For example, the hot word heat information and heat trend information can also be weighted and summed, or weighted and multiplied, etc., to obtain the comprehensive score of the hot trend of the new word.

[0051] Furthermore, in the case of multiple historical periods, multiple historical periods can be aggregated and regarded as one historical period, and the first heat information, the second heat information, the heat information and heat trend information of the new word, and the comprehensive score of the hot trend can be calculated in the manner given in the above example, but not limited thereto.

[0052] After obtaining the comprehensive score of the hot trend corresponding to each new word, the hot trend words can be selected from each new word according to the comprehensive score of the hot trend corresponding to each new word. Optionally, at least one new word with the highest comprehensive score of the hot trend can be selected as the hot trend word, or at least one new word with a comprehensive score of the hot trend greater than the set score threshold can be selected as the hot trend word. It should be noted that the number of hot trend words can be one or multiple. Preferably, multiple hot trend words can be selected.

[0053] Furthermore, since the new words mined may have the characteristics of uneven semantic distribution and long-tail problems, for example, there may be new words with the same or similar semantics. In order to reduce the number of hot trend words, after obtaining the comprehensive score of the hot trend corresponding to each new word, in the process of selecting hot trend words, multiple candidate trend words are selected from each new word according to the comprehensive score of the hot trend corresponding to each new word; then, according to the text information where the multiple candidate trend words are located, semantic clustering is performed on the multiple candidate trend words to obtain the hot trend words. Among them, clustering is used to cluster the candidate trend words so that the candidate trend words with the same or similar semantics are clustered together, and one of them is selected as the representative of this type of candidate trend word, that is, the hot trend word, which is beneficial to reducing the number of hot trend words, thereby reducing the computational complexity of subsequent event information extraction based on hot trend words and improving the efficiency of obtaining event description information.

[0054] In this embodiment, when semantic clustering is performed on multiple candidate trend words based on the text information where the multiple candidate trend words are located to obtain hot trend words, the simhash algorithm can be used, but it is not limited thereto. Among them, when using the simhash algorithm, first, the text information where the multiple candidate trend words are located is mapped into a hash value; according to the hash values corresponding to the text information where the multiple candidate trend words are located, the similarity of the multiple candidate trend words is calculated, and the candidate trend words with the similarity meeting the requirements are aggregated into one category, and one is selected from the same category as the hot trend word. That is, first, the text information where the candidate trend word is located is obtained, and this text information is mapped into a 64-bit simhash value through the simhash algorithm. By calculating the similarity between the simhash values corresponding to the text information of any two candidate trend words, this similarity is used as the similarity between the candidate trend words; then, the candidate trend words are aggregated according to the similarity between the candidate trend words, and the candidate trend words with the similarity greater than the set similarity threshold are aggregated into one category to achieve the effect of removing duplicates. More specifically, when using the simhash algorithm, weights are assigned to the multiple candidate trend words, the text information where the multiple candidate trend words are located is mapped into a hash value, and these hash values are weighted and merged for dimensionality reduction processing according to the corresponding weights, mapping the high-dimensional feature vectors into low-dimensional feature vectors, mapping the text information into a 64-bit simhash value, and then determining whether the semantics of the corresponding two candidate trend words are repeated or highly similar through the Hamming distance between the two low-dimensional feature vectors. For example, in the e-commerce origin operation and maintenance scenario, for candidate trend words such as "backhaul", "form", "publish", "fill in", "publish form", "send", etc., they can be clustered into one category, and the word "backhaul" among them can be used as the representative of this category of words, that is, as a hot trend word; correspondingly, for candidate trend words such as "press conference", "press conference invitation", "invite press conference", "invite", "whether", "demand point", "whether there is an investment awareness", "invite home time", "pain point", etc., they can be clustered into one category, and the word "press conference" among them can be used as the representative of this category of words, that is, as a hot trend word. Of course, it is possible that only one candidate trend word is selected. For the case where there is only one candidate trend word, clustering processing may not be performed, but this candidate trend word can be directly used as the hot trend word. Of course, clustering operations can also be performed, but the clustering result is still to use this candidate trend word as the hot trend word. For example, in the e-commerce origin operation and maintenance scenario, single candidate trend words such as "participate in meeting", "meeting", "brand", "renewal fee", etc. may appear, and these single candidate trend words, such as "participate in meeting", "meeting", "brand", "renewal fee", etc. are the hot trend words.

[0055] Further, after obtaining the hot trend words, since the determined hot trend words are generally the smallest units of the new words corresponding to the longest word chain and usually cannot reflect the actual semantic information of the event, based on this, in the embodiments of the present application, diverse part-of-speech templates are pre-constructed. Based on the diverse part-of-speech templates, event information extraction is performed on the text information where the hot trend words are located to obtain the event description information corresponding to the hot trend words, that is, to obtain the actual semantic information of the hot event. In this embodiment, there are at least two or more part-of-speech templates, and each part-of-speech template is used to extract a phrase combination that can reflect event information. In this embodiment, the part-of-speech template can extract a phrase combination with the same part of speech, or can extract a phrase combination with two or more parts of speech, and no limitation is made in this regard. Each part-of-speech template will define the relevant and mutually exclusive relationships between parts of speech, such as which parts of speech can be combined or collocated with each other, and which parts of speech are prohibited from being combined or collocated, etc. Based on this, an optional implementation manner of performing event information extraction on the text information where the hot trend words are located based on the diverse part-of-speech templates is as follows: For each hot trend word, first, based on the relevant and mutually exclusive relationships between parts of speech defined by the diverse part-of-speech templates, multiple initial phrase combinations are extracted from the text information where the hot trend word is located. Among them, the initial phrase combination includes one or more words and has a part-of-speech attribute, and this part-of-speech attribute is used to reflect the parts of speech of the words included in the initial phrase combination and the part-of-speech collocation; further, according to the frequency of occurrence of the multiple initial phrase combinations in the text information where the hot trend word is located, a target phrase combination is selected from the multiple initial phrase combinations as the event description information corresponding to the hot trend word.

[0056] In this embodiment, from the perspective of part-of-speech collocation, the diverse part-of-speech templates include at least two or more of the part-of-speech templates of preposition-noun collocation, preposition-verb collocation, verb-verb collocation, noun-verb collocation, and noun-noun collocation. From the perspective of event description, the diverse part-of-speech templates include at least two or more of the SV part-of-speech template, SVC part-of-speech template, SVO part-of-speech template, SVOO part-of-speech template, and SVOA part-of-speech template. Among them, the SV part-of-speech template is Subject (subject) + Verb (predicate), that is, the subject-predicate structure, referring to a sentence pattern composed of one or several subjects plus one or several predicates, usually involving the collocation of nouns or prepositions and verbs. The SVC part-of-speech template, that is, the SVP template, is Subject (subject) + Link.V (linking verb) + predicate (predicative), that is, the subject-linking verb-predicative structure, referring to a sentence pattern composed of a subject, a linking verb, and a predicative, usually involving the collocation of nouns, verbs, and nouns or prepositions. The SVO part-of-speech template is Subject (subject) + Verb (predicate) + Object (object), that is, the subject-predicate-object structure, referring to a sentence pattern composed of a subject, a predicate, and an object, usually involving the collocation of nouns, verbs, and nouns. The SVOO part-of-speech template is Subject (subject) + Verb (predicate) + Indirect Object (indirect object) + Direct Object (direct object), that is, the subject-double object structure, referring to a sentence pattern composed of a subject, a predicate, a direct object, and an indirect object, usually involving the collocation of nouns, verbs, and nouns. The SVOA part-of-speech template is Subject (subject) + Verb (predicate) + Object (object) + Object Complement (object complement), that is, the subject-predicate-object-object complement structure, referring to a sentence pattern composed of a subject, a predicate, an object, and an object complement, usually involving the collocation of nouns, verbs, nouns, and prepositions.

[0057] For each hot trend word, after obtaining multiple initial phrase combinations corresponding to the hot trend word, the frequency of each initial phrase combination appearing in the text information where the hot trend word is located can be counted; then, the frequencies of the multiple initial phrase combinations appearing in the text information where the hot trend word is located are sorted, and at least one target phrase combination whose appearance frequency meets the frequency threshold is selected, or at least one target phrase combination ranked at the top is selected as the event description information corresponding to the hot trend word.

[0058] In the embodiment of the present application, for the target text information generated in the target application scenario during the current period, new words are first discovered by using the solidification degree and the longest word chain to obtain the current new word set, which can ensure the tightness between individual characters in each new word and exclude words with ambiguous meanings, thereby improving the accuracy of new words. Then, combined with the popularity information and popularity trend information of each new word in the new word set, hot trend word mining is carried out, taking into account the popularity and popularity trend of new words, and hot trend words with a certain popularity and popularity trend are obtained from the current new word set, thereby ensuring the utilization rate of hot trend words; further, using diverse part-of-speech templates, event information extraction is performed on multiple text information where the hot trend words are located to obtain representative information as the event description information corresponding to the corresponding hot trend words, thus achieving the purpose of timely and accurately capturing effective hot event information.

[0059] In this embodiment, taking the e-commerce origin operation scenario as the target application scenario as an example, in this scenario, the target text information includes the operation process information recorded by the e-commerce origin marketers about their operations with merchants and / or the text information converted from the voice information generated by the online audio and video communication between the e-commerce origin marketers and merchants. Then, the event description information in the e-commerce origin operation scenario can be mined according to the event information processing method provided in the above embodiment. In the following scenario embodiments, taking the mining of the description information of the production and power rationing event in the e-commerce origin operation scenario as an example for illustration, but the mined events are not limited to the production and power rationing event.

[0060] Specifically, in the e-commerce origin operation scenario, on the one hand, the note information uploaded by the e-commerce origin marketers during the current period is obtained. The operation process information recorded by the e-commerce origin marketers about their operations with merchants is recorded in these note information, and the voice information generated by the online audio and video communication between the e-commerce origin marketers and each merchant during the current period is obtained. The ASR technology is used to perform text conversion on these voice information to obtain text information, and the operation process information recorded by the e-commerce origin marketers about their operations with merchants is also recorded in these text information.

[0061] Next, a new word mining operation is performed, that is, new words are mined from the operation process information during the current period by using the fusion of the solidification degree and the longest word chain to obtain the current new word set.

[0062] Taking the case of production and power rationing as an example, the operation process information in the current period will include multiple new words related to power rationing. For example, during the communication process, merchants will mention that they are busy during specific periods (such as holidays), lack manpower, have insufficient product supply or postponed delivery, etc. In this case, the following new words can be mined from these operation process information, but not limited to: very busy, individual, no manpower, time is tight, relatively busy, customers decide to renew the contract, customers cannot cooperate, cannot meet the cooperation requirements, can participate in activities, reduced supply quantity, postponed supply time, agency operation, production capacity cannot keep up, factory power rationing, rising raw materials, power rationing, great impact, not considered, power rationing, holiday, vacation, brand, Tuesday, next week, logic, etc.

[0063] Next, perform the operation of mining hot trend words, that is, determine the heat information and heat trend information of each new word found above, and perform hot trend word mining based on the heat information and heat trend information of each new word to obtain the hot trend words in the current period.

[0064] For example, the heat information and heat trend change information of new words can be calculated by combining the new words in the current period and the new words in the historical period, and finally the hot trend words that appear in the current period can be obtained. Further, in this process, the chart information of the heat change of each new word can also be drawn, and with the help of the chart information, the heat change information of each new word can be better identified, as Figure 1b shown. In Figure 1b the horizontal axis represents the time dimension, which can be the current day, the previous day, or the same period of last week, the same period of last month, etc., and the specific time dimension can be defined by oneself; the vertical axis represents the heat change information of the new word in the corresponding time dimension, and this curve represents the fluctuation of the heat information. Continuing with the new words found in the above production and power rationing situation, the hot trend words in the current period can finally be mined, such as holiday, power rationing, great impact, logic, etc.

[0065] Next, based on a diverse set of part-of-speech templates, event information extraction is performed on the text information where the hot trend words mined in the current period are located, and event description information that can reflect the event trend in the current period can be obtained. Continuing with the above production and power rationing situation, in the current period, for example, at the end of September, it can be obtained that some merchants may have insufficient production capacity or insufficient manpower due to production and power rationing during the upcoming holiday, such as the National Day holiday, resulting in the inability to supply goods in a timely manner or insufficient supply. Based on this event information, the e-commerce platform can make corresponding strategy adjustments in a timely manner, such as replacing other merchants, or suspending cooperation with this merchant during the next holiday, or renegotiating a new cooperation model with this merchant, such as adjusting cooperation fees, cycles, etc., to ensure that the e-commerce platform can have enough merchants and products to participate in activities during the holiday, ensure the service quality provided by the platform to consumer users, and guarantee the user traffic and stickiness of the platform.

[0066] Figure 2a The structural schematic diagram of the event information processing system provided by an exemplary embodiment of the present application Figure 2b The internal structural schematic diagram of each module in the system given by an exemplary embodiment of the present application. As Figure 2a and Figure 2b shown, the system includes:

[0067] A text acquisition module 21, configured to acquire target text information generated in a current period of a target application scenario, where the target text information includes original text information and / or text information converted from voice information;

[0068] A new word discovery module 22, configured to discover new words in the target text information by integrating the cohesion degree and the longest word chain, so as to obtain a current new word set;

[0069] A hot trend mining module 23, configured to mine hot trend words in the current new word set based on the heat information and heat trend information of each new word in the current new word set, so as to obtain hot trend words in the current new word set;

[0070] An event extraction module 24, configured to extract event information from the text information where the hot trend words are located based on a diversified part-of-speech template, so as to obtain event description information corresponding to the hot trend words.

[0071] Further, when the new word discovery module 22 is used to discover new words in the target text information by integrating the cohesion degree and the longest word chain to obtain a current new word set, it is specifically configured to: use an n-gram language model to identify words in the target text information, and perform cohesion degree filtering on the identified words to obtain an initial word set, where n is an integer greater than or equal to 2; use the longest word chain algorithm to perform word segmentation on the target text information based on the initial word set to obtain candidate word segments; perform backtracking filtering on the candidate word segments based on the initial word set to obtain valid word segments, and the valid word segments form the current new word set.

[0072] Further, when the new word discovery module 22 is used to perform word segmentation on the target text information based on the initial word set by using the longest word chain algorithm to obtain candidate word segments, it is specifically configured to: construct a dictionary tree based on the initial word segmentation set, and use a depth-first search algorithm to search the dictionary tree for the target text information to obtain the longest word chain, and use the longest word chain as the candidate word segment.

[0073] Further, when the new word discovery module 22 is used to backtrack and filter candidate word segmentations based on the initial word set to obtain valid word segmentations, it is specifically used for: for each candidate word segmentation, if the length of the candidate word segmentation is less than or equal to n and the candidate word segmentation appears in the initial word set, then the candidate word segmentation is used as a valid word segmentation; if the length of the candidate word segmentation is greater than n, then the candidate word segmentation is sequentially segmented into multiple word segmentation segments with a length of n, and the multiple word segmentation segments are respectively used as valid word segmentations.

[0074] Further, when the hot trend mining module 23 is used to mine hot trend words in the current new word set based on the heat information and heat trend information of each new word in the current new word set to obtain the hot trend words in the current new word set, it is specifically used for: combining the historical new word set in the historical time period to determine the heat information and heat trend information of each new word in the current new word set; calculating the comprehensive heat trend score corresponding to each new word based on the heat information and heat trend information of each new word in the current new word set; selecting hot trend words from each new word according to the comprehensive heat trend score corresponding to each new word.

[0075] Further, when the hot trend mining module 23 is used to combine the historical new word set in the historical time period to determine the heat information and heat trend information of each new word in the current new word set, it is specifically used for: for each new word in the current new word set, determining the first heat information of the new word in the current time period according to the word frequency of the new word in the current time period and the total word frequency of the current new word set in the current time period; determining the second heat information of the new word in the historical time period according to the word frequency of the new word in the historical time period and the total word frequency of the historical new word set in the historical time period; generating the heat information of the new word according to the larger value of the first heat information and the second heat information, and the smaller value of the minimum heat information in the current time period and the minimum heat information in the historical time period; generating the heat trend information of the new word according to the heat difference between the first heat information and the second heat information, and the larger value of the first heat information and the second heat information.

[0076] Further, when the hot trend mining module 23 is used to select hot trend words from each new word according to the comprehensive heat trend score corresponding to each new word, it is specifically used for: selecting multiple candidate trend words from each new word according to the comprehensive heat trend score corresponding to each new word; performing semantic clustering on the multiple candidate trend words according to the text information where the multiple candidate trend words are located to obtain hot trend words.

[0077] Further, the event extraction module 24 is configured to extract event information from the text information where the hot trend words are located based on diversified part-of-speech templates, so as to obtain event description information corresponding to the hot trend words, including: for each hot trend word, based on the correlation and mutual exclusion relationships between the part-of-speech words defined by the diversified part-of-speech templates, extracting multiple initial phrase combinations from the text information where the hot trend words are located, and the initial phrase combinations have part-of-speech attributes; according to the frequencies of occurrence of the multiple initial phrase combinations in the text information where the hot trend words are located, selecting a target phrase combination from the multiple initial phrase combinations as the event description information corresponding to the hot trend word.

[0078] Further, the event extraction module 24 is configured to pre-generate diversified part-of-speech templates, and the diversified part-of-speech templates include at least two of the SV part-of-speech template, SVC part-of-speech template, SVO part-of-speech template, SVOO part-of-speech template, and SVOA part-of-speech template for describing event information.

[0079] Further, the target application scenario is an e-commerce origin operation scenario, and the target text information includes operation process information recorded by e-commerce origin marketers with merchants and / or text information converted from voice information generated by voice communication between e-commerce origin marketers and merchants.

[0080] It should be noted here that: the event information processing system provided in this embodiment can implement the technical solutions described in the above Figure 1a method embodiments. The specific implementation principles of the above modules or units can be referred to the corresponding content in the above method embodiments, and will not be elaborated here.

[0081] Figure 3 This is a schematic structural diagram of an electronic device provided in an exemplary embodiment of the present application. As Figure 3 shown, the electronic device includes: a memory 30a and a processor 30b. The memory 30a is used to store a computer program, and the processor 30b is coupled to the memory 30a and is configured to execute the computer program to implement the following steps:

[0082] Obtain target text information generated in the current period of the target application scenario, where the target text information includes original text information and / or text information converted from voice information; perform new word discovery on the target text information by using a method that combines coagulation degree and the longest word chain to obtain a current new word set; perform hot trend word mining based on the heat information and heat trend information of each new word in the current new word set to obtain hot trend words in the current new word set; extract event information from the text information where the hot trend words are located based on diversified part-of-speech templates to obtain event description information corresponding to the hot trend words.

[0083] Further, when the processor 30b is used to discover new words from the target text information by fusing the solidification degree and the longest word chain to obtain the current new word set, it is specifically used for: identifying words from the target text information by using an n-gram language model, and filtering the identified words by the solidification degree to obtain an initial word set, where n is an integer greater than or equal to 2; using the longest word chain algorithm to segment the target text information based on the initial word set to obtain candidate segmentations; performing backtracking filtering on the candidate segmentations based on the initial word set to obtain valid segmentations, and the valid segmentations form the current new word set.

[0084] Further, when the processor 30b is used to segment the target text information based on the initial word set by using the longest word chain algorithm to obtain candidate segmentations, it is specifically used for: constructing a trie tree based on the initial segmentation set, and searching the trie tree for the longest word chain by using a depth-first search algorithm for the target text information, and taking the longest word chain as the candidate segmentation.

[0085] Further, when the processor 30b is used to perform backtracking filtering on the candidate segmentations based on the initial word set to obtain valid segmentations, it is specifically used for: for each candidate segmentation, if the length of the candidate segmentation is less than or equal to n and the candidate segmentation appears in the initial word set, then taking the candidate segmentation as a valid segmentation; if the length of the candidate segmentation is greater than n, then sequentially splitting the candidate segmentation into multiple segmentation segments with a length of n, and taking the multiple segmentation segments as valid segmentations respectively.

[0086] Further, when the processor 30b is used to mine hot trend words based on the heat information and heat trend information of each new word in the current new word set to obtain the hot trend words in the current new word set, it is specifically used for: combining the historical new word set in the historical time period to determine the heat information and heat trend information of each new word in the current new word set; calculating the comprehensive hot trend score corresponding to each new word based on the heat information and heat trend information of each new word in the current new word set; selecting hot trend words from each new word according to the comprehensive hot trend score corresponding to each new word.

[0087] Further, when the processor 30b is used to determine the heat information and heat trend information of each new word in the current new word set in combination with the historical new word set within the historical period, it is specifically used for: for each new word in the current new word set, determining the first heat information of the new word in the current period according to the word frequency of the new word in the current period and the total word frequency of the current new word set in the current period; determining the second heat information of the new word in the historical period according to the word frequency of the new word in the historical period and the total word frequency of the historical new word set in the historical period; generating the heat information of the new word according to the larger value of the first heat information and the second heat information, and the smaller value of the minimum heat information in the current period and the minimum heat information in the historical period; generating the heat trend information of the new word according to the heat difference between the first heat information and the second heat information, and the larger value of the first heat information and the second heat information.

[0088] Further, when the processor 30b is used to select hot trend words from each new word according to the comprehensive hot trend score corresponding to each new word, it is specifically used for: selecting multiple candidate trend words from each new word according to the comprehensive hot trend score corresponding to each new word; performing semantic clustering on the multiple candidate trend words according to the text information where the multiple candidate trend words are located to obtain the hot trend words.

[0089] Further, when the processor 30b is used to extract event information from the text information where the hot trend words are located based on diverse part-of-speech templates to obtain the event description information corresponding to the hot trend words, it includes: for each hot trend word, extracting multiple initial phrase combinations from the text information where the hot trend word is located based on the correlation and mutual exclusion relationships between the parts of speech defined by the diverse part-of-speech templates, and the initial phrase combinations have part-of-speech attributes; selecting a target phrase combination from the multiple initial phrase combinations as the event description information corresponding to the hot trend word according to the frequency of occurrence of the multiple initial phrase combinations in the text information where the hot trend word is located.

[0090] Further, when the processor 30b is used to pre-generate diverse part-of-speech templates, the diverse part-of-speech templates include at least two of the SV part-of-speech template, SVC part-of-speech template, SVO part-of-speech template, SVOO part-of-speech template, and SVOA part-of-speech template for describing event information.

[0091] Further, the target application scenario is an e-commerce origin operation scenario, and the target text information includes the operation process information recorded by the e-commerce origin marketers with the merchants and / or the text information converted from the voice information generated by the voice communication between the e-commerce origin marketers and the merchants.

[0092] Further, as Figure 3 shown, the electronic device further includes: a communication component 30c, a power supply component 30d, and other components. Figure 3Only some components are schematically shown, which does not mean that the electronic device only includes Figure 3 the components shown.

[0093] It should be noted here that: The electronic device provided in this embodiment can implement the above Figure 1a technical solutions described in the method embodiments. For the specific implementation principles of the above modules or units, reference can be made to the corresponding content in the above method embodiments, which will not be elaborated here.

[0094] An exemplary embodiment of the present application provides a computer-readable storage medium storing computer programs / instructions. When the computer programs / instructions are executed by a processor, the processor is enabled to implement the steps in the above-mentioned method, which will not be elaborated here.

[0095] An exemplary embodiment of the present application provides a computer program product including computer programs / instructions. When the computer programs / instructions are executed by a processor, the processor is enabled to implement the steps in the above-mentioned method, which will not be elaborated here.

[0096] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0097] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in one or more processes in the flowchart and / or one or more blocks in the block diagram.

[0098] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in one or more processes in the flowchart and / or one or more blocks in the block diagram.

[0099] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process or multiple processes in the flowchart and / or one block or multiple blocks in the block diagram.

[0100] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0101] Memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0102] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0103] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.

[0104] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. An event information processing method, characterized in that, it includes: Obtain target text information generated in the current period for a target application scenario, where the target text information includes original text information and / or text information converted from voice information; Perform new word discovery on the target text information by combining the coagulation degree and the longest word chain to obtain a current new word set, including: use an n-gram language model to identify words in the target text information, and perform coagulation degree filtering on the identified words to obtain an initial word set, where n is an integer greater than or equal to 2; use the longest word chain algorithm to segment the target text information based on the initial word set to obtain candidate segmentations; perform backtracking filtering on the candidate segmentations based on the initial word set to obtain valid segmentations, and the valid segmentations form the current new word set; Mine hot trend words based on the heat information and heat trend information of each new word in the current new word set to obtain hot trend words in the current new word set; Based on diverse part-of-speech templates, extract event information from the text information where the hot trend words are located to obtain a target phrase combination as the event description information corresponding to the hot trend words. The diverse part-of-speech templates are used to extract a phrase combination reflecting the event information based on the relevant and mutually exclusive relationships between the defined parts of speech.

2. The method according to claim 1, characterized in that, using the longest word chain algorithm to segment the target text information based on the initial word set to obtain candidate segmentations, including: Construct a dictionary tree based on the initial word set, and use a depth-first search algorithm to search the dictionary tree for the target text information to obtain the longest word chain, and use the longest word chain as the candidate segmentation.

3. The method according to claim 1, characterized in that, performing backtracking filtering on the candidate segmentations based on the initial word set to obtain valid segmentations, including: For each candidate segmentation, if the length of the candidate segmentation is less than or equal to n and the candidate segmentation appears in the initial word set, then use the candidate segmentation as a valid segmentation; If the length of the candidate segmentation is greater than n, then sequentially segment the candidate segmentation into multiple segmentation segments of length n, and use the multiple segmentation segments as valid segmentations respectively.

4. The method according to claim 1, characterized in that, mining hot trend words based on the heat information and heat trend information of each new word in the current new word set to obtain hot trend words in the current new word set, including: Combining the historical new word set in the historical period to determine the heat information and heat trend information of each new word in the current new word set; Based on the heat information and heat trend information of each new word in the current new word set, calculate the comprehensive hot trend score corresponding to each new word; Select hot trend words from each new word according to the comprehensive hot trend score corresponding to each new word.

5. The method according to claim 4, characterized in that, combining the historical new word set in the historical period to determine the heat information and heat trend information of each new word in the current new word set, including: For each new word in the current new word set, determine the first heat information of the new word in the current time period according to the word frequency of the new word in the current time period and the total word frequency of the current new word set in the current time period; Determine the second heat information of the new word in the historical time period according to the word frequency of the new word in the historical time period and the total word frequency of the historical new word set in the historical time period; Generate the heat information of the new word according to the larger value of the first heat information and the second heat information, and the smaller value of the minimum heat information in the current time period and the minimum heat information in the historical time period; Generate the heat trend information of the new word according to the heat difference between the first heat information and the second heat information, and the larger value of the first heat information and the second heat information.

6. The method according to claim 4, wherein, Selecting hot trend words from each new word according to the comprehensive score of the hot trend corresponding to each new word, including: Selecting a plurality of candidate trend words from each new word according to the comprehensive score of the hot trend corresponding to each new word; Performing semantic clustering on the plurality of candidate trend words according to the text information where the plurality of candidate trend words are located to obtain hot trend words.

7. The method according to any one of claims 1-6, wherein, Performing event information extraction on the text information where the hot trend word is located based on a diversified part-of-speech template to obtain a target phrase combination as the event description information corresponding to the hot trend word, including: Extracting a plurality of initial phrase combinations from the text information where the hot trend word is located based on the correlation and mutual exclusion relationships between the parts of speech defined by the diversified part-of-speech template, and the initial phrase combinations have part-of-speech attributes; Selecting a target phrase combination from the plurality of initial phrase combinations according to the frequency of occurrence of the plurality of initial phrase combinations in the text information where the hot trend word is located as the event description information corresponding to the hot trend word.

8. The method according to claim 7, wherein, further comprising: Pre-generating diversified part-of-speech templates, and the diversified part-of-speech templates include at least two of the SV part-of-speech template, SVC part-of-speech template, SVO part-of-speech template, SVOO part-of-speech template, and SVOA part-of-speech template for describing event information.

9. The method according to any one of claims 1-6, wherein, The target application scenario is an e-commerce origin operation scenario, and the target text information includes the operation process information recorded by e-commerce origin marketers with merchants and / or the text information converted from the voice information generated by the online audio and video communication between e-commerce origin marketers and merchants.

10. An event information processing system, wherein, comprising: A text acquisition module for acquiring target text information generated in the current time period in the target application scenario, and the target text information includes original text information and / or text information converted from voice information; A new word discovery module for discovering new words in the target text information by integrating coagulation degree and the longest word chain to obtain a current new word set; A hot trend mining module, configured to perform hot trend word mining based on the popularity information and popularity trend information of each new word in the current new word set, so as to obtain the hot trend words in the current new word set; An event extraction module, configured to perform event information extraction on the text information where the hot trend words are located based on a diversified part-of-speech template, so as to obtain a target phrase combination as the event description information corresponding to the hot trend words, and the diversified part-of-speech template is used to extract a phrase combination reflecting the event information based on the relevant and mutually exclusive relationships between the defined parts of speech; Among them, when the new word discovery module performs new word discovery on the target text information in a manner that combines solidification degree and the longest word chain to obtain the current new word set, it is configured to: perform word recognition on the target text information using an n-gram language model, and perform solidification degree filtering on the recognized words to obtain an initial word set, where n is an integer greater than or equal to 2; use the longest word chain algorithm to perform word segmentation on the target text information based on the initial word set to obtain candidate word segments; perform backtracking filtering on the candidate word segments based on the initial word set to obtain valid word segments, and the valid word segments form the current new word set.

11. An electronic device Characterized in that It includes: A memory and a processor, the memory is used to store a computer program, and the processor is coupled with the memory and is configured to execute the computer program to implement the steps in the method according to any one of claims 1-9.

12. A computer-readable storage medium storing a computer program Characterized in that When the computer program is executed by a processor, the processor is caused to implement the steps in the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Hot spot information acquisition method and device, server and medium

    CN111159557A

  • Data processing method and device, storage medium and equipment

    CN113836914A