Event context generation method and device, computer equipment and storage medium

By using a large language model to analyze topic description information, generate keyword lists and logical expressions, extract and standardize time feature text, group and cluster, and finally generate event sequence data and event context, the problem of missing non-standard format time text in the prior art is solved, and the accuracy and readability of event context is improved.

CN120123602APending Publication Date: 2025-06-10BEIJING ZHIHUI XINGGUANG INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510177053.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art tends to miss text containing non-standard format time when sorting out the event context, resulting in poor readability and accuracy of the event development context.

Method used

By determining the keyword list based on the topic description information input by the customer, generating logical expressions, collecting matching news data, extracting feature sentences containing time feature text, standardizing processing, grouping and clustering, generating event sequence data, and finally generating event context based on event sequence data and logical expressions.

Benefits of technology

It improves the comprehensiveness and accuracy of data collection, enhances the processing ability of non-standard time text, and improves the accuracy and readability of event context.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123602A_ABST
    Figure CN120123602A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data analysis, and discloses an event context generation method and device, computer equipment and a storage medium, and aims to analyze theme description information, recommend synonyms and generate a keyword list through a large language model, simplify the operation difficulty of clients and improve the data comprehensiveness. According to the method, the news data are collected through the logic expression corresponding to the keyword list, the time feature text in the news data is standardized, the target short sentence only containing the standard time is obtained, the standard time corresponding to various texts is deduced through the context reasoning ability of the large language model, and the data recall rate and accuracy are improved. The target short sentences are grouped and clustered by using a large language model vector, so that the merging capability of similar events is improved. And selecting the core short sentences of each cluster to generate event sequence data, processing the event sequence data in combination with a logic expression, and filtering uncorrelated events. And finally, generating event venation with high readability by utilizing text output and summarization capability of the large language model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data analysis, and in particular to an event context generation method, apparatus, computer device, and storage medium. Background Art

[0002] With the development of the Internet, information has gradually become complex. When a social hot spot appears, due to the interference of other content, it is difficult for users to understand the event evolution information such as the start, development, and end of the event.

[0003] In the prior art, after a customer sets a search keyword, the system will collect event-related data according to the keyword and use a deep learning model to extract the event summary information to generate the final event development context. However, the above method is likely to miss some non-standard format times in the event-related data, resulting in the lack of time context information of the event and poor readability and accuracy of the generated event development context. Summary of the Invention

[0004] In view of this, the present invention provides an event context generation method, apparatus, computer device, and storage medium to solve the problem that the prior art is likely to miss the text containing non-standard format times when sorting out the event context, resulting in poor readability and accuracy of the generated event development context.

[0005] In a first aspect, the present invention provides an event context generation method, which includes:

[0006] Determining a keyword list based on the theme description information input by the customer, generating a logical expression based on the keyword list, and collecting news data that matches the logical expression;

[0007] Extracting candidate short sentences containing time feature characters from the news data, and screening out feature short sentences containing time feature text from the candidate short sentences; wherein, the time feature text is text arranged according to a preset time format by the time feature characters;

[0008] Normalizing the time feature text in the feature short sentences to obtain target short sentences containing only standard times, and grouping the target short sentences according to the standard times in the target short sentences;

[0009] Clustering the target short sentences in each group, and determining the core short sentence in the target short sentences of each cluster, wherein each group contains at least one cluster;

[0010] Arranging the core short sentences based on the standard times in the core short sentences to generate event sequence data, and generating an event context based on the event sequence data and the logical expression.

[0011] Beneficial effects: The present invention analyzes the topic description information through a large language model to help customers generate a keyword list, simplifies the operation difficulty of customers, and recommends a batch of synonyms to generate a keyword list by utilizing the generalization ability of the large language model, which is beneficial to improving the comprehensiveness of data collection. Then, the news data that matches is collected by using the logical expression corresponding to the keyword list, improving the accuracy of the collected data. Next, the characteristic short sentences containing time characteristic texts in the news data are extracted, the time characteristic texts are standardized to obtain target short sentences containing only standard times, and the context reasoning ability of the large language model is used to infer the standard times corresponding to various complex non-standard time texts in the short sentences, improving the data recall rate and accuracy. Subsequently, the target short sentences are grouped according to the standard time and clustered using the large language model vectors, and the same events at the same time are merged together. Among them, by using the large language model that can express complex semantic relationships, the merging ability of similar events can be further improved. Finally, the core short sentences of each clustering cluster are selected, event sequence data is generated based on the core short sentences, and the large language model combines the logical expression to process the event sequence data, which can more accurately determine the main line of event development, filter out irrelevant events, and improve the accuracy of the event context. Then, the powerful text output and summarization ability of the large language model is used to generate the event context, improving the readability of the event context.

[0012] In an alternative embodiment, the preset time formats include standard time formats and non-standard time formats. Among them, the standard time formats include at least some of the following items: year, month, day, hour, minute, second, and time nouns used to represent time periods in a day; the time characteristic texts include standard time texts in which time characteristic characters are arranged according to the standard time format and non-standard time texts in which time characteristic characters are arranged according to the non-standard time format;

[0013] Standardizing the time characteristic texts in the characteristic short sentences to obtain target short sentences containing only standard times includes:

[0014] If it is detected that the time characteristic text in the characteristic short sentence is a standard time text, the standard time text in the characteristic short sentence is extracted using a regular expression, and the standard time text in the characteristic short sentence is replaced with the standard time corresponding to the standard time text to obtain a target short sentence containing only the standard time;

[0015] If it is detected that the time characteristic text in the characteristic short sentence is a non-standard time text, the trained large language model is used to replace the non-standard time text in the characteristic short sentence with the standard time corresponding to the non-standard time text to obtain a target short sentence containing only the standard time.

[0016] Beneficial effects: In the present invention, the extracted standard time is used to replace the original standard time text or non-standard time text in the feature short sentence, so as to reduce the difficulty of analysis when using a large language model to extract time subsequently. Moreover, by replacing the non-standard time text in the short sentence with the extracted standard time, all the times in the short sentences are made into standard times. When calculating similarity, the similarity degree of similar short sentences will be higher. When using a large language model to analyze the event context, there is no need for the large language model to deduce the specific time anymore, and both the accuracy and performance can be improved.

[0017] In an alternative embodiment, using the trained large language model to replace the non-standard time text in the feature short sentence with the standard time corresponding to the non-standard time text, obtaining a target short sentence containing only standard time, includes:

[0018] Determine the context-related feature short sentences of the feature short sentence containing non-standard time text in the news data, and the context-related feature short sentences are feature short sentences containing standard time text;

[0019] Using the trained large language model based on the standard time text in the context-related feature short sentences to predict the standard time corresponding to the non-standard time text in the feature short sentence containing non-standard time text;

[0020] According to the standard time corresponding to the non-standard time text, replace the non-standard time text in the feature short sentence to obtain a target short sentence containing only standard time.

[0021] Beneficial effects: For the feature short sentence containing non-standard time text, the present invention uses a large language model to extract the context-related feature short sentences of this feature short sentence in the news data, and uses the standard time text in the context-related feature short sentences to infer the standard time corresponding to the non-standard time text. In this way, firstly, it avoids missing related events and prevents the loss of event context information; secondly, in subsequent clustering, it can be aggregated according to the standard time corresponding to the time feature short sentence, and has a better aggregation effect.

[0022] In an alternative embodiment, clustering the target short sentences in each group and determining the core short sentence in the target short sentences of each cluster, includes:

[0023] Using the trained large language model to generate short sentence vectors of the target short sentences, clustering the target short sentences in each group based on the short sentence vectors, and determining the target short sentences included in each cluster in each group;

[0024] According to the arithmetic average of the short sentence vectors of all the target short sentences in each cluster, calculate the center vector of each cluster;

[0025] According to the target short sentences with the shortest cosine distance from the center vector in each cluster, the core short sentences of each cluster are obtained.

[0026] Beneficial effects: The present invention uses short sentence vectors to cluster all target short sentences in each standard time group, so that short sentences with complex semantics are aggregated together, and then the same events are aggregated in one cluster, and the core events of the cluster are represented by selecting the core short sentences of each cluster, avoiding repeated descriptions of the same event when generating the event context later, and improving the accuracy of the event context description.

[0027] In an alternative embodiment, the candidate short sentences containing time feature characters in the news data are extracted, including:

[0028] The news data is split using delimiters to obtain news short sentences; wherein, the delimiters are terminating punctuation marks used to indicate the end of a sentence.

[0029] The news short sentences containing time feature characters are detected to obtain candidate short sentences.

[0030] Beneficial effects: The present invention performs some filtering work before formally extracting the time feature text in the short sentences. First, the candidate short sentences containing relevant time feature characters are extracted, which facilitates the subsequent extraction and processing of the time feature text and improves the overall processing speed.

[0031] In an alternative embodiment, an event context is generated based on the event sequence data and logical expressions, including:

[0032] Filter the core short sentences in the event sequence data based on the logical expressions to obtain target core short sentences.

[0033] Generate an event context based on the target core short sentences.

[0034] Beneficial effects: The present invention first filters the core short sentences using the logical expressions used in news data collection, thereby filtering out the irrelevant data in the event sequence data and improving the accuracy of the event context data. Then, the large language model is used to merge and summarize the relevant target core short sentences to generate the event context, improving the readability of the event context data.

[0035] In an alternative embodiment, a keyword list is determined based on the theme description information input by the customer, including:

[0036] Use the trained large language model to extract the key information in the theme description information and generate the near-synonym association information of the key information.

[0037] In response to the user's selection operation on the key information and near-synonym association information, determine at least one keyword selected by the user and generate a keyword list.

[0038] Beneficial effects: The present invention uses a trained large language model to extract key information and generate near-synonym related information from the topic description information input by the user, preventing the omission of event-related information and improving the comprehensiveness of the collected information. Moreover, the user can directly select the final keywords from the key information and near-synonym related information generated by the large language model, thereby determining the final keyword list, reducing the usage threshold for the user to set keywords, and being beneficial to improving the user experience.

[0039] In a second aspect, the present invention provides an event context generation device, which includes:

[0040] A first processing module, configured to determine a keyword list based on the topic description information input by the customer, generate a logical expression based on the keyword list, and collect news data that matches the logical expression;

[0041] A second processing module, configured to extract candidate short sentences containing time feature characters from the news data, and screen out feature short sentences containing time feature texts from the candidate short sentences; wherein, the time feature text is a text in which time feature characters are arranged according to a preset time format;

[0042] A third processing module, configured to standardize the time feature text in the feature short sentences to obtain target short sentences containing only standard times, and group the target short sentences according to the standard times in the target short sentences;

[0043] A fourth processing module, configured to cluster the target short sentences in each group, and determine the core short sentences in the target short sentences of each cluster, wherein each group contains at least one cluster;

[0044] A fifth processing module, configured to arrange the core short sentences based on the standard times in the core short sentences to generate event sequence data, and generate an event context based on the event sequence data and the logical expression.

[0045] In a third aspect, the present invention provides a computer device, including: a memory and a processor, which are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the event context generation method according to the first aspect or any corresponding embodiment thereof.

[0046] In a fourth aspect, the present invention provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the event context generation method according to the first aspect or any corresponding embodiment thereof. Description of the Drawings

[0047] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the accompanying drawings required for the description of the specific embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0048] Figure 1 It is a schematic flowchart of an event context generation method according to an embodiment of the present invention;

[0049] Figure 2 It is a schematic flowchart of another event context generation method according to an embodiment of the present invention;

[0050] Figure 3 It is a schematic flowchart of yet another event context generation method according to an embodiment of the present invention;

[0051] Figure 4 It is a structural block diagram of an event context generation system according to an embodiment of the present invention;

[0052] Figure 5 It is a structural block diagram of an event context generation device according to an embodiment of the present invention;

[0053] Figure 6 It is a schematic hardware structure diagram of a computer device according to an embodiment of the present invention. Specific Embodiments

[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0055] In recent years, with the continuous improvement of the Internet ecosystem, the speed of information dissemination has become faster and faster, and the amount of information data has also shown an explosive growth. People are increasingly likely to obtain information. However, this has also brought problems such as content duplication, the mixture of junk content, and information fragmentation. When a certain hot event occurs in society, various news reports and interpretations are flooding, and it has become more difficult for people to understand the event evolution information such as the beginning, development, and end of a certain event from the news.

[0056] The related technologies mainly generate the event development context through the following steps:

[0057] 1) The customer sets keywords, and the system collects event-related data according to the keywords.

[0058] 2) Use means such as regular expressions to obtain time-related text in event-related data and standardize the time format.

[0059] 3) Adopt deep learning models such as BERT to extract the keyword of the event subject and form the event summary information.

[0060] 4) Use word2vec vectors to establish an event theme vector through event-related data, calculate the similarity with the event summary, and filter out events irrelevant to the theme.

[0061] 5) Cluster the event summary information using word2vec vectors and remove duplicates for similar events.

[0062] 6) Generate the summary information of each specific event on the context through deep learning models such as BURT.

[0063] 7) Sort the summary information according to time to generate the final event development context.

[0064] However, in the above method, customers input keywords based on their own understanding, which may result in omissions and cause the problem of incomplete information collection. In addition, the keyword setting function of many such products is also very complex. It takes a long time for customers to write a group of keywords, and the usage threshold is high.

[0065] Moreover, for the extraction of non-standard format times such as "last Monday at 3:00" and "2:00 pm on the second day", it is difficult for the above method to exhaust all combinations, resulting in omissions and recognition errors. In addition, such times often need to be deduced based on the context time. The omission and recognition errors of information also lead to incorrect deductions, resulting in poor accuracy and readability of the deduced event context.

[0066] In addition, it is difficult to cluster texts with complex semantics using word2vec vectors. The same events are not easily clustered together, resulting in duplicate descriptions of the same event in the generated event context. It is also difficult to discover the associated events of a certain event that require reasoning, resulting in missing context information and poor readability of the extracted event context description.

[0067] Therefore, the embodiment of the present invention provides an event context generation method, which combines a large language model to optimize the method of generating an event development context. First, group the news short sentences according to the standardized time, and then improve the clustering effect of similar events through the embedding vector of the large language model, reducing the interference of non-standard format times on clustering. Moreover, use the large language model to assist in extracting event keywords, extracting non-standard times, mining associated events, filtering out non-associated events, generating event contexts, etc., greatly improving the data recall and accuracy of the event context.

[0068] According to an embodiment of the present invention, an embodiment of an event context generation method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0069] In this embodiment, an event context generation method is provided, which can be used for devices that sort out event contexts, such as mobile phones, tablets, etc. Figure 1 It is a flowchart of the event context generation method according to an embodiment of the present invention, as Figure 1 shown, and this process includes the following steps:

[0070] Step S101, determine a keyword list based on the theme description information input by the customer, generate a logical expression based on the keyword list, and collect news data that matches the logical expression.

[0071] Specifically, a user interface is provided to receive the theme description information input by the customer, and the trained large language model is used to extract the keywords in the theme description information and return them to the customer for selection. The customer can select the semantic association information between the keywords according to the keywords provided by the system, and then generate a keyword list. Further, the system can generate a logical expression corresponding to the keyword list according to the keyword list submitted by the customer, and collect the matching news data through the logical expression. It should be noted that a large language model is a model that can understand and generate human language by training on a large-scale text dataset. The number of parameters of these models is very large, and a large number of parameters enable the model to learn rich language knowledge and patterns.

[0072] In this way, since the logical expression contains the semantic association information between the keywords, the accuracy of data collection can be improved during data collection, and news data that contains keywords but whose semantic structure does not match the customer's needs can be avoided. For example, based on the event information that the user wants to understand, both keyword 1 and keyword 2 need to be included. If only retrieved based on keywords, it is very likely to collect news data that only contains keyword 1 or only contains keyword 2, resulting in a very large amount of news data to be analyzed subsequently and not matching the customer's needs. By generating the logical expression (keyword 1) and (keyword 2) and collecting data, it can be ensured that the collected news data contains both keyword 1 and keyword 2, thus avoiding collecting garbage data, improving the accuracy of data collection, and effectively reducing the amount of data analysis.

[0073] In some alternative embodiments, after obtaining the topic description information input by the user, the trained large language model is used to extract keywords from the topic description information, and the generalization ability of the large language model is utilized to recommend multiple groups of synonyms for the keywords. A keyword list and a logical expression are generated based on the keywords and their multiple groups of synonyms, thereby helping the system collect as much relevant event information as possible, avoiding omission, and improving the comprehensiveness of data collection.

[0074] In this embodiment, news data is collected using the keyword list and the logical expression, without the need for the customer to set keywords themselves, avoiding the situation of missing keywords input by the customer, thereby improving the comprehensiveness of news data collection. Moreover, the operation process of setting keywords by the customer is simplified, and the usage threshold of the user is reduced.

[0075] Step S102: Extract the candidate short sentences containing time feature characters from the news data, and screen out the feature short sentences containing time feature text from the candidate short sentences; wherein, the time feature text is the text in which the time feature characters are arranged according to a preset time format.

[0076] Specifically, the time feature characters are characters with implicit time features such as year, month, day, hour, minute, second, morning, day, etc. The multi-pattern string matching algorithm can be used to extract the short sentences containing time feature characters in the news data to obtain the candidate short sentences containing time feature characters.

[0077] It should be noted that even if some candidate short sentences contain time feature characters, it does not necessarily mean that these short sentences contain extractable time information. For example, the short sentence text "red sun" contains the time feature character "day", and this short sentence will be extracted as a candidate short sentence, but "red sun" is not a time feature text that can be analyzed for time information.

[0078] Therefore, it is necessary to analyze and detect the arrangement format of the time feature characters in the candidate short sentences to determine which of the candidate short sentences are feature short sentences containing time feature text.

[0079] In some alternative embodiments, there are also various arrangement formats for the time feature characters that conform to the time feature text, and the arrangement format of the time feature characters can be preset. For example, the preset time format includes a standard time format and a non-standard time format, wherein the standard time format includes at least some of the following items: year, month, day, hour, minute, second, and time nouns used to represent time periods in a day. The time feature text includes standard time text in which the time feature characters are arranged according to the standard time format and non-standard time text in which the time feature characters are arranged according to the non-standard time format.

[0080] Exemplarily, the time nouns used to characterize time periods in a day can be special time periods in a day such as evening, dusk, midnight, morning, noon, afternoon, early morning, etc., which can be preset according to actual scenario requirements. That is to say, texts such as "yyyy-MM-dd HH:mm:ss", "yyyy-MM-dd", "yyyy-MM-dd HH:mm", "yyyy-MM-dd HH", and "yyyy-MM-dd morning" are all standard time texts in which time characteristic characters are arranged in the standard time format.

[0081] It should be noted that since the year-month-day information of the data can generally be extracted during data extraction, the time accuracy of the standard time format is at least accurate to the year-month-day. However, for some data with higher time accuracy, it can be accurate to the year-month-day-hour-minute-second, which can be specifically selected according to the actual scenario.

[0082] In addition, time characteristic characters that are not arranged in the standard time format are not necessarily non-standard time texts. Therefore, it is necessary to preset the non-standard time format. For example, texts such as tomorrow, the second day, next Monday, the day after tomorrow, the day before yesterday, the second day HH:mm, New Year's Day, and the second day afternoon, which can deduce the specific date based on the context, are regarded as non-standard time texts in which time characteristic characters are arranged in the non-standard time format.

[0083] Step S103, standardize the time characteristic texts in the characteristic short sentences to obtain target short sentences containing only standard time, and group the target short sentences according to the standard time in the target short sentences.

[0084] Specifically, the time characteristic text may be a standard time text such as "yyyy-MM-dd", or it may be a non-standard time text such as "the second day". It is necessary to select the corresponding method according to the type of the time characteristic text to standardize the time characteristic text into a standard time, and group the target short sentences according to the standard time. The target short sentences representing the same standard time will be grouped into the same group.

[0085] Exemplarily, if the characteristic short sentence with the standard time text such as "yyyy-MM-dd" and the characteristic short sentence with the non-standard time text such as "the second day" both describe the same event at the same time, by standardizing the time characteristic texts in these two characteristic short sentences and grouping them according to the standardized standard time, in this way, when clustering subsequently, events that actually occur at the same time are very likely to be clustered together, preventing the same type of events from being clustered into different clusters due to different styles of time characteristic texts in the text. Thus, the same type of events can be identified according to the standard time, preventing the omission or misidentification of events and improving the accuracy of data extraction.

[0086] Step S104: Cluster the target short sentences in each group, and determine the core short sentence in the target short sentences of each cluster. Each group contains at least one cluster.

[0087] Specifically, by clustering the target short sentences in the group representing the same standard time and determining the core short sentence that can represent the core event of each cluster, short sentence texts with complex semantics can be aggregated together, avoiding repeated descriptions of the same event when the same events are not aggregated together.

[0088] Step S105: Arrange the core short sentences based on the standard time in the core short sentences to generate event sequence data, and generate an event context based on the event sequence data and logical expressions.

[0089] Specifically, process the event sequence data through logical expressions to avoid unrelated events in the event context, thereby improving the presentation effect of the event context.

[0090] The event context generation method provided in this embodiment analyzes the topic description information through a large language model to help the customer generate a keyword list, simplifies the operation difficulty of the customer, and uses the generalization ability of the large language model to recommend a batch of synonyms to generate a keyword list, which is beneficial to improving the comprehensiveness of data collection. Then, collect the news data that matches the logical expression corresponding to the keyword list to improve the accuracy of the collected data. Then, extract the feature short sentences containing time feature texts in the news data, standardize the time feature texts to obtain target short sentences containing only standard time, and use the context reasoning ability of the large language model to infer the standard time corresponding to various complex non-standard time texts in the short sentences, improving the data recall rate and accuracy. Next, group the target short sentences according to the standard time and use the large language model vector for clustering to merge the same events at the same time. Among them, by using a large language model that can express complex semantic relationships, the merging ability of similar events can be further improved. Finally, select the core short sentence of each cluster, generate event sequence data based on the core short sentence, and the large language model processes the event sequence data in combination with logical expressions, which can more accurately determine the main line of event development, filter out unrelated events, improve the accuracy of the event context, and then use the powerful text output and summarization ability of the large language model to generate the event context, improving the readability of the event context.

[0091] In this embodiment, an event context generation method is provided, which can be used for devices for sorting out event contexts, such as mobile phones, tablet computers, etc. Figure 2 It is a flowchart of the event context generation method according to the embodiment of the present invention, as Figure 2 shown. This process includes the following steps:

[0092] Step S201: Determine a keyword list based on the topic description information input by the customer, generate a logical expression based on the keyword list, and collect news data that matches the logical expression.

[0093] Specifically, the above step S201 includes:

[0094] Step S2011: Use the trained large language model to extract key information from the topic description information, generate near-synonym association information for the key information, and determine at least one keyword selected by the user and generate a keyword list in response to the user's selection operation on the key information and near-synonym association information.

[0095] Specifically, the customer inputs the topic description information of an event to be concerned about in the setting interface. The topic description information can be a sentence or keywords related to the event, including content such as time, place, subject, and behavior. The trained large language model extracts key information such as keywords based on this content, and recommends multiple groups of near-synonyms according to the keywords to help the system collect as much relevant event information as possible and avoid omission.

[0096] Furthermore, after the generated keywords and recommended near-synonyms are returned to the customer, the customer can modify them as needed to form the final keyword list.

[0097] Step S2012: Generate a logical expression based on the keyword list, and collect news data that matches the logical expression.

[0098] Exemplarily, if the keyword list includes {[keyword1, near-synonym of keyword1], [keyword2, near-synonym of keyword2], [keyword3]}, then the generated logical expression is: (keyword1 or near-synonym of keyword1) and (keyword2 or near-synonym of keyword2) and (keyword3). Through the logical expression corresponding to the keyword list, the semantic structure in the user's topic description information is retained to collect matching news data, improving the accuracy of data collection and avoiding collecting data that contains keywords but has inconsistent semantic structures.

[0099] In the embodiment of the present application, the trained large language model is used to extract key information and generate near-synonym association information for the topic description information input by the user, preventing the omission of event association information and improving the comprehensiveness of the collected information. Moreover, the user can directly select the final keywords from the key information and near-synonym association information generated by the large language model, thereby determining the final keyword list, reducing the usage threshold for the user to set keywords, and being beneficial to improving the user experience.

[0100] Step S202: Extract candidate short sentences containing time feature characters from the news data, and screen out feature short sentences containing time feature texts from the candidate short sentences; wherein, the time feature text is a text formed by arranging time feature characters according to a preset time format.

[0101] Specifically, the above step S202 includes:

[0102] Step S2021: Split the news data using delimiters to obtain news short sentences, and detect the news short sentences containing time feature characters to obtain candidate short sentences.

[0103] Specifically, the delimiter is a terminating punctuation mark used to indicate the end of a sentence, such as a period, a question mark, and an exclamation mark. First, the content of the collected news data is segmented into several news short sentences through sentence terminating punctuation marks such as periods, question marks, and exclamation marks, and then a multi-pattern string matching algorithm such as the Aho-Corasick algorithm is used to detect whether the news short sentences contain time feature characters, and only the short sentences containing feature characters are retained to obtain candidate short sentences.

[0104] In this embodiment, the multi-pattern string matching algorithm has very high string retrieval performance and is very suitable for the system to perform some filtering work before formally extracting the time feature text in the short sentences, first extract the candidate short sentences containing relevant time feature characters, so as to facilitate subsequent text processing and overall improve the processing speed. After detecting the candidate short sentences containing time feature characters, the system sends the candidate short sentences and the list of time feature characters contained in the candidate short sentences to the subsequent module for extraction of time feature texts.

[0105] Step S2022: Screen out feature short sentences containing time feature texts from the candidate short sentences; wherein, the time feature text is a text formed by arranging time feature characters according to a preset time format. For details, please refer to Figure 1 the description of step S102 in the illustrated embodiment, which will not be elaborated here.

[0106] The above embodiment splits the news data into news short sentences and extracts candidate short sentences containing time feature characters, and then screens out feature short sentences containing time feature texts from the candidate short sentences, which facilitates subsequent time feature text processing and improves data processing efficiency.

[0107] Step S203: Standardize the time feature text in the feature short sentences to obtain target short sentences containing only standard times, and group the target short sentences according to the standard times in the target short sentences.

[0108] Specifically, the above step S203 includes:

[0109] Step S2031, if the time feature text in the feature short sentence is a standard time text, use regular expressions to extract the standard time text in the feature short sentence, and replace the standard time text in the feature short sentence with the corresponding standard time, obtaining a target short sentence that only contains the standard time.

[0110] Specifically, if the time feature text in the feature short sentence is a standard time text, it can be a standard time in the format of "yyyy-MM-dd HH:mm:ss", or a fuzzy time without specific hours, minutes, and seconds such as "yyyy-MM-dd noon". Use regular expressions to try to extract the standard time text in the feature short sentence (such as xx year, xx month xx day, 2:00 pm on xx day, etc.) and generate the corresponding standard time.

[0111] It should be noted that the precision of the standard time needs to be the same as that of the original time feature text, and the 24-hour time system is adopted. For example, if the time feature characters contained in the original time feature text of the feature short sentence are "May 2024", its formatted standard time is "May 2024". Another example is formatting "2 pm on May 1, 2024" into "14:00 on May 1, 2024", and formatting "morning of May 1, 2024" into "morning of May 1, 2024". In addition, for text that does not contain the year, the year included in the news data release time can be used to obtain the standard time.

[0112] Step S2032, if the time feature text in the feature short sentence is a non-standard time text, use the trained large language model to replace the non-standard time text in the feature short sentence with the corresponding standard time, obtaining a target short sentence that only contains the standard time.

[0113] Specifically, determine the context-related feature short sentences of the feature short sentence containing the non-standard time text in the news data. The context-related feature short sentences are feature short sentences containing standard time text. Then, use the trained large language model to predict the standard time corresponding to the non-standard time text in the feature short sentence containing the non-standard time text based on the standard time text in the context-related feature short sentences. Finally, replace the non-standard time text in the feature short sentence with the standard time corresponding to the non-standard time text, obtaining a target short sentence that only contains the standard time.

[0114] In this embodiment, for the characteristic short sentences in which the standard time text cannot be normally extracted using regular expressions, or the characteristic short sentences containing non-standard time text, a large language model is used for extraction. In this case, the context-related time in a news data will be used. First, all candidate short sentences containing time characteristic characters in a news data are arranged in the order of appearance in the news data. The large language model first extracts the times in various formats (including standard time text and non-standard time text) from these candidate short sentences, and infers the standard times corresponding to all non-standard time texts that do not meet the standard time format based on the characteristic short sentences with clear time in the context, that is, the characteristic short sentences containing standard time text. For example, for non-standard time texts such as "next year" and "tomorrow", their corresponding standard times can be inferred through the accurate time in the context.

[0115] The present invention aims at the characteristic short sentences containing non-standard time text, uses a large language model to extract the context-related characteristic short sentences of the characteristic short sentences in the news data, and uses the standard time text in the context-related characteristic short sentences to infer the standard time corresponding to the non-standard time text. In this way, first, the omission of related events is avoided, and the loss of the context information of the events is prevented; second, in subsequent clustering, aggregation can be performed according to the standard time corresponding to the time characteristic short sentences, and better aggregation effects can be obtained.

[0116] In the embodiment of the present invention, the extracted standard time is used to replace the original standard time text or non-standard time text in the characteristic short sentences, so as to reduce the difficulty of analysis when using a large language model to extract time subsequently. Moreover, by using the extracted standard time to replace the non-standard time text in the short sentences, the times in all short sentences are changed into standard times. When calculating the similarity, the similarity degree of similar short sentences will be higher. When using a large language model to analyze the event context, the large language model does not need to deduce the specific time anymore, and both the accuracy and performance can be improved.

[0117] Exemplarily, for two characteristic short sentences "At 14:00 on September 1st, a car accident occurred on Jianshe Street" and "At 2:00 pm on Monday, a traffic accident occurred on Jianshe Street", if the article release time is 2024, after replacement with the standard time, they become "At 14:00 on September 1st, 2024, a car accident occurred on Jianshe Street" and "At 14:00 on September 1st, 2024, a traffic accident occurred on Jianshe Street". In this way, the target short sentences obtained after replacement have a higher similarity, which is beneficial to the clustering and analysis and judgment of the model.

[0118] Step S204: Cluster the target short sentences in each group, and determine the core short sentence in the target short sentences of each cluster, where each group includes at least one cluster.

[0119] Specifically, the above step S204 includes:

[0120] Step S2041: Use the trained large language model to generate the short sentence vectors of the target short sentences, and cluster the target short sentences in each group based on the short sentence vectors to determine the target short sentences included in each cluster in each group.

[0121] Specifically, after grouping all the target short sentences according to the standard time, for each group, generate the short sentence vectors corresponding to the target short sentences through the large language model, such as embedding vectors, and cluster the short sentences in the same group to merge similar events. The clustering method can compare the target short sentences pairwise through the cosine distance. When the cosine distance is within the threshold (such as 0.2), it is considered that the two short sentences describe the same event and are clustered into the same cluster. Among them, the cosine distance formula between the short sentence vectors of any two target short sentences is as follows:

[0122]

[0123] where A and B are the short sentence vectors of two target short sentences to be compared, A·B is the dot product of the two short sentence vectors, AB is the product of the norms of the two short sentence vectors, and d is the cosine distance between the two short sentence vectors.

[0124] It should be noted that the number of clusters formed by each group can be one, two, or multiple, and the present invention is not limited thereto. Generally, the more accurate the generated logical expression is, the stronger the correlation between the collected news data is, and the fewer the clusters formed in each group when clustering according to the standard time.

[0125] Step S2042: Calculate the central vector of each cluster according to the arithmetic mean of the short sentence vectors of all the target short sentences in each cluster, and obtain the core short sentence of each cluster according to the target short sentence with the shortest cosine distance from the central vector in each cluster.

[0126] Specifically, when the clustering of each group is completed, a certain target short sentence needs to be selected in each cluster as the core short sentence of the current cluster to represent the event represented by the cluster. The calculation method is to obtain the central vector of the cluster by taking the arithmetic mean of all the short sentence vectors in each cluster, and then calculate the cosine distance from each target short sentence to the central vector of its respective cluster, and select the target short sentence with the smallest cosine distance as the core short sentence of the cluster.

[0127] In some alternative embodiments, there are m k-dimensional short sentence vectors, and the i-th short sentence vector is represented as v i (v i1 ,v i2 ,v i3 ,...,v ik )(0 < i ≤ m), and the central vector is The formula for calculating the j-th dimension (0 < j ≤ k) of the central vector through arithmetic mean is as follows:

[0128]

[0129] In the embodiments of the present invention, short sentence vectors are used to cluster all target short sentences in each standard time group, so that short sentences with complex semantics are aggregated together, and then the same events are aggregated in a clustering cluster. By selecting the core short sentence of each clustering cluster to represent the core event of the clustering cluster, repeated descriptions of the same event are avoided when generating the event context later, and the accuracy of the event context description is improved.

[0130] Step S205: Arrange the core short sentences based on the standard time in the core short sentences to generate event sequence data, and generate an event context based on the event sequence data and logical expressions.

[0131] Specifically, the above step S205 includes:

[0132] Step S2051: Arrange the core short sentences based on the standard time in the core short sentences to generate event sequence data.

[0133] Specifically, after obtaining the core short sentences, corresponding timestamps are generated according to their standard time. For cases including fuzzy times such as "morning" and "dawn", they are unified to specific times according to rules and then timestamps are generated. For example, "morning" corresponds to 9 o'clock and "dawn" corresponds to 2 o'clock, etc., which can be specifically set according to actual accuracy requirements. After generating the timestamps of the core short sentences, all the clustered core short sentences are sorted by the timestamps to generate event sequence data in the format of a series of "timestamp + core short sentence".

[0134] Step S2052: Filter the core short sentences in the event sequence data based on logical expressions to obtain target core short sentences, and generate an event context based on the target core short sentences.

[0135] Specifically, by using a large language model and combining the logical expressions used when collecting news data to process the event sequence data, the final event context can be obtained.

[0136] In some alternative embodiments, the event sequence data is filtered according to the logical expressions used when collecting news data. The large language model will first understand the event subject information such as the characters, locations, times, and actions mentioned in the logical expressions, and then compare them with the core short sentences in the current event sequence data. When there is no association between the event subject information mentioned in the core short sentences and the subject information mentioned in the logical expressions, the irrelevant core short sentences will be filtered to obtain the final target core short sentences.

[0137] In this embodiment, the reason for processing the event sequence data using logical expressions is to provide a reference for the large language model to analyze the event subject, find the associations between events from a set of events, identify which events are associated events and which are unassociated events. After filtering out the unassociated events, different events at the same time are merged. Finally, relying on the powerful text generation ability of the large language model, based on the target core short sentence, an event development context data including the time and a summary description of the events is generated, thus completing the entire analysis process and outputting it for the customer to use.

[0138] The present invention first filters the core short sentences using the logical expressions used in news data collection, thereby filtering out the unassociated data in the event sequence data and improving the accuracy of the event context data. Then, the large language model is used to merge and summarize the associated target core short sentences to generate the event context and improve the readability of the event context data.

[0139] The event context generation method provided in this embodiment optimizes the event context generation process through the large language model. By analyzing the theme description information through the large language model, it helps the customer generate collection keywords, greatly simplifying the difficulty of customer operation. Moreover, by using the large language model for context reasoning, various complex non-standard format times can be inferred, improving the data recall rate and accuracy and avoiding missing event context information. Through clustering of short sentence vectors, the ability to merge similar events can be improved at the semantic level. By using the reasoning ability of the large language model, the main line of event development can be found more accurately, and events unrelated to the main line can be filtered out. Finally, by using the powerful text output and summarization ability of the large language model, the readability of the event context is improved.

[0140] The following combines a specific application example to detail the event context generation method of the present invention, as Figure 3 shown. This application example includes the following steps from the start of setting, processing a concerned event to outputting the final event context:

[0141] Step 1, the customer inputs the theme description information and uses the large language model to generate a keyword list.

[0142] Exemplarily, the prompt for the input large language model to analyze the topic description information and generate a keyword list can be as follows. The analysis objective is: You are a senior public opinion analysis expert. Please carefully analyze the following description information of an event and generate a keyword list to search for relevant news on the Internet that matches the keyword list. The expression generation rule is: Please execute step by step according to the following rules; tokenize the expression, remove meaningless words such as modal particles and auxiliary words, and extract keywords such as time, location, subject, and action; to match more content, try to split into shorter keywords as much as possible; for each keyword, find up to 5 synonyms, enclose the synonyms in [], the first one is the original keyword, and the following ones are the recommended synonyms, and if there are none, no recommendation is made.

[0143] Exemplarily, the customer inputs the topic description information "Rainwater accumulation causes traffic accidents" about an event. The system splits the keywords according to the description information and provides multiple recommended keywords. The generated keyword list is as follows: [rain, precipitation], [water accumulation, water storage], [traffic accident, traffic accident].

[0144] Step 2, collect news data based on the keyword list.

[0145] Exemplarily, generate a logical expression according to the above keyword list: (rain or precipitation) and (water accumulation or water storage) and (traffic accident or traffic accident). This logical expression is used for the system to collect matching news data, and the logical expression is not directly presented to the customer to reduce the customer's usage difficulty.

[0146] Step 3, split the news, extract candidate short sentences containing time feature characters based on the time feature dictionary, and screen out feature short sentences containing time feature texts from the candidate short sentences. Among them, the time feature text includes standard time text and non-standard time text.

[0147] Step 4, use regular expressions to extract the standard time text in the feature short sentence and convert it into the standard time.

[0148] Step 5, use the large language model to extract the non-standard time text in the feature short sentence and convert it into the standard time.

[0149] Exemplarily, the prompt for using a large language model to determine the standard time corresponding to non-standard time text is as follows. The extraction objective is: You are a senior public opinion analysis expert. Please carefully analyze the following news information, and based on the context of the information, infer the specific time corresponding to the non-standard time text and generate the standard time. Precautions for inputting into the large language model: The converted standard time format is xxxx year xx month xx day xx hour xx minute xx second or xxxx year xx month xx day (morning|afternoon|evening, etc.); the output time adopts the 24-hour system and only shows the minimum precision that the original time text itself has; if it can be accurate to the hour level, do not display ambiguous times such as morning and afternoon.

[0150] For example, the text is "It started to rain on September 1, 2024, causing traffic congestion in the city. 5 traffic accidents occurred in the afternoon of the same day. The rainfall continued until 2 pm on the next day." Extraction result: afternoon of the same day -> afternoon of September 1, 2024, 2 pm on the next day -> 14:00 on September 2, 2024.

[0151] Step 6, use the embedding vectors of the large language model to merge similar events.

[0152] Step 7, use the large language model to filter non-related events and output the development context of the events.

[0153] Exemplarily, the prompt for the entire analysis process using the large language model is as follows. The analysis objective is: You are a senior public opinion analysis expert. Please, according to the following logical expression given, mine the development context information of the events from the event sequence data. Analysis rules: Carefully compare the correlation between each event and the logical expression, and filter out information unrelated to the event; carefully judge the correlation between information, and merge duplicate information for the same event; merge different events at the same time and describe them uniformly. Output: Describe the entire process of event development in chronological order. Please describe it from an objective perspective, without subjective evaluation and explanatory content, and without expanding non-existent information.

[0154] For example, the logical expression is: (rain or precipitation) and (ponding or water storage) and (traffic accident or traffic accident). Event sequence data: At 9 am on September 1, 2024, it started to rain and there was ponding on many roads in the city; At 9 am on September 1, 2024, the ponding caused traffic congestion in the city; 5 traffic accidents occurred in the afternoon of September 1, 2024; As of 5 pm on September 1, 2024, a total of 5 traffic accidents occurred in the city; At 9 pm on September 1, 2024, a resident in Xiangyang Community called the police saying that his home had been burglarized; The rainfall continued until 2 pm on September 2, 2024.

[0155] Further, output the event context: At 9:00 on September 1, 2024, it started to rain in the city, and there was waterlogging on many roads in the city, causing traffic congestion. At 17:00 on September 1, 2024, as of 17:00 in the afternoon, a total of 5 traffic accidents occurred in the whole city. At 14:00 on September 2, 2024, the rainfall continued until 14:00 on the 2nd and ended.

[0156] At the same time, the "At 21:00 on September 1, 2024, a resident in Xiangyang Community reported that his home was burglarized" in the event sequence data was filtered in the final result because it belongs to an uncorrelated event.

[0157] As Figure 4 shown, this application example provides an event context generation system, which includes:

[0158] A keyword generation module, which is responsible for generating a keyword list and a logical expression for collecting relevant data according to the topic description information input by the user.

[0159] A news collection module, which triggers the collection system to collect news data according to the logical expression and saves the collected news data into the system.

[0160] A time text extraction module: extracts candidate short sentences containing time feature characters from the collected news, extracts the feature short sentences containing time feature text from all candidate short sentences, and converts the time feature text into a standard time through regular expressions and a large language model.

[0161] An event aggregation module, which merges short sentences describing the same event together to generate event sequence data to be analyzed.

[0162] An event development context extraction module, which analyzes the core events according to the logical expression and the event sequence data, filters out irrelevant events, generates a complete event development context, and outputs the final result to the customer.

[0163] The above application example provides a complete set of prompting words for analyzing and processing data in each link, and includes a complete process from creating a topic, extracting time, event deduplication, filtering uncorrelated events, to outputting the event development context, which can improve the final presentation effect of the event development context, improve the accuracy and readability of the event context, and avoid people misinterpreting the process of event development.

[0164] In this embodiment, an event context generation device is also provided. This device is used to implement the above embodiment and the preferred implementation manner, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0165] This embodiment provides an event context generation device, as Figure 5 shown, including:

[0166] A first processing module 501, configured to determine a keyword list based on the topic description information input by the customer, generate a logical expression based on the keyword list, and collect news data that matches the logical expression;

[0167] A second processing module 502, configured to extract candidate short sentences containing time feature characters from the news data, and screen out feature short sentences containing time feature text; wherein, the time feature text is text arranged according to a preset time format by the time feature characters;

[0168] A third processing module 503, configured to standardize the time feature text in the feature short sentences to obtain target short sentences containing only standard times, and group the target short sentences according to the standard times in the target short sentences;

[0169] A fourth processing module 504, configured to cluster the target short sentences in each group, and determine the core short sentences in the target short sentences of each cluster, wherein each group contains at least one cluster;

[0170] A fifth processing module 505, configured to arrange the core short sentences based on the standard times in the core short sentences to generate event sequence data, and generate an event context based on the event sequence data and the logical expression.

[0171] In some alternative embodiments, the first processing module 501 is further configured to:

[0172] Use a trained large language model to extract key information from the topic description information and generate synonymous associated information of the key information;

[0173] In response to the user's selection operation on the key information and the synonymous associated information, determine at least one keyword selected by the user and generate a keyword list.

[0174] In some alternative embodiments, the second processing module 502 is further configured to:

[0175] Use a delimiter to split the news data to obtain news short sentences; wherein, the delimiter is a terminating punctuation mark used to indicate the end of a sentence;

[0176] Detect news short sentences containing time feature characters to obtain candidate short sentences.

[0177] In some alternative embodiments, the preset time format includes a standard time format and a non-standard time format. Among them, the standard time format includes at least some of the following items: year, month, day, hour, minute, second, and time nouns used to represent time periods within a day; the time feature text includes standard time text in which time feature characters are arranged according to the standard time format and non-standard time text in which time feature characters are arranged according to the non-standard time format; the third processing module 503 is further configured to:

[0178] If it is detected that the time feature text in the feature short sentence is standard time text, use a regular expression to extract the standard time text in the feature short sentence, and replace the standard time text in the feature short sentence with the corresponding standard time, so as to obtain a target short sentence containing only standard time;

[0179] If it is detected that the time feature text in the feature short sentence is non-standard time text, use the trained large language model to replace the non-standard time text in the feature short sentence with the corresponding standard time, so as to obtain a target short sentence containing only standard time.

[0180] In some alternative embodiments, the third processing module 503 is further configured to:

[0181] Determine the context-related feature short sentences of the feature short sentence containing non-standard time text in the news data, and the context-related feature short sentences are feature short sentences containing standard time text;

[0182] Use the trained large language model to predict the standard time corresponding to the non-standard time text in the feature short sentence containing non-standard time text based on the standard time text in the context-related feature short sentence;

[0183] Replace the non-standard time text in the feature short sentence according to the standard time corresponding to the non-standard time text, so as to obtain a target short sentence containing only standard time.

[0184] In some alternative embodiments, the fourth processing module 504 is further configured to:

[0185] Use the trained large language model to generate a short sentence vector of the target short sentence, cluster the target short sentences in each group based on the short sentence vector, and determine the target short sentences included in each cluster in each group;

[0186] Calculate the central vector of each cluster according to the arithmetic mean of the short sentence vectors of all the target short sentences in each cluster;

[0187] Obtain the core short sentence of each cluster according to the target short sentence with the shortest cosine distance from the central vector in each cluster.

[0188] In some alternative embodiments, the fifth processing module 505 is further configured to:

[0189] Filter the core short sentences in the event sequence data based on a logical expression to obtain target core short sentences;

[0190] Generate an event context based on the target core short sentences.

[0191] The further function descriptions of the above-mentioned various modules and units are the same as those in the corresponding embodiments above, and will not be elaborated here.

[0192] The event context generation device in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0193] This embodiment of the present invention further provides a computer device having the above-mentioned Figure 5 shown event context generation device.

[0194] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of a computer device provided by an alternative embodiment of the present invention. As shown in Figure 6 , the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a set of blade servers, or a multi-processor system). Figure 6 In

[0195] FIG. 24, a single processor 10 is taken as an example.

[0196] Among them, the memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiments.

[0197] The memory 20 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may optionally include a memory remotely provided with respect to the processor 10, and these remote memories may be connected to the computer device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0198] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk, or a solid-state drive; the memory 20 may also include a combination of the above types of memories.

[0199] The computer device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30, and the output device 40 may be connected through a bus or other means. Figure 6 Taking connection through a bus as an example.

[0200] The input device 30 may receive input digital or character information, and generate key signal inputs related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., an LED), and a haptic feedback device (e.g., a vibration motor), etc. The above display device includes but is not limited to a liquid crystal display, a light-emitting diode, a display, and a plasma display. In some alternative embodiments, the display device may be a touch screen.

[0201] Embodiments of the present invention also provide a computer-readable storage medium. The method according to the embodiments of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code that is originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiments is implemented.

[0202] A part of the present invention can be applied as a computer program product, such as computer program instructions. When executed by a computer, through the operation of the computer, the method and / or technical solution according to the present invention can be called or provided. Those skilled in the art should be able to understand that the forms of existence of computer program instructions in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to the computer.

[0203] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A method for generating an event context, characterized in that: The method comprises: Determine a keyword list based on the subject description information input by the customer, generate a logical expression based on the keyword list, and collect news data matching the logical expression; Extracting short sentences containing time characteristic characters from news data, and selecting characteristic short sentences containing time characteristic text from the short sentences; wherein the time characteristic text is a text in which the time characteristic characters are arranged according to a preset time format; Standardizing the time feature text in the feature short sentences to obtain target short sentences containing only standard time, and grouping the target short sentences according to the standard time in the target short sentences; Clustering the target short sentences in each group, and determining the core short sentences in the target short sentences of each cluster, wherein each group contains at least one cluster; The core sentences are arranged based on the standard time in the core sentences to generate event sequence data, and an event context is generated based on the event sequence data and the logical expression.

2. The method according to claim 1, characterized in that The preset time format includes a standard time format and a non-standard time format, wherein the standard time format includes at least part of the following items: year, month, day, hour, minute, second, and a time noun used to characterize a time period in a day; the time feature text includes a standard time text in which time feature characters are arranged according to the standard time format and a non-standard time text in which time feature characters are arranged according to the non-standard time format; The step of standardizing the time feature text in the feature short sentence to obtain a target short sentence containing only the standard time includes: If it is detected that the time feature text in the feature short sentence is the standard time text, the standard time text in the feature short sentence is extracted using a regular expression, and the standard time text in the feature short sentence is replaced with the standard time corresponding to the standard time text, so as to obtain a target short sentence containing only the standard time; If it is detected that the time feature text in the feature sentence is non-standard time text, the non-standard time text in the feature sentence is replaced with the standard time corresponding to the non-standard time text using the trained large language model to obtain a target sentence containing only the standard time.

3. The method according to claim 2, characterized in that The method of using the trained large language model to replace the non-standard time text in the feature short sentence with the standard time corresponding to the non-standard time text to obtain a target short sentence containing only the standard time includes: Determine a context-related feature short sentence containing a feature short sentence of non-standard time text in the news data, wherein the context-related feature short sentence is a feature short sentence containing a standard time text; Using the trained large language model, based on the standard time text in the context-related feature sentences, predict the standard time corresponding to the non-standard time text in the feature sentences containing non-standard time text; The non-standard time text in the characteristic short sentence is replaced according to the standard time corresponding to the non-standard time text, so as to obtain a target short sentence containing only the standard time.

4. The method according to claim 1, characterized in that: The step of clustering the target sentences in each group and determining the core sentences in the target sentences of each cluster includes: Generate a short sentence vector of the target short sentence using the trained large language model, cluster the target short sentences in each group based on the short sentence vector, and determine the target short sentences contained in each cluster in each group; The center vector of each cluster is calculated based on the arithmetic mean of the short sentence vectors of all target short sentences in each cluster; According to the target short sentence with the shortest cosine distance from the center vector in each cluster, the core short sentence of each cluster is obtained.

5. The method according to any one of claims 1 to 4, characterized in that The step of extracting short sentences containing time characteristic characters from news data includes: The news data is split using separators to obtain news short sentences; wherein the separator is a terminal punctuation mark used to indicate the end of a sentence; Detect news sentences containing time feature characters to obtain short sentences to be selected.

6. The method according to any one of claims 1 to 4, characterized in that The generating an event context based on the event sequence data and the logical expression comprises: Filtering the core short sentences in the event sequence data based on the logical expression to obtain target core short sentences; Generate event context based on the target core sentence.

7. The method according to any one of claims 1 to 4, characterized in that The step of determining a keyword list based on the subject description information input by the customer includes: Extract key information from the topic description information using a trained large language model, and generate synonymous related information of the key information; In response to the user's selection operation on the key information and the synonymous associated information, at least one keyword selected by the user is determined and a keyword list is generated.

8. An event context generating device, characterized in that: The device comprises: A first processing module, for determining a keyword list based on the subject description information input by the customer, generating a logical expression based on the keyword list, and collecting news data matching the logical expression; The second processing module is used to extract short sentences containing time characteristic characters from the news data, and select characteristic short sentences containing time characteristic text from the short sentences; wherein the time characteristic text is a text in which the time characteristic characters are arranged according to a preset time format; The third processing module is used to standardize the time feature text in the feature short sentences to obtain target short sentences containing only standard time, and group the target short sentences according to the standard time in the target short sentences; A fourth processing module is used to cluster the target short sentences in each group and determine the core short sentences in each cluster of the target short sentences, wherein each group contains at least one cluster; The fifth processing module is used to arrange the core sentences based on the standard time in the core sentences to generate event sequence data, and to generate event context based on the event sequence data and the logical expression.

9. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the event context generation method according to any one of claims 1 to 7 by executing the computer instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the event context generation method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Event context reduction method and system based on vector retrieval and large language model

    CN120780759A