Information processing method and device based on GPT model, storage medium and electronic device

Through the information processing method based on the GPT model, event topic profile information is automatically generated, which solves the problem of wasting time when users find event topic profile information in a large amount of text information, and achieves the effect of quickly understanding the development trend of event topics.

CN120067312APending Publication Date: 2025-05-30QINGDAO HAIER TECH +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311584777.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-24
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

It takes a lot of time for users to find event topic profile information in a large amount of text information, and it is difficult for the prior art to automatically generate effective event topic profile information.

Method used

The information processing method based on the GPT model is adopted, and multiple text documents in the current time period are obtained, trend hot words are extracted, trend hot words co-occurrence matrix is ​​constructed, trend hot words map is generated, clustered, and the clustered trend hot words are extracted through the second GPT model to generate event topic overview information.

Benefits of technology

It realizes the rapid and automatic generation of event topic overview information, helping users quickly understand the development trends of event topics, and improving user experience and satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067312A_ABST
    Figure CN120067312A_ABST
Patent Text Reader

Abstract

The invention discloses an information processing method and device based on a GPT model, a storage medium and an electronic device, and relates to the technical field of smart home / smart home, and the method comprises the steps: obtaining a plurality of text documents in a current time period, and carrying out the extraction processing of trend hot words in the text documents based on a preset first GPT model; based on the multiple trend hot words, constructing a co-occurrence matrix of the trend hot words in the current time period, and obtaining matrix elements of the co-occurrence matrix; constructing a trend hot word graph of trend hot words in the current time period based on the matrix elements; and clustering the trend hot words based on a trend hot word graph to obtain clustered trend hot words, and performing abstract extraction on the target text document corresponding to the clustered trend hot words through a preset second GPT model to obtain event topic general situation information in the current time period. The event topic general situation information is automatically generated, and a user can conveniently and quickly know the development trend of the event topic.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of smart home, and particularly to an information processing method, device, storage medium and electronic device based on the GPT model. Background Art

[0002] With the development of network technology, text information about various events floods the network media every day. If a user wants to understand the general information of the event topic of a certain event, they need to read a large amount of text information among numerous texts to achieve this, which will take a lot of time for the user.

[0003] Therefore, currently, finding a method that can automatically generate the general information of the event topic representing a certain event has become a research hotspot. Summary of the Invention

[0004] This application provides an information processing method, device, storage medium and electronic device based on the GPT model, which realizes the ability to automatically generate the general information of the event topic, thereby facilitating the user to quickly understand the development trend of the event topic and improving the user experience and satisfaction.

[0005] This application provides an information processing method based on the GPT model. The method includes: obtaining a plurality of text documents within the current time period, and extracting trend hot words in the text documents based on a preset first GPT model, where the trend hot words represent words whose occurrence frequency meets the set requirements within a preset statistical time; constructing a co-occurrence matrix of the trend hot words within the current time period based on the plurality of trend hot words, and obtaining matrix elements of the co-occurrence matrix, where the matrix elements represent the co-occurrence frequency of different trend hot words in the same text document; constructing a trend hot word graph of the trend hot words within the current time period based on the matrix elements, where the trend hot word graph includes graph nodes and edges, the graph nodes in the trend hot word graph are the trend hot words within the current time period, and the edges connecting the graph nodes are the matrix elements corresponding to the trend hot words within the current time period; clustering the trend hot words based on the trend hot word graph to obtain the clustered trend hot words, and extracting a summary of the target text documents corresponding to the clustered trend hot words through a preset second GPT model to obtain the general information of the event topic within the current time period.

[0006] According to the information processing method based on the GPT model provided by this application, constructing a co-occurrence matrix of the trend hot words within the current time period based on the plurality of trend hot words specifically includes: determining the occurrence frequency of each trend hot word and other trend hot words in the same text; using the frequency as matrix elements to construct a co-occurrence matrix of the trend hot words within the current time period.

[0007] According to the information processing method based on the GPT model provided by the present application, clustering the trending hot words based on the trending hot word graph to obtain the clustered trending hot words specifically includes: based on the trending hot word graph, clustering the graph nodes in the trending hot word graph through a community clustering algorithm to obtain the clustered graph nodes; clustering the trending hot words corresponding to the clustered graph nodes to obtain the clustered trending hot words.

[0008] According to the information processing method based on the GPT model provided by the present application, before extracting an abstract of the target text document corresponding to the clustered trending hot words through a preset second GPT model to obtain the event topic overview information in the current time period, the method further includes: based on the clustered trending hot words, screening out the target text document from multiple text documents, where the clustered trending hot words are extracted from the target text document; determining a preset second text prompt, where the second text prompt is used to guide the second GPT model to output the prompt information for the event topic overview information; the extracting an abstract of the target text document corresponding to the clustered trending hot words through a preset second GPT model to obtain the event topic overview information in the current time period specifically includes: inputting the target text document and the second text prompt into the second GPT model, so that the second GPT model extracts an abstract of the target text document to obtain the event topic overview information in the current time period.

[0009] According to the information processing method based on the GPT model provided by the present application, before extracting the trending hot words in the text document through a preset first GPT model, the method further includes: determining a preset first text prompt, where the first text prompt is used to guide the first GPT model to output the prompt information for the output result in a preset format, where the output result in the preset format includes the trending hot words of the text document and the document identifier of the text document; the extracting the trending hot words in the text document through a preset first GPT model specifically includes: inputting the text document and the first text prompt into the first GPT model, so that the first GPT model extracts the trending hot words of the text document and the document identifier of the text document; the screening out the target text document from multiple text documents based on the clustered trending hot words specifically includes: based on the clustered trending hot words and the document identifier of the text document corresponding to the trending hot words, determining the document identifier of the target text document corresponding to the clustered trending hot words; screening out the target text document from the multiple text documents based on the document identifier of the target text document.

[0010] According to the information processing method based on the GPT model provided by the present application, after obtaining the event topic overview information in the current time period, the method further includes: obtaining a plurality of text documents in a plurality of subsequent time periods; and obtaining the event topic overview information in the current time period according to the steps of combining a plurality of text documents in the current time period as described above, obtaining a plurality of event topic overview information in a plurality of subsequent time periods based on the plurality of text documents in the plurality of subsequent time periods; and forming an event chain representing the development trend of the event topic based on the event topic overview information in the current time period and the plurality of event topic overview information in the plurality of subsequent time periods.

[0011] According to the information processing method based on the GPT model provided by the present application, forming an event chain representing the development trend of the event topic based on the event topic overview information in the current time period and the plurality of event topic overview information in the plurality of subsequent time periods specifically includes: classifying the event topics in the current time period and the event topics in the subsequent time periods having an event association relationship through bipartite graph matching based on the event topic overview information in the current time period and the plurality of event topic overview information in the plurality of subsequent time periods, to obtain the classified event topics; and sorting the event topic overview information corresponding to the classified event topics in chronological order to form an event chain representing the development trend of the event topic.

[0012] The present application also provides an information processing device based on the GPT model. The device includes: an acquisition module, configured to acquire a plurality of text documents in the current time period, and perform extraction processing on the trend hot words in the text documents based on a preset first GPT model, where the trend hot words represent words whose occurrence frequencies meet the set requirements within a preset statistical time; a matrix construction module, configured to construct a co-occurrence matrix of the trend hot words in the current time period based on the plurality of trend hot words, and obtain the matrix elements of the co-occurrence matrix, where the matrix elements represent the co-occurrence frequencies of different trend hot words in the same text document; a graph construction module, configured to construct a trend hot word graph of the trend hot words in the current time period based on the matrix elements, where the trend hot word graph includes graph nodes and edges, the graph nodes in the trend hot word graph are the trend hot words in the current time period, and the edges connecting the graph nodes are the matrix elements corresponding to the trend hot words in the current time period; and an extraction module, configured to perform clustering on the trend hot words based on the trend hot word graph to obtain the clustered trend hot words, and perform abstract extraction on the target text documents corresponding to the clustered trend hot words through a preset second GPT model to obtain the event topic overview information in the current time period.

[0013] The present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to execute and implement any one of the above-mentioned GPT model-based information processing methods through the computer program.

[0014] The present application further provides a computer-readable storage medium. The computer-readable storage medium includes a stored program, wherein the program, when running, executes and implements any one of the above-mentioned GPT model-based information processing methods.

[0015] The present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements any one of the above-mentioned GPT model-based information processing methods.

[0016] An information processing method, device, storage medium and electronic device based on the GPT model provided by the present application. Multiple text documents within the current time period are obtained, and trend hot words in the text documents are extracted based on a preset first GPT model. Based on the multiple trend hot words, a co-occurrence matrix of the trend hot words within the current time period is constructed, and matrix elements of the co-occurrence matrix are obtained. A trend hot word graph of the trend hot words within the current time period is constructed based on the matrix elements. The trend hot words are clustered based on the trend hot word graph to obtain the clustered trend hot words, and abstract extraction is performed on the target text documents corresponding to the clustered trend hot words through a preset second GPT model to obtain the event topic overview information within the current time period. It realizes the ability to automatically generate event topic overview information, thereby facilitating users to quickly understand the development trend of event topics and improving the user experience and satisfaction. Description of the Drawings

[0017] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 It is a schematic diagram of the hardware environment of an information processing method based on the GPT model according to an embodiment of the present application;

[0020] Figure 2 It is one of the flow diagrams of the information processing method based on the GPT model provided by the present application;

[0021] Figure 3It is a schematic flowchart of the process provided by this application for extracting a summary of the target text documents corresponding to the clustered trending hot words through a pre-set second GPT model to obtain the event topic overview information within the current time period;

[0022] Figure 4 It is the second schematic flowchart of the information processing method based on the GPT model provided by this application;

[0023] Figure 5 It is a schematic structural diagram of the information processing device based on the GPT model provided by this application;

[0024] Figure 6 It is a schematic structural diagram of the electronic device provided by this application. Detailed implementation manners

[0025] In order to enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0027] According to one aspect of the embodiments of this application, an information processing method based on the GPT model is provided. The information processing method based on the GPT model is widely applied to whole-house intelligent digital control application scenarios such as Smart Home, smart home, smart home device ecosystem, and Intelligence House ecosystem. Optionally, in this embodiment, the above-mentioned information processing method based on the GPT model can be applied to, for example Figure 1 the hardware environment composed of the terminal device 102 and the server 104 as shown. As Figure 1As shown, the server 104 is connected to the terminal device 102 via a network and can be used to provide services (such as application services, etc.) for the terminal or the client installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for the server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data operation services for the server 104.

[0028] The above network can include but is not limited to at least one of the following: wired network, wireless network. The above wired network can include but is not limited to at least one of the following: wide area network, metropolitan area network, local area network. The above wireless network can include but is not limited to at least one of the following: WIFI (Wireless Fidelity), Bluetooth. The terminal device 102 is not limited to being a PC, mobile phone, tablet computer, smart air conditioner, smart range hood, smart refrigerator, smart oven, smart stove, smart washing machine, smart water heater, smart washing equipment, smart dishwasher, smart projection device, smart TV, smart clothes hanger, smart curtain, smart audio and video, smart socket, smart speaker, smart sound box, smart fresh air device, smart kitchen and bathroom equipment, smart bathroom equipment, smart floor sweeping robot, smart window cleaning robot, smart mopping robot, smart air purification device, smart steam box, smart microwave oven, smart kitchen water heater, smart purifier, smart water dispenser, smart door lock, etc.

[0029] In another embodiment, the information processing method based on the GPT model provided by the present application can be applied to smart home appliances. Among them, smart home appliances refer to home appliance products formed after introducing microprocessor, sensor technology, and network communication technology into home appliance devices, which have the ability to automatically sense the state of the residential space, the state of the home appliances themselves, and the service state of the home appliances, and can automatically control and receive control instructions from residential users inside or remotely from the residence. It can be understood that smart home appliances are components of smart homes.

[0030] Figure 2 It is a schematic flowchart of the information processing method based on the GPT model provided by the present application.

[0031] Next, in combination with Figure 2 the process of the information processing method based on the GPT model will be described.

[0032] In an exemplary embodiment of the present application, in combination with Figure 2 it can be seen that the information processing method based on the GPT model can include steps 210 to 240, and each step will be introduced separately below.

[0033] In step 210, multiple text documents within the current time period are obtained, and trend hot words in the text documents are extracted based on a preset first GPT model.

[0034] In one embodiment, the trending hot words in the text document can be extracted based on a pre-set first GPT model, so as to obtain the trending hot words. Among them, the trending hot words refer to the words whose occurrence frequencies meet the set requirements within the preset statistical time. The preset requirements can be adjusted according to the actual situation. In one example, the trending hot words can be the words whose mentioned frequencies are greater than the frequency threshold within the preset statistical time. The trending hot words can be regarded as the hot topics or hot events within the preset time period.

[0035] In one embodiment, the source of the text document can be the news text of a news media website, and the source of the text document can also be the posts on social media. In one instance, the text document can be collected within the current time period. Further, based on the text document within the current time period, the trending hot words within the current time period can be obtained. Among them, the time period can be 1 day. In this embodiment, the time period is not specifically limited and can be adjusted according to the actual situation.

[0036] In step 220, based on multiple trending hot words, a co-occurrence matrix of the trending hot words within the current time period is constructed, and the matrix elements of the co-occurrence matrix are obtained, where the matrix elements represent the co-occurrence frequencies of different trending hot words in the same text document.

[0037] In another exemplary embodiment of the present application, continuing with the embodiment described above, based on multiple trending hot words, constructing a co-occurrence matrix of the trending hot words within the current time period (which can correspond to step 220) can be implemented in the following manner:

[0038] Determine the occurrence frequencies of each trending hot word and other trending hot words in the same text;

[0039] Using the frequencies as matrix elements, construct a co-occurrence matrix of the trending hot words within the current time period.

[0040] It should be noted that the co-occurrence matrix can describe the situation where different trending hot words appear in the same text segment. The elements of the co-occurrence matrix can represent the co-occurrence frequencies of the two trending hot words corresponding to the row and column in the same text document. In one example, if the element of the co-occurrence matrix is 0, it can indicate that the trending hot word represented by the row corresponding to the element and the trending hot word represented by the column corresponding to the element have not appeared in the same text document.

[0041] In one embodiment, a co-occurrence matrix of the trending hot words within the current time period can be constructed based on multiple trending hot words, and the matrix elements of the co-occurrence matrix are obtained. Among them, the matrix elements in the obtained co-occurrence matrix lay the foundation for clustering the trending hot words.

[0042] In step 230, based on the matrix elements, a trend hot word graph of the trend hot words in the current time period is constructed. The trend hot word graph includes graph nodes and edges. The graph nodes in the trend hot word graph are the trend hot words in the current time period, and the edges connecting the graph nodes are the matrix elements corresponding to the trend hot words in the current time period.

[0043] In step 240, the trend hot words are clustered based on the trend hot word graph to obtain the clustered trend hot words, and the target text documents corresponding to the clustered trend hot words are summarized by a preset second GPT model to obtain the event topic overview information in the current time period.

[0044] In one embodiment, a trend hot word graph regarding the trend hot words in the current time period can be constructed based on the matrix elements of the co-occurrence matrix. The trend hot word graph includes graph nodes and edges. The graph nodes in the trend hot word graph are the trend hot words in the current time period, and the edges connecting the graph nodes are the elements corresponding to the trend hot words in the current time period. In other words, the graph nodes in the trend hot word graph are the trend hot words in each current time period. If the elements corresponding to two trend hot words in the current time period in the co-occurrence matrix are not 0, then there is a relationship edge between these two trend hot words in the current time period, and the magnitude of the matrix element value represents the relationship strength, the larger the data, the stronger the relationship.

[0045] It should be noted that constructing the trend hot word graph can lay a foundation for clustering the graph nodes in the trend hot word graph of the current time period based on the community clustering algorithm to obtain the clustered graph nodes. It can be understood that the clustered graph nodes correspond to the clustered trend hot words.

[0046] In another embodiment, after obtaining the clustered trend hot words, the target text documents corresponding to the clustered trend hot words can also be summarized by a preset second GPT model to obtain the event topic overview information in the current time period. The clustered trend hot words are extracted from the target text documents.

[0047] In another example, the event topics corresponding to the clustered trend hot words can also be obtained based on the clustered trend hot words. In one example, the set composed of the clustered trend hot words can be used as the event topic.

[0048] An information processing method based on the GPT model provided by this application obtains multiple text documents within the current time period, extracts trend hot words in the text documents based on a preset first GPT model, constructs a co-occurrence matrix of the trend hot words within the current time period based on the multiple trend hot words, and obtains the matrix elements of the co-occurrence matrix; constructs a trend hot word graph of the trend hot words within the current time period based on the matrix elements; clusters the trend hot words based on the trend hot word graph to obtain the clustered trend hot words, and extracts the abstracts of the target text documents corresponding to the clustered trend hot words through a preset second GPT model to obtain the event topic overview information within the current time period. It realizes the ability to automatically generate event topic overview information, thereby facilitating users to quickly understand the development trend of event topics and improving the user experience and satisfaction.

[0049] In another exemplary embodiment of this application, continuing with the foregoing embodiment as an example for illustration, clustering the trend hot words based on the trend hot word graph to obtain the clustered trend hot words (which may correspond to step 240) can be implemented in the following manner:

[0050] Based on the trend hot word graph, cluster the graph nodes in the trend hot word graph through a community clustering algorithm to obtain the clustered graph nodes;

[0051] Cluster the trend hot words corresponding to the clustered graph nodes to obtain the clustered trend hot words.

[0052] The community clustering algorithm can be the louvain algorithm, and the community clustering algorithm is not specifically limited in this embodiment.

[0053] In another embodiment, based on the trend hot word graph, the graph nodes of the trend hot word graph within the current time period can be clustered through a community clustering algorithm. Among them, the graph nodes with higher co-occurrence frequencies also tend to be clustered in one cluster, which can be understood as the clustered graph nodes. Among them, each formed clustered graph node can correspond to an event topic.

[0054] Further, clustering processing based on the trend hot words within the current time period corresponding to the clustered graph nodes can obtain the clustered trend hot words within the current time period. It can be understood that clustering the trend hot words within multiple next time periods based on the elements of the co-occurrence matrix to obtain the clustered trend hot words within the next time period can also be implemented in the foregoing manner.

[0055] In another embodiment, after constructing the clustered graph nodes, it can also have a positive impact on the process of classifying event topics with event association relationships.

[0056] Figure 3It is a schematic flowchart of the process of extracting a summary of the target text document corresponding to the clustered trend hot words through a pre-set second GPT model to obtain the event topic overview information within the current time period provided by this application.

[0057] The following will be combined with Figure 3 to illustrate the process of the information processing method based on the GPT model of this application.

[0058] In an exemplary embodiment of this application, combined with Figure 3 it can be known that extracting a summary of the target text document corresponding to the clustered trend hot words through a pre-set second GPT model to obtain the event topic overview information within the current time period may include steps 310 to 330. The following will introduce each step separately.

[0059] In step 310, based on the clustered trend hot words, target text documents are screened out from multiple text documents, where the clustered trend hot words are extracted from the target text documents.

[0060] In step 320, a pre-set second text prompt is determined, where the second text prompt is used to guide the second GPT model to output the prompt information of the event topic overview information.

[0061] In step 330, the target text document and the second text prompt are input into the second GPT model, so that the second GPT model extracts a summary of the target text document to obtain the event topic overview information within the current time period.

[0062] In one embodiment, multiple text documents corresponding to the clustered trend hot words within the current time period can be recalled according to the clustered trend hot words within the current time period. In another example, when the number of recalled text documents is large, text documents with a higher word density containing the clustered trend hot words can be selected as the target text documents, so as to effectively improve the quality of the target text documents and will not affect the quality of the finally obtained event overview information.

[0063] In addition, screening out the target text documents is also to meet the input requirements of the second GPT model, thus laying a foundation for quickly and accurately outputting the event topic overview information based on the second GPT model.

[0064] In another embodiment, the second text prompt can be a prompt for outputting the event overview information corresponding to the event topic. In other words, the second text prompt is the prompt information for guiding the second GPT model to output the event topic overview information.

[0065] During the application process, the target text document and the second text prompt can be input into the second GPT model, so that the event topic overview information corresponding to the target text document within the current time period output by the second GPT model can be obtained quickly and accurately.

[0066] In another exemplary embodiment of the present application, taking the foregoing embodiment as an example, before extracting the trend hot words in the text document based on the preset first GPT model, the information processing method based on the GPT model may further include:

[0067] Determine a preset first text prompt, where the first text prompt is used to guide the first GPT model to output prompt information in a preset format, and the preset format output result includes the trend hot words of the text document and the document identifier of the text document;

[0068] Furthermore, extracting the trend hot words in the text document based on the preset first GPT model can be implemented in the following manner:

[0069] Input the text document and the first text prompt into the first GPT model, so that the first GPT model extracts the text document to obtain the trend hot words of the text document and the document identifier of the text document.

[0070] In one embodiment, in order to quickly extract the trend hot words in the text document, it can also be implemented based on the GPT language model. In one example, since the document text input by the GPT language model has a character limit, during the application process, the text document can be cleaned and segmented to meet the requirements of the GPT language model (corresponding to the first GPT model) for the input characters.

[0071] In one embodiment, the first text prompt can be a prompt. Since the prompt restricts the output result to include the trend hot words and the document identifier of the text document corresponding to the trend hot words, further, when inputting the text document within the current time period and the first text prompt (prompt) into the first GPT model, the trend hot words of the text document and the document identifier of the text document output by the first GPT model can be obtained. Since the text document has a document identifier, it can lay a foundation for quickly indexing the source text document based on the trend hot words.

[0072] In another embodiment, the format of the output result can be (doc_id, <trend hot words>), where doc_id represents the document identifier of the text document.

[0073] In yet another embodiment, based on the above-described embodiment, on the basis of the clustered trending hotwords, to screen out the target text documents from multiple text documents, the following method may also be adopted:

[0074] Based on the clustered trending hotwords and the document identifiers of the text documents corresponding to the trending hotwords, determine the document identifiers of the target text documents corresponding to the clustered trending hotwords;

[0075] Based on the document identifiers of the target text documents, screen out the target text documents from the multiple text documents.

[0076] In one embodiment, since multiple sets of document identifiers of text documents corresponding to the trending hotwords within the current time period can be obtained, that is, multiple sets (doc_id, <trending hotword>) are obtained. Moreover, since the clustered trending hotwords are derived from the trending hotwords, the document identifiers of the target text documents corresponding to the clustered trending hotwords can be determined from the obtained multiple sets (doc_id, <trending hotword>). Further, based on the document identifiers of the target text documents, the target text documents can be quickly screened out from the multiple text documents. Thus, it can lay a foundation for quickly and efficiently generating the event topic profile information within the current time period.

[0077] Figure 4 It is the second flowchart of the information processing method based on the GPT model provided by the present application.

[0078] To further introduce the information processing method based on the GPT model provided by the present application, the following will be combined with Figure 4 for illustration.

[0079] In an exemplary embodiment of the present application, combined with Figure 4 it can be known that after obtaining the event topic profile information within the current time period, the information processing method based on the GPT model may further include steps 410 to 430. The following will introduce each step separately.

[0080] In step 410, obtain multiple text documents within multiple subsequent time periods;

[0081] In step 420, according to the steps of obtaining the event topic profile information within the current time period by combining the multiple text documents within the current time period, based on the multiple text documents within multiple subsequent time periods, obtain multiple event topic profile information within multiple subsequent time periods;

[0082] In step 430, based on the event topic profile information within the current time period and the multiple event topic profile information within multiple subsequent time periods, form an event chain representing the development trend of the event topic.

[0083] In one embodiment, continuing with the embodiment described above, it is also possible to obtain multiple text documents within multiple subsequent time periods, and based on the steps of combining multiple text documents within the current time period as described above to obtain the event topic profile information within the current time period, that is, based on multiple text documents within multiple subsequent time periods, obtain multiple event topic profile information within multiple subsequent time periods according to steps 210 to 240. In other words, it is only necessary to replace the multiple text documents in the current time period in steps 210 to 240 with multiple text documents within multiple subsequent time periods.

[0084] It can be understood that regardless of the time period, the processing process of text documents is the same. In the application process, only the text documents in different time periods are changed.

[0085] In another embodiment, it is possible to classify event topics based on the event topic profile information within the current time period and the multiple event topic profile information within multiple subsequent time periods. Sort the profile information belonging to the same time topic in chronological order to form an event chain representing the development trend of the event topic. This facilitates users to quickly understand the development trend of the event topic and improves the user experience and satisfaction.

[0086] In another exemplary embodiment of the present application, continuing with Figure 3 the embodiment described above, to form an event chain representing the development trend of the event topic based on the event topic profile information within the current time period and the multiple event topic profile information within multiple subsequent time periods, the following method can be used:

[0087] Based on the event topic profile information within the current time period and the multiple event topic profile information within multiple subsequent time periods, classify the event topics within the current time period and the event topics within the subsequent time periods with event association relationships through binary graph matching to obtain the classified event topics;

[0088] Sort the event topic profile information corresponding to the classified event topics in chronological order to form an event chain representing the development trend of the event topic.

[0089] In one embodiment, event topics within the current time period having an event association relationship and event topics within multiple subsequent time periods can be classified to obtain classified event topics. It can be understood that the classified event topics are topics that discuss or express the same topic event in different time periods. Further, according to the chronological order in which the classified event topics occur, the event topic profile information corresponding to the classified event topics is sorted, so that an event chain representing the development trend of the event topic can be automatically generated, thereby facilitating the user to quickly understand the development trend of the event topic and improving the user experience and satisfaction.

[0090] In another embodiment, event topics within the current time period having an event association relationship and event topics within a subsequent time period can be classified through bipartite graph matching to obtain classified event topics.

[0091] Among them, a bipartite graph, also known as a two-part graph, is a special model in graph theory. Its characteristic is that assuming all vertices of a graph are divided into two non-overlapping sets, the two vertices associated with any one edge are distributed in different sets. This model structure is more suitable for analyzing event relationships in the time domain. Assuming each event is a node, multiple events on the same day are in different communities through community clustering (opposite to the community after clustering), and it is considered that there is no relationship. Then, for events on adjacent days, those belonging to the same event are connected by relationship edges. All events on adjacent days form a bipartite graph structure.

[0092] According to this rule, the time when the event nodes on each adjacent two days are connected is determined through bipartite graph matching. Then, on a larger time unit, an event development chain is formed. Thus, it lays a foundation for facilitating the user to quickly understand the development trend of the event topic and improving the user experience and satisfaction. According to the foregoing description, an information processing method based on a GPT model provided by the present application obtains multiple text documents within the current time period, extracts trend hot words in the text documents based on a preset first GPT model, constructs a co-occurrence matrix of the trend hot words within the current time period based on the multiple trend hot words, and obtains the matrix elements of the co-occurrence matrix; constructs a trend hot word graph of the trend hot words within the current time period based on the matrix elements; clusters the trend hot words based on the trend hot word graph to obtain the clustered trend hot words, and extracts a summary of the target text documents corresponding to the clustered trend hot words through a preset second GPT model to obtain event topic profile information within the current time period. It realizes the ability to automatically generate event topic profile information, thereby facilitating the user to quickly understand the development trend of the event topic and improving the user experience and satisfaction.

[0093] Figure 5 It is a schematic structural diagram of an information processing device based on a GPT model provided by the present application.

[0094] The information processing device based on the GPT model provided by the present application will be described below. The information processing device based on the GPT model described below can be correspondingly referred to the information processing method based on the GPT model described above.

[0095] In an exemplary embodiment of the present application, the information processing device based on the GPT model may include an acquisition module 510, a matrix construction module 520, a graph construction module 530, and an extraction module 540. Each module will be introduced separately below.

[0096] The acquisition module 510 can be configured to acquire multiple text documents within the current time period, and perform extraction processing on the trending hot words in the text documents based on a preset first GPT model, where the trending hot words refer to words whose occurrence frequency meets the set requirements within the preset statistical time;

[0097] The matrix construction module 520 can be configured to construct a co-occurrence matrix of the trending hot words within the current time period based on multiple trending hot words, and obtain the matrix elements of the co-occurrence matrix, where the matrix elements represent the co-occurrence frequency of different trending hot words in the same text document;

[0098] The graph construction module 530 can be configured to construct a trending hot word graph of the trending hot words within the current time period based on the matrix elements, where the trending hot word graph includes graph nodes and edges. The graph nodes in the trending hot word graph are the trending hot words within the current time period, and the edges connecting the graph nodes are the matrix elements corresponding to the trending hot words within the current time period;

[0099] The extraction module 540 can be configured to cluster the trending hot words based on the trending hot word graph to obtain the clustered trending hot words, and perform abstract extraction on the target text documents corresponding to the clustered trending hot words through a preset second GPT model to obtain the event topic overview information within the current time period.

[0100] In an exemplary embodiment of the present application, the graph construction module 530 adopts the following method to construct a co-occurrence matrix of the trending hot words within the current time period based on multiple trending hot words:

[0101] Determine the co-occurrence frequency of each trending hot word and other trending hot words in the same text;

[0102] Construct a co-occurrence matrix of the trending hot words within the current time period with the frequency as the matrix elements.

[0103] In an exemplary embodiment of the present application, the extraction module 540 can adopt the following method to cluster the trending hot words based on the trending hot word graph to obtain the clustered trending hot words:

[0104] Based on the trend hot word graph, cluster the graph nodes in the trend hot word graph through a community clustering algorithm to obtain the clustered graph nodes;

[0105] Cluster the trend hot words corresponding to the clustered graph nodes to obtain the clustered trend hot words.

[0106] In an exemplary embodiment of the present application, the extraction module 540 may further be configured to:

[0107] Based on the clustered trend hot words, screen out target text documents from multiple text documents, where the clustered trend hot words are extracted from the target text documents;

[0108] Determine a preset second text prompt, where the second text prompt is used to guide the second GPT model to output prompt information for the event topic overview information;

[0109] The extraction module 540 may also be implemented in the following manner to perform abstract extraction on the target text document corresponding to the clustered trend hot words through a preset second GPT model to obtain the event topic overview information in the current time period:

[0110] Input the target text document and the second text prompt into the second GPT model, so that the second GPT model performs abstract extraction on the target text document to obtain the event topic overview information in the current time period.

[0111] In an exemplary embodiment of the present application, the acquisition module 510 may further be configured to:

[0112] Determine a preset first text prompt, where the first text prompt is used to guide the first GPT model to output the prompt information of the preset format output result, where the preset format output result includes the trend hot words of the text document and the document identifier of the text document;

[0113] The acquisition module 510 may also be implemented in the following manner to perform extraction processing on the trend hot words in the text document based on a preset first GPT model:

[0114] Input the text document and the first text prompt into the first GPT model, so that the first GPT model performs extraction on the text document to obtain the trend hot words of the text document and the document identifier of the text document;

[0115] The extraction module 540 may also be implemented in the following manner to screen out the target text document from multiple text documents based on the clustered trend hot words:

[0116] Determine the document identifiers of the target text documents corresponding to the clustered trending hotwords based on the clustered trending hotwords and the document identifiers of the text documents corresponding to the trending hotwords.

[0117] Based on the document identifiers of the target text documents, screen out the target text documents from the multiple text documents.

[0118] In an exemplary embodiment of the present application, the extraction module 540 may also be configured to:

[0119] Obtain multiple text documents within multiple subsequent time periods;

[0120] According to the steps of obtaining the event topic profile information within the current time period by combining the multiple text documents within the current time period as described above, obtain multiple event topic profile information within multiple subsequent time periods based on the multiple text documents within multiple subsequent time periods;

[0121] Based on the event topic profile information within the current time period and the multiple event topic profile information within multiple subsequent time periods, form an event chain representing the development trend of the event topic.

[0122] In an exemplary embodiment of the present application, the extraction module 540 may also implement forming an event chain representing the development trend of the event topic based on the event topic profile information within the current time period and the multiple event topic profile information within multiple subsequent time periods in the following manner:

[0123] Based on the event topic profile information within the current time period and the multiple event topic profile information within multiple subsequent time periods, classify the event topics within the current time period and the event topics within the subsequent time periods having an event association relationship through bipartite graph matching to obtain the classified event topics;

[0124] Sort the event topic profile information corresponding to the classified event topics in chronological order to form an event chain representing the development trend of the event topic.

[0125] Figure 6 Illustrate a schematic diagram of the physical structure of an electronic device, such as Figure 6As shown, the electronic device may include: a processor 610, a communications interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communications interface 620, and the memory 630 complete communication with each other through the communication bus 640. The processor 610 may call logical instructions in the memory 630 to execute an information processing method based on the GPT model. The method includes: obtaining a plurality of text documents within the current time period, and extracting trend hot words in the text documents based on a preset first GPT model, where the trend hot words represent words whose occurrence frequency meets a set requirement within a preset statistical time; based on the plurality of trend hot words, constructing a co-occurrence matrix of the trend hot words within the current time period, and obtaining matrix elements of the co-occurrence matrix, where the matrix elements represent the co-occurrence frequency of different trend hot words in the same text document; based on the matrix elements, constructing a trend hot word graph of the trend hot words within the current time period, where the trend hot word graph includes graph nodes and edges, the graph nodes in the trend hot word graph are the trend hot words within the current time period, and the edges connecting the graph nodes are the matrix elements corresponding to the trend hot words within the current time period; clustering the trend hot words based on the trend hot word graph to obtain the clustered trend hot words, and extracting an abstract of the target text documents corresponding to the clustered trend hot words through a preset second GPT model to obtain event topic profile information within the current time period.

[0126] In addition, when the logical instructions in the above-mentioned memory 630 are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0127] On the other hand, the present application also provides a computer program product, which includes a computer program. The computer program can be stored on a computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the information processing method based on the GPT model provided by the above-mentioned various methods. The method includes: obtaining a plurality of text documents within the current time period, and performing extraction processing on the trend hot words in the text documents based on a preset first GPT model, where the trend hot words refer to words whose occurrence frequencies meet the set requirements within a preset statistical time; based on the plurality of trend hot words, constructing a co-occurrence matrix of the trend hot words within the current time period, and obtaining matrix elements of the co-occurrence matrix, where the matrix elements represent the co-occurrence frequencies of different trend hot words in the same text document; based on the matrix elements, constructing a trend hot word graph of the trend hot words within the current time period, where the trend hot word graph includes graph nodes and edges, the graph nodes in the trend hot word graph are the trend hot words within the current time period, and the edges connecting the graph nodes are the matrix elements corresponding to the trend hot words within the current time period; clustering the trend hot words based on the trend hot word graph to obtain the clustered trend hot words, and performing abstract extraction on the target text documents corresponding to the clustered trend hot words through a preset second GPT model to obtain the event topic overview information within the current time period.

[0128] On the other hand, the present application also provides a computer-readable storage medium, which includes a stored program. When the program runs, it executes the information processing method based on the GPT model provided by the above-mentioned various methods. The method includes: obtaining a plurality of text documents within the current time period, and performing extraction processing on the trend hot words in the text documents based on a preset first GPT model, where the trend hot words refer to words whose occurrence frequencies meet the set requirements within a preset statistical time; based on the plurality of trend hot words, constructing a co-occurrence matrix of the trend hot words within the current time period, and obtaining matrix elements of the co-occurrence matrix, where the matrix elements represent the co-occurrence frequencies of different trend hot words in the same text document; based on the matrix elements, constructing a trend hot word graph of the trend hot words within the current time period, where the trend hot word graph includes graph nodes and edges, the graph nodes in the trend hot word graph are the trend hot words within the current time period, and the edges connecting the graph nodes are the matrix elements corresponding to the trend hot words within the current time period; clustering the trend hot words based on the trend hot word graph to obtain the clustered trend hot words, and performing abstract extraction on the target text documents corresponding to the clustered trend hot words through a preset second GPT model to obtain the event topic overview information within the current time period.

[0129] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.

[0130] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0131] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. An information processing method based on the GPT model, characterized in that, the method includes: Obtain multiple text documents within the current time period, and perform extraction processing on the trend hot words in the text documents based on a preset first GPT model, where the trend hot words refer to words whose occurrence frequency meets the set requirements within a preset statistical time; Based on multiple said trend hot words, construct a co-occurrence matrix of the trend hot words within the current time period, and obtain the matrix elements of the co-occurrence matrix, where the matrix elements represent the co-occurrence frequency of different trend hot words in the same text document; Based on the matrix elements, construct a trend hot word graph of the trend hot words within the current time period, where the trend hot word graph includes graph nodes and edges, the graph nodes in the trend hot word graph are the trend hot words within the current time period, and the edges connecting the graph nodes are the matrix elements corresponding to the trend hot words within the current time period; Cluster the trend hot words based on the trend hot word graph to obtain the clustered trend hot words, and perform abstract extraction on the target text documents corresponding to the clustered trend hot words through a preset second GPT model to obtain the event topic overview information within the current time period.

2. The information processing method based on the GPT model according to claim 1, characterized in that, the constructing a co-occurrence matrix of the trend hot words within the current time period based on multiple said trend hot words specifically includes: Determine the occurrence frequency of each said trend hot word and other trend hot words in the same text; Construct a co-occurrence matrix of the trend hot words within the current time period with the frequency as the matrix elements.

3. The information processing method based on the GPT model according to claim 1 or 2, characterized in that, the clustering the trend hot words based on the trend hot word graph to obtain the clustered trend hot words specifically includes: Based on the trend hot word graph, cluster the graph nodes in the trend hot word graph through a community clustering algorithm to obtain the clustered graph nodes; Cluster the trend hot words corresponding to the clustered graph nodes to obtain the clustered trend hot words.

4. The information processing method based on the GPT model according to claim 1, characterized in that, before the performing abstract extraction on the target text documents corresponding to the clustered trend hot words through a preset second GPT model to obtain the event topic overview information within the current time period, the method further includes: Based on the clustered trend hot words, screen out the target text documents from multiple said text documents, where the clustered trend hot words are extracted from the target text documents; Determine a preset second text prompt, where the second text prompt is used to guide the second GPT model to output prompt information for the event topic overview information; the performing abstract extraction on the target text documents corresponding to the clustered trend hot words through a preset second GPT model to obtain the event topic overview information within the current time period specifically includes: Input the target text document and the second text prompt into the second GPT model, so that the second GPT model extracts a summary of the target text document to obtain event topic overview information for the current time period.

5. The information processing method based on the GPT model according to claim 4, wherein, before the method extracts trend hot words in the text document based on the preset first GPT model, the method further includes: determining a preset first text prompt, where the first text prompt is used to guide the first GPT model to output prompt information in a preset format, and the preset format output result includes the trend hot words of the text document and the document identifier of the text document; The process of extracting trend hot words in the text document based on the preset first GPT model specifically includes: Input the text document and the first text prompt into the first GPT model, so that the first GPT model extracts the text document to obtain the trend hot words of the text document and the document identifier of the text document; The process of screening out the target text document from multiple text documents based on the clustered trend hot words specifically includes: Based on the clustered trend hot words and the document identifiers of the text documents corresponding to the trend hot words, determine the document identifier of the target text document corresponding to the clustered trend hot words; Based on the document identifier of the target text document, screen out the target text document from the multiple text documents.

6. The information processing method based on the GPT model according to claim 1, wherein, after obtaining the event topic overview information for the current time period, the method further includes: obtaining multiple text documents in multiple subsequent time periods; According to the steps of combining multiple text documents in the current time period to obtain the event topic overview information for the current time period, based on multiple text documents in multiple subsequent time periods, obtain multiple event topic overview information for multiple subsequent time periods; Based on the event topic overview information for the current time period and the multiple event topic overview information for multiple subsequent time periods, form an event chain representing the development trend of the event topic.

7. The information processing method based on the GPT model according to claim 6, wherein, The process of forming an event chain representing the development trend of the event topic based on the event topic overview information for the current time period and the multiple event topic overview information for multiple subsequent time periods specifically includes: Based on the event topic overview information for the current time period and the multiple event topic overview information for multiple subsequent time periods, classify the event topics in the current time period and the event topics in the subsequent time periods with event association relationships through bipartite graph matching to obtain the classified event topics; Arrange the event topic overview information corresponding to the classified event topics in chronological order to form an event chain representing the development trend of the event topic.

8. An information processing device based on the GPT model, characterized in that, the device includes: an acquisition module, configured to acquire a plurality of text documents within a current time period, and perform extraction processing on trend hot words in the text documents based on a preset first GPT model, where the trend hot words represent words whose occurrence frequencies meet set requirements within a preset statistical time; a matrix construction module, configured to construct a co-occurrence matrix of trend hot words within the current time period based on the plurality of trend hot words, and obtain matrix elements of the co-occurrence matrix, where the matrix elements represent the co-occurrence frequencies of different trend hot words in the same text document; a graph construction module, configured to construct a trend hot word graph of trend hot words within the current time period based on the matrix elements, where the trend hot word graph includes graph nodes and edges, the graph nodes in the trend hot word graph are the trend hot words within the current time period, and the edges connecting the graph nodes are the matrix elements corresponding to the trend hot words within the current time period; an extraction module, configured to cluster the trend hot words based on the trend hot word graph to obtain the clustered trend hot words, and perform abstract extraction on the target text documents corresponding to the clustered trend hot words through a preset second GPT model to obtain event topic overview information within the current time period.

9. A computer-readable storage medium, characterized in that, the computer-readable storage medium includes a stored program, where the program, when running, executes the information processing method based on the GPT model according to any one of claims 1 to 7.

10. An electronic device, including a memory and a processor, characterized in that, a computer program is stored in the memory, and the processor is configured to execute the information processing method based on the GPT model according to any one of claims 1 to 7 through the computer program.