Industrial information processing method and device
Through the large language model, the automatic matching, classification, summary and summary of industry information is achieved, and the problems of insufficient automation and low processing accuracy in the existing technology are solved, and efficient and accurate processing of industry information is achieved, which is suitable for multiple industries.
Patent Information
- Application Number
- CN202510321224.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-13
AI Technical Summary
The existing technology is insufficient in the processing of industry information, which is prone to subjective deviations, and is difficult to quickly process massive information, poor processing accuracy and scalability, and is difficult to directly apply across industries and fields.
By obtaining the initial information collection and using the large language model, industry matching, classification, summary and summary are carried out based on pre-designed prompts to form the target information collection and corresponding information.
It realizes efficient and accurate processing of industry information, reduces manual intervention, improves processing efficiency and accuracy, is suitable for various industries, and has cross-industry and cross-field application capabilities.
Smart Images

Figure CN120144756A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of big data technology, and in particular, to a method, apparatus, computer device, computer-readable storage medium, and computer program product for processing industry information. Background Art
[0002] With the rapid development of Internet technology, the amount of information has increased explosively. Enterprises urgently need to obtain and process industry information efficiently and accurately to support strategic decision-making. From information-intensive industries such as technology, finance, and healthcare to manufacturing industries such as photovoltaic and materials, how to extract valuable information from a large amount of publicly available information and process it efficiently has become a common issue for all industries.
[0003] However, the existing industry information processing methods still have the following defects: (1) Through manual collection, screening, and classification, the degree of automation is insufficient, time-consuming and laborious, prone to subjective biases, and unable to process a large amount of information quickly; (2) Using machine learning models trained with high-quality labeled data for information processing, the processing accuracy is insufficient, the intelligent analysis is lacking, and the industry scalability is poor, making it difficult to be directly applied across industries and fields.
[0004] It should be noted that the above content is not necessarily prior art and is not used to limit the patent protection scope of the present application. Summary of the Invention
[0005] Embodiments of the present application provide a method, apparatus, computer device, computer-readable storage medium, and computer program product for processing industry information to solve or alleviate one or more of the above technical problems.
[0006] One aspect of the embodiments of the present application provides a method for processing industry information, the method including: Obtaining an initial information set, where the initial information set includes multiple pieces of initial information corresponding to a target industry; Based on a pre-designed prompt, the initial information set, and a large language model, obtaining a target information set and target information corresponding to the target information set; Wherein, the target information set includes multiple pieces of target information, and the target information includes initial information with a matching degree greater than a preset threshold with the target industry; the target information includes: the information category of each piece of target information, the summary information of each piece of target information, and the summary information of multiple pieces of target information for the same information category; the pre-designed prompt is used to guide the large language model to perform industry matching, classification, summarization, and summary on the multiple pieces of initial information.
[0007] Optionally, based on pre-designed prompts, the initial information set, and a large language model, obtain a target information set and target information corresponding to the target information set, including: Input the pre-designed first prompt and the initial information set into the first large language model to determine the multiple pieces of target information from the multiple pieces of initial information through the first large language model; Obtain the target information based on the multiple pieces of target information; Wherein, the first prompt includes the industry background, industry theme, and / or industry terms of the target industry, and the first prompt is used to guide the first large language model: obtain the industry information of the target industry according to the industry background, industry theme, and / or industry terms of the target industry; perform semantic matching between the industry information and the multiple pieces of initial information, and determine the multiple pieces of target information from the multiple pieces of initial information.
[0008] Optionally, performing semantic matching between the industry information and the multiple pieces of initial information, and determining the multiple pieces of target information from the multiple pieces of initial information, includes: Perform semantic matching between the industry information and the multiple pieces of initial information to obtain multiple pieces of valid information, forming a valid information set, where the semantic matching degree between the valid information and the industry information is greater than the preset threshold; Based on a preset historical information database, compare and remove duplicates from the valid information set to obtain a deduplicated information set; Perform self-deduplication on the deduplicated information set to obtain the target information set, and the target information set includes the multiple non-repeated pieces of target information.
[0009] Optionally, obtaining the target information based on the multiple pieces of target information includes: Input the pre-designed second prompt and the multiple pieces of target information into the second large language model to determine the information category of each piece of target information through the second large language model; Wherein, the second prompt includes a preset multiple information categories and the definition of each information category, and the second prompt is used to guide the second large language model: classify the multiple pieces of target information according to the preset multiple information categories and the definition of each information category to obtain the information category of each piece of target information.
[0010] Optionally, obtaining the target information based on the multiple pieces of target information further includes: Determine the multiple pieces of target information corresponding to each information category; Input the pre-designed third prompt and the corresponding multiple pieces of target information into the third large language model to obtain summary information of the multiple pieces of target information for the same information category through the third large language model; Among them, the third prompt includes summary rules and / or summary examples, and the third prompt is used to guide the third large language model to generate summary information for multiple pieces of target information in the same information category according to the summary rules and / or summary examples.
[0011] Optionally, obtaining the target information, based on the multiple pieces of target information, includes: Inputting a pre-designed fourth prompt and the multiple pieces of target information into a fourth large language model to generate summary information for each piece of target information through the fourth large language model; Among them, the fourth prompt includes summary rules and / or summary examples, and the fourth prompt is used to guide the fourth large language model to generate summary information for each piece of target information according to the summary rules and / or summary examples.
[0012] Optionally, the industry information processing method further includes: Determining the type of the target object; Based on the type of the target object, obtaining a target push template from multiple pre-configured push templates, where the target push template includes target layout rules; Generating a target information report based on the target information and the target layout rules; Pushing the target information report to the target object.
[0013] Optionally, obtaining the initial information set includes: Collecting multiple pieces of initial information corresponding to the target industry through web crawler technology based on a preset information acquisition rule to form the initial information set; Among them, the information acquisition rule includes the collection frequency, time period, information source, and / or information type.
[0014] Another aspect of the embodiments of the present application provides an information processing device, and the device includes: A first acquisition module, configured to acquire an initial information set, where the initial information set includes multiple pieces of initial information corresponding to a target industry; A second acquisition module, configured to acquire a target information set and target information corresponding to the target information set based on a pre-designed prompt, the initial information set, and a large language model; Among them, the target information set includes multiple pieces of target information, and the target information includes initial information with a matching degree greater than a preset threshold with the target industry; the target information includes: the information category of each piece of target information, the summary information of each piece of target information, and the summary information of multiple pieces of target information for the same information category; the pre-designed prompt is used to guide the large language model to perform industry matching, classification, summarization and summarization on the multiple pieces of initial information.
[0015] Another aspect of the embodiments of the present application provides a computer device, including: At least one processor; and A memory communicatively connected to the at least one processor; Wherein: the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as described above.
[0016] Another aspect of the embodiments of the present application provides a computer-readable storage medium, in which computer instructions are stored, and when the computer instructions are executed by a processor, the method as described above is implemented.
[0017] Another aspect of the embodiments of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method as described above is implemented.
[0018] The embodiments of the present application adopting the above technical solutions may include the following advantages: Obtain multiple pieces of initial information related to the target industry. Input the pre-designed prompt and multiple pieces of initial information into the large language model, and the large language model, under the guidance of the prompt: obtain multiple pieces of target information with a matching degree greater than a preset threshold with the target industry from the multiple pieces of initial information to form a target information set; determine the information category and generate summary information for each piece of target information; and generate summary information for multiple pieces of target information under the same information category. It can be seen that the embodiments of the present application perform automated information processing through the large language model, reduce manual intervention, improve processing efficiency and accuracy, and are applicable to various industries. Description of the Drawings
[0019] The drawings exemplarily show the embodiments and form a part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments of the embodiments. The shown embodiments are only for illustrative purposes and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0020] Figure 1 Schematically shows a flowchart of the industry information processing method according to Embodiment 1 of the present application; Figure 2 Schematically shows a flowchart of the industry information processing method according to Embodiment 1 of the present application; Figure 3 Schematically shows Figure 2 the sub-step flowchart of step S200 in; Figure 4 Schematically shows the sub-step flowchart of step S400 according to Embodiment 1 of the present application; Figure 5 Schematically shows Figure 3 the sub-step flowchart of step S302 in; Figure 6 Schematically shows the new flowchart of the industry information processing method according to Embodiment 1 of the present application; Figure 7 Schematically shows the block diagram of the information processing device according to Embodiment 2 of the present application; and Figure 8 Schematically shows the hardware architecture diagram of the computer device according to Embodiment 3 of the present application. Detailed implementation manners
[0021] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0022] It should be noted that the descriptions involving "first", "second", etc. in the embodiments of the present application are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between the various embodiments may be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present application.
[0023] In the description of the present application, it should be understood that the numerical labels before the steps do not identify the order of execution of the steps, but are only used to facilitate the description of the present application and distinguish each step, and thus cannot be understood as a limitation to the present application.
[0024] First, provide the term explanations involved in the present application: SVM (Support Vector Machine): Support Vector Machine, a machine learning algorithm for classification and regression. It finds a hyperplane in a high-dimensional space to separate data points into different classes. SVM is especially suitable for handling small-sample data and can effectively deal with non-linear problems.
[0025] RNN (Recurrent Neural Network): Recurrent Neural Network, a deep learning model used to process sequential data (such as time series, text, or speech). By feeding the information from the previous moment back to the current moment, it can capture the temporal dependence and context relationship of the data.
[0026] Transformer: A deep learning model based on the attention mechanism, used for natural language processing tasks (such as translation, text generation). Different from RNN, Transformer can process sequences in parallel with higher efficiency and can capture global dependencies, which is the basic architecture of large language models.
[0027] API (Application Programming Interface): Application Programming Interface, an interface that allows different software applications to communicate with each other, providing predefined methods and data formats. APIs can be used for system integration and service calls.
[0028] OA (Office Automation): Office Automation, which refers to using information technology and systems to automate office processes, such as Feishu and DingTalk.
[0029] The embodiments of this application provide a technical solution for information processing. In this technical solution: (1) Based on the powerful semantic understanding and generation capabilities of large language models, automatic screening and classification of information can be carried out, which can greatly improve the accuracy of industry information processing and reduce information processing errors. (2) Achieve full-process automation from information crawling to classification, deduplication, summarization, and summary, reduce manual intervention, significantly improve efficiency, and with large language models as the core, information processing can be completed under zero-sample or small-sample conditions, reducing the need for manual annotation; (3) Through intelligent analysis of information by large language models, it is also possible to automatically adjust the information processing strategy according to industry-specific needs, so as to provide more personalized information recommendations; (4) Provide an intelligent and accurate information processing and pushing mechanism, and integrate it into the client of the enterprise OA system through the API method, so as to timely provide the target object (user) with industry information that better meets their needs and improve the user experience. See the following for details.
[0030] The technical solutions of the present application will be introduced through multiple embodiments. It should be noted that these embodiments can be implemented in various different forms and should not be construed as being limited only to the embodiments described herein.
[0031] Embodiment 1 Figure 1 A flowchart of the industry information processing method according to Embodiment 1 of the present application is schematically shown.
[0032] As Figure 1 shown, the industry information processing method may include steps S100 to S102, where: Step S100, obtaining an initial information set, where the initial information set includes multiple pieces of initial information corresponding to a target industry.
[0033] Step S102, based on a pre-designed prompt, the initial information set, and a large language model, obtaining a target information set and target information corresponding to the target information set. Among them, the target information set includes multiple pieces of target information, and the target information includes initial information with a matching degree greater than a preset threshold with the target industry; the target information includes: the information category of each piece of target information, the abstract information of each piece of target information, and the summary information of multiple pieces of target information for the same information category; the pre-designed prompt is used to guide the large language model to perform industry matching, classification, abstraction, and summary on the multiple pieces of initial information.
[0034] The industry information processing method provided in this embodiment performs automated information processing through a large language model, reduces manual intervention, improves processing efficiency and accuracy, and is applicable to various industries. The following will be combined with Figure 1 to elaborate in detail on each step in steps S100 to S102 and optional other steps.
[0035] Step S100, obtaining an initial information set, where the initial information set includes multiple pieces of initial information corresponding to a target industry.
[0036] The information can be news reports, academic papers, forum posts, dynamics on social media, blog articles, informal reports, etc. The industry can be technology, finance, healthcare, photovoltaic, materials, etc. Hereinafter, the photovoltaic industry will be regarded as the target industry to give an exemplary introduction to the industry information processing method provided in the embodiments of the present application. For example, multiple pieces of initial information can be obtained from multiple public information sources on the Internet (such as news websites, industry forums, social media, etc.) to form an initial information set. Among them, the initial information can be information with a certain relevance to the photovoltaic industry. In some embodiments, the obtained initial information can also be desensitized to ensure data compliance. The following provides an exemplary solution.
[0037] In an optional embodiment, step S100 may include: based on a preset information acquisition rule, collecting multiple pieces of initial information corresponding to the target industry through web crawler technology to form the initial information set; wherein, the information acquisition rule includes the acquisition frequency, time period, information source, and / or information type.
[0038] Exemplarily, a custom configuration function may be provided, and a user (such as an administrator) can flexibly configure the information acquisition rule according to actual needs, such as: acquisition frequency, time period, information source, and / or information type, etc. Dynamic adjustment of the information acquisition rule is also supported, such as: adding, deleting, or replacing information sources according to requirements. Based on the configured information acquisition rule, a large number of pieces of information related to the target industry, that is, initial information, are automatically crawled on the Internet through web crawler technology to form an initial information set.
[0039] In this embodiment, the information acquisition rule can be customized and can also be flexibly adjusted according to industry needs and information timeliness to ensure that the most relevant and timely industry information can be obtained. Support for multi-channel information acquisition can integrate information from different platforms and improve the breadth and comprehensiveness of information. The latest industry information on the Internet is obtained through automated crawling technology, reducing the time and energy consumption of manual collection.
[0040] Step S102, based on a pre-designed prompt, the initial information set, and a large language model, obtain a target information set and target information corresponding to the target information set. Wherein, the target information set includes multiple pieces of target information, and the target information includes initial information with a matching degree greater than a preset threshold with the target industry; the target information includes: the information category of each piece of target information, the abstract information of each piece of target information, and the summary information of multiple pieces of target information for the same information category; the pre-designed prompt is used to guide the large language model to perform industry matching, classification, summarization, and summary on the multiple pieces of initial information.
[0041] Large Language Model (LLM): A large-scale natural language processing model trained based on deep learning, with parameters that can reach billions or even hundreds of billions, capable of understanding and generating natural language, and having a wide range of application scenarios, including: text generation, translation, question answering, etc.
[0042] A prompt is input text, an image, audio, etc. used to guide a large language model to generate a specific output when interacting with the large language model. For example, asking questions or providing context information to guide the model to generate the expected answer. The design of the prompt directly affects the quality of the model's response. Predesigned prompts can include industry information, as well as rules or examples for classification, summarization, and abstraction. Inputting the predesigned prompt and multiple initial pieces of information into the large language model can guide the large language model to perform industry matching (screening), classification, summarization, and abstraction on the multiple initial pieces of information, obtain multiple target pieces of information with a matching degree greater than a preset threshold with the target industry, form a target information set, determine the information category and summary information of each target piece of information, generate summary information for multiple target pieces of information under the same information category, and obtain the target information corresponding to the target information set. In practical applications, under the guidance of the predesigned prompt, the large language model can complete the entire process of information processing (screening + classification + summarization + abstraction) at one time, from input to output, significantly improving the information processing efficiency. Of course, according to different processing tasks (screening, classification, summarization, abstraction, deduplication, etc.), the information processing process can be split into multiple concurrently executed modules, as Figure 2 shown. Each module can apply the same or different large language models to process one or more tasks respectively. For example, classification and summarization can be divided into one module, or classification and abstraction can be divided into one module. The modules can be set according to actual needs, which can effectively balance the efficiency and quality of information processing. Of course, each module can also be replaced with other algorithms instead of the large language model. For example, an automated information aggregation platform (such as a news crawler system, an information aggregation tool, etc.) can be used to crawl information from multiple sources and perform simple classification and summary generation. Thus, the processing quality can be optimized for different tasks, the accuracy and reliability of information processing can be improved, and the processing strategy for each task can be adjusted according to actual needs, enhancing the flexibility and scalability of information processing. The following will be introduced exemplarily in combination with multiple embodiments.
[0043] In an alternative embodiment, as Figure 3 shown, step S200 may include: Step S300, input the predesigned first prompt and the initial information set into the first large language model to determine the multiple target pieces of information from the multiple initial pieces of information through the first large language model. Wherein, the first prompt includes the industry background, industry theme, and / or industry terms of the target industry, and the first prompt is used to guide the first large language model: S400: Obtain the industry information of the target industry according to the industry background, industry theme, and / or industry terms of the target industry; S402: Perform semantic matching between the industry information and the multiple initial pieces of information to determine the multiple target pieces of information from the multiple initial pieces of information.
[0044] Step S302: Obtain the target information based on the multiple pieces of target information.
[0045] Exemplarily, the industry background, industry themes (such as frequently mentioned words), and industry terms (such as professional terms) of the target industry can be provided as the first prompt. The first prompt and multiple pieces of initial information are input into the first large language model. It should be noted that the first large language model, the second large language model, the third large language model, and the fourth large language model mentioned later can be the same large language model, different large language models, or the same large language model, which is not limited herein. The first large language model can automatically analyze the first prompt, understand the target industry, and determine the industry information. It can also automatically analyze the multiple pieces of initial information crawled and identify the implicitly industry-related information therein. For example, the relationship between "improving the charging efficiency of lithium batteries" and the new energy industry. The industry information is semantically matched with each piece of target information, and only the information with a high relevance to the target industry (such as: the semantic matching degree is greater than the preset threshold) is retained to achieve intelligent screening of the information. Based on the multiple pieces of target information, the corresponding target information can be obtained through a large language model or other algorithms (such as: natural language processing technology, etc.). In some embodiments, intelligent screening of the information can also be achieved through a classifier (such as: SVM).
[0046] In this embodiment, based on the custom prompt words, the intelligent screening of the information is realized through the semantic understanding ability of the large language model, automatically judging the quality of the information and its relevance to the target industry, not limited to simple keyword matching, ensuring that irrelevant or low-quality information is filtered out, avoiding manual intervention, and improving the efficiency and accuracy of information processing.
[0047] In an alternative embodiment, as Figure 4 shown, step S400 may include: Step S500: Semantically match the industry information with the multiple pieces of initial information to obtain multiple pieces of valid information, forming a valid information set, where the semantic matching degree between the valid information and the industry information is greater than the preset threshold.
[0048] Step S502: Based on the preset historical information database, compare and remove duplicates from the valid information set to obtain a deduplicated information set.
[0049] Step S504: Perform self-deduplication on the deduplicated information set to obtain the target information set, where the target information set includes the multiple pieces of non-repeated target information.
[0050] Exemplarily, the large language model performs semantic matching between industry information and multiple initial news items, and can determine multiple valid news items to form a set of valid news items. Among them, the valid news items can be high-quality initial news items with a semantic matching degree greater than a preset threshold with the industry information. Compare the set of valid news items with multiple historical news items in the pre-configured historical news database to remove duplicates, and automatically eliminate duplicate valid news items to obtain a deduplicated news item set. Further, based on the content such as the title and body of the news items, the deduplicated news item set can be self-deduplicated to obtain a target news item set. The target news item set includes multiple target news items with a matching degree greater than a preset threshold with the target industry and no duplicates.
[0051] In this embodiment, by combining self-deduplication and historical comparison deduplication, non-duplicate news items with high industry relevance are retained, improving the quality of news items and ensuring that the news items used for subsequent pushing have high quality and high value.
[0052] In an alternative embodiment, step S300 may include: inputting a pre-designed second prompt and the multiple target news items into a second large language model to determine the news item category of each target news item through the second large language model; wherein, the second prompt includes multiple preset news item categories and the definition of each news item category, and the second prompt is used to guide the second large language model to classify the multiple target news items according to the multiple preset news item categories and the definition of each news item category to obtain the news item category of each target news item.
[0053] Exemplarily, multiple preset news item categories (such as industry news, market trends, investment and financing, technology frontiers, etc.) and the definition of each news item category (such as: a text description, a set of images or audio and video, etc.) can be provided as the second prompt. Input the second prompt and multiple target news items into the second large language model. The second large language model can automatically analyze the second prompt, deeply understand each news item category, and clarify the classification basis. Through the powerful reasoning and generation ability of the large language model, automatically analyze each target news item and match the corresponding news item category, which can achieve efficient classification of news items. Among them, the news item category can also be dynamically adjusted according to actual needs, for example: adding, deleting or replacing, and the definition of the news item category can also be modified in real time to improve the flexibility and adaptability of news item processing.
[0054] In this embodiment, using the large language model to automatically classify news items can greatly reduce manual intervention, achieve accurate classification, and improve the user experience.
[0055] In an alternative embodiment, as Figure 5 shown, step S302 may further include: Step S600, determining multiple target news items corresponding to each news item category.
[0056] Step S602: Input the pre-designed third prompt and the corresponding multiple pieces of target information into a third large language model to obtain summary information for the multiple pieces of target information in the same information category through the third large language model. The third prompt includes a summary rule and / or a summary example, and the third prompt is used to guide the third large language model to generate summary information for the multiple pieces of target information in the same information category according to the summary rule and / or the summary example.
[0057] Exemplarily, first obtain multiple pieces of target information corresponding to each information category, and at the same time, a summary rule and / or a summary example can be provided as the third prompt. Input the third prompt and the multiple pieces of target information corresponding to one information category into the third large language model. The third large language model can automatically analyze the third prompt and learn the inductive summary method. Through the powerful reasoning and generation ability of the large language model, automatically analyze each piece of target information in the same information category, extract the key information of the multiple pieces of target information for inductive summary, and the summary information for the multiple pieces of target information in the same information category can be obtained. Among them, the summary rule and / or the summary example can also be dynamically adjusted according to actual needs, such as adding, deleting, or replacing, to further improve the flexibility and adaptability of information processing.
[0058] In this embodiment, using the large model to automatically perform intelligent inductive summary of information can help users quickly capture the key information corresponding to each information category and improve the usability of information.
[0059] In an alternative embodiment, step S302 may include: Input the pre-designed fourth prompt and the multiple pieces of target information into a fourth large language model to generate summary information for each piece of target information through the fourth large language model; where the fourth prompt includes a summary rule and / or a summary example, and the fourth prompt is used to guide the fourth large language model to generate summary information for each piece of target information according to the summary rule and / or the summary example.
[0060] Exemplarily, a summary rule and / or a summary example can be provided as the fourth prompt. Input the fourth prompt and multiple pieces of target information into the fourth large language model. The fourth large language model can automatically analyze the fourth prompt and learn the summary generation method. Through the powerful reasoning and generation ability of the large language model, automatically analyze each piece of target information, simplify and generate a summary, and the summary information for each piece of target information can be obtained. Among them, the summary rule and / or the summary example can also be dynamically adjusted according to actual needs, such as adding, deleting, or replacing, to further improve the flexibility and adaptability of information processing. In some embodiments, a deep learning model (such as RNN or Transformer) can also be used to process the target information to generate corresponding summary information. The deep learning model can be pre-trained through high-quality labeled data.
[0061] In this embodiment, the large language model is used to automatically generate abstracts, which greatly reduces manual intervention and processing time, obtains concise information abstracts, helps users quickly understand the core content of the information, and optimizes the user experience.
[0062] In an alternative embodiment, as Figure 6 shown, the industry information processing method may further include: Step S700, determining the type of the target object.
[0063] Step S702, based on the type of the target object, obtaining a target push template from multiple pre-configured push templates, where the target push template includes target layout rules.
[0064] Step S704, generating a target information report based on the target information and the target layout rules.
[0065] Step S706, pushing the target information report to the target object.
[0066] Exemplarily, considering the differences in the industries or user groups the information faces, multiple push templates can be pre-configured. One push template can correspond to one type of industry or one type of user. The push template can include layout rules, such as set styles, formats, colors, fonts, etc., for indicating how information is combined and arranged to generate a structured report. Specifically, the type of the target object (such as industry, user) can be determined first. According to the type of the target object, the target push template can be obtained from multiple pre-configured push templates. The target push template includes target layout rules. Based on the target layout rules, the target information, such as the information category, abstract information, and / or summary information of the target information, is combined and arranged to generate a target information report with high readability. Finally, through the API interface integrated in the OA system, the target information report can be pushed to the client of the target object in real time to ensure that users can receive the latest industry dynamics in the first time.
[0067] In this embodiment, the format and content presentation method of the report can be flexibly adjusted according to different industries and user requirements, taking into account the personalization and applicability of information push. Through intelligent layout, it is ensured that the finally generated industry information report has good structure and readability, improving the professionalism and visual effect of the report. Through the integration with the OA system, the processed industry information can be efficiently and timely pushed to the target object, ensuring the instant transmission of information, and further optimizing the user experience through real-time push.
[0068] The industry information processing method provided by the embodiments of the present application can be directly applied across industries and fields, and can also process complex knowledge across fields.
[0069] To make this application easier to understand, an exemplary application is provided below.
[0070] S1: According to the pre-set information acquisition rules (including acquisition frequency, time period, information source, and / or information type), automatically crawl information related to the target industry through web crawler technology, that is, initial information.
[0071] S2: Guide the large language model through pre-designed prompts: automatically analyze the crawled initial information, and then intelligently screen the initial information according to the matching degree with the target industry, only retaining the information with a high degree of relevance to the target industry. It also performs self-duplicate removal based on the title and text of the information, and compares and removes duplicates with the existing information in the historical information database, automatically eliminating duplicate information to obtain target information.
[0072] S3: Guide the large language model to automatically classify, generate summaries, and intelligently summarize and generalize the screened information (target information) through pre-designed prompts to obtain corresponding target information. Among them, the target information can be further customized according to user needs to ensure that the presentation form of the information meets industry requirements and user habits.
[0073] S4: According to the customized push template, combine and arrange the target information to generate a target information report with high readability and push it to the target object.
[0074] In this exemplary application: combining the large language model with the industry information processing process, through intelligent screening, classification, summarization, generalization, and typesetting, greatly improves the processing efficiency and accuracy of industry information, and at the same time realizes the functions of automated and personalized push.
[0075] Embodiment 2 Figure 7 Schematically shows a block diagram of an information processing device according to Embodiment 2 of the present application. The device can be divided into one or more program modules. One or more program modules are stored in a storage medium and executed by one or more processors to complete the embodiments of the present application. The program modules referred to in the embodiments of the present application refer to a series of computer program instruction segments that can complete specific functions. The following description will specifically introduce the functions of each program module in this embodiment. As Figure 7 shown, the device 1000 may include: a first acquisition module 1100 and a second acquisition module 1200, where: The first acquisition module 1100 is used to acquire an initial information set, and the initial information set includes multiple pieces of initial information corresponding to the target industry; The second acquisition module 1200 is configured to obtain a target information set and target information corresponding to the target information set based on a pre-designed prompt, the initial information set, and a large language model; Wherein, the target information set includes multiple pieces of target information, and the target information includes initial information with a matching degree greater than a preset threshold with the target industry; the target information includes: the information category of each piece of target information, the summary information of each piece of target information, and the summary information of multiple pieces of target information belonging to the same information category; the pre-designed prompt is used to guide the large language model to perform industry matching, classification, summarization, and summarization on the multiple pieces of initial information.
[0076] As an optional embodiment, obtaining a target information set and target information corresponding to the target information set based on a pre-designed prompt, the initial information set, and a large language model includes: Inputting a pre-designed first prompt and the initial information set into a first large language model to determine the multiple pieces of target information from the multiple pieces of initial information through the first large language model; Obtaining the target information based on the multiple pieces of target information; Wherein, the first prompt includes the industry background, industry theme, and / or industry terms of the target industry, and the first prompt is used to guide the first large language model to: obtain the industry information of the target industry according to the industry background, industry theme, and / or industry terms of the target industry; perform semantic matching between the industry information and the multiple pieces of initial information, and determine the multiple pieces of target information from the multiple pieces of initial information.
[0077] As an optional embodiment, performing semantic matching between the industry information and the multiple pieces of initial information, and determining the multiple pieces of target information from the multiple pieces of initial information includes: Performing semantic matching between the industry information and the multiple pieces of initial information to obtain multiple pieces of valid information, forming a valid information set, where the semantic matching degree between the valid information and the industry information is greater than the preset threshold; Based on a preset historical information database, comparing and removing duplicates from the valid information set to obtain a deduplicated information set; Performing self-deduplication on the deduplicated information set to obtain the target information set, where the target information set includes the multiple non-duplicate pieces of target information.
[0078] As an optional embodiment, obtaining the target information based on the multiple pieces of target information includes: Inputting a pre-designed second prompt and the multiple pieces of target information into a second large language model to determine the information category of each piece of target information through the second large language model; Among them, the second prompt includes a plurality of preset information categories and the definition of each information category, and the second prompt is used to guide the second large language model: according to the plurality of preset information categories and the definition of each information category, classify the multiple target information to obtain the information category of each target information.
[0079] As an optional embodiment, based on the multiple target information, obtaining the target information further includes: Determine multiple target information corresponding to each information category; Input a preset third prompt and the corresponding multiple target information into a third large language model to obtain summary information of the multiple target information for the same information category through the third large language model; Among them, the third prompt includes a summary rule and / or a summary example, and the third prompt is used to guide the third large language model: according to the summary rule and / or the summary example, generate summary information of the multiple target information for the same information category.
[0080] As an optional embodiment, based on the multiple target information, obtaining the target information includes: Input a preset fourth prompt and the multiple target information into a fourth large language model to generate summary information of each target information through the fourth large language model; Among them, the fourth prompt includes a summary rule and / or a summary example, and the fourth prompt is used to guide the fourth large language model: according to the summary rule and / or the summary example, generate summary information of each target information.
[0081] As an optional embodiment, the apparatus 1000 is further configured to: Determine the type of the target object; Based on the type of the target object, obtain a target push template from a plurality of pre-configured push templates, and the target push template includes a target layout rule; Generate a target information report based on the target information and the target layout rule; Push the target information report to the target object.
[0082] As an optional embodiment, obtaining the initial information set includes: Based on a preset information acquisition rule, collect multiple initial information corresponding to the target industry through web crawler technology to form the initial information set; Among them, the information acquisition rule includes the acquisition frequency, time period, information source, and / or information type.
[0083] Embodiment III Figure 8 FIG. schematically shows a hardware architecture diagram of a computer device 10000 suitable for implementing the industry information processing method according to Embodiment 3 of the present application. In some embodiments, the computer device 10000 may be a terminal device such as a smart phone, a wearable device, a tablet computer, a personal computer, a vehicle-mounted terminal, a game console, a virtual device, a workbench, a digital assistant, a set-top box, a robot, etc. In other embodiments, the computer device 10000 may be a rack server, a blade server, a tower server, or a cabinet server (including an independent server or a server cluster composed of multiple servers), etc. As Figure 8 shown, the computer device 10000 includes, but is not limited to: a memory 10010, a processor 10020, and a network interface 10030 that can communicate with each other through a system bus. Among them: The memory 10010 includes at least one type of computer-readable storage medium. The readable storage medium includes flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 10010 may be an internal storage module of the computer device 10000, such as the hard disk or memory of the computer device 10000. In other embodiments, the memory 10010 may also be an external storage device of the computer device 10000, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device 10000. Of course, the memory 10010 may also include both the internal storage module and the external storage device of the computer device 10000. In this embodiment, the memory 10010 is generally used to store the operating system and various application software installed on the computer device 10000, such as the program code of the industry information processing method. In addition, the memory 10010 may also be used to temporarily store various data that have been output or will be output.
[0084] In some embodiments, the processor 10020 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other chips. The processor 10020 is generally used to control the overall operation of the computer device 10000, such as performing control and processing related to data interaction or communication with the computer device 10000. In this embodiment, the processor 10020 is used to run the program code stored in the memory 10010 or process data.
[0085] The network interface 10030 may include a wireless network interface or a wired network interface, which is generally used to establish a communication link between the computer device 10000 and other computer devices. For example, the network interface 10030 is used to connect the computer device 10000 to an external terminal through a network, and establish a data transmission channel and a communication link between the computer device 10000 and the external terminal. The network may be a wireless or wired network such as an enterprise intranet (Intranet), the Internet, the Global System of Mobile communication (GSM for short), Wideband Code Division Multiple Access (WCDMA for short), 4G network, 5G network, Bluetooth, Wi-Fi, etc.
[0086] It should be noted that Figure 8 only the computer device with components 10010 - 10030 is shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components may be implemented alternatively.
[0087] In this embodiment, the industry information processing method stored in the memory 10010 may also be divided into one or more program modules and executed by one or more processors (such as the processor 10020) to complete the embodiments of this application.
[0088] Embodiment 4 The embodiments of this application also provide a computer - readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the industry information processing method in the embodiment are implemented.
[0089] In this embodiment, the computer-readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the computer-readable storage medium may be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., equipped on the computer device. Of course, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the computer-readable storage medium is generally used to store the operating system installed on the computer device and various application software, such as the program code of the industry information processing method in the embodiment. In addition, the computer-readable storage medium can also be used to temporarily store various data that have been output or will be output.
[0090] Embodiment Five The embodiment of the present application further provides a computer program product, including a computer program, which when executed by a processor implements the method in the above embodiment.
[0091] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the embodiments of the present application can be implemented by a general computer device. They can be concentrated on a single computer device or distributed on a network composed of multiple computer devices. Optionally, they can be implemented by program codes executable by the computer device, so that they can be stored in a storage device and executed by the computer device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.
[0092] It should be noted that the above are only the preferred embodiments of the present application, and do not limit the patent protection scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A method for processing industry information, characterized in that: The method comprises: Acquire an initial information set, wherein the initial information set includes a plurality of initial information corresponding to a target industry; Based on the pre-designed prompt, the initial information set and the large language model, a target information set and target information information corresponding to the target information set are obtained; Among them, the target information set includes multiple target information, and the target information includes initial information whose matching degree with the target industry is greater than a preset threshold; the target information information includes: the information category of each target information, the summary information of each target information, and the summary information of multiple target information for the same information category; the pre-designed prompts are used to guide the large language model to perform industry matching, classification, summary and summary on the multiple initial information.
2. The method according to claim 1, characterized in that Based on the pre-designed prompt, the initial information set and the large language model, a target information set and target information corresponding to the target information set are obtained, including: Inputting the pre-designed first prompt and the initial information set into a first large language model, so as to determine the plurality of target information from the plurality of initial information by using the first large language model; Based on the multiple pieces of target information, obtaining the target information; Among them, the first prompt includes the industry background, industry theme and / or industry terminology of the target industry, and the first prompt is used to guide the first large language model: obtain the industry information of the target industry according to the industry background, industry theme and / or industry terminology of the target industry; semantically match the industry information with the multiple initial information, and determine the multiple target information from the multiple initial information.
3. The method according to claim 1, characterized in that Semantically matching the industry information with the multiple pieces of initial information, and determining the multiple pieces of target information from the multiple pieces of initial information, including: Perform semantic matching between the industry information and the multiple pieces of initial information to obtain multiple pieces of valid information and form a valid information set, wherein the semantic matching degree between the valid information and the industry information is greater than the preset threshold; Based on a preset historical information database, the valid information set is compared and deduplicated to obtain a deduplicated information set; The deduplicated information set is self-deduplicated to obtain the target information set, wherein the target information set includes the multiple non-repetitive target information.
4. The method according to claim 2, characterized in that: Based on the plurality of target information, obtaining the target information includes: Inputting the pre-designed second prompt and the plurality of target information into the second largest language model to determine the information category of each target information through the second largest language model; The second prompt includes a plurality of preset information categories and a definition of each information category, and the second prompt is used to guide the second large language model: according to the plurality of preset information categories and the definition of each information category, the plurality of target information are classified to obtain an information category of each target information.
5. The method according to claim 4, characterized in that Based on the plurality of target information, obtaining the target information also includes: Determine multiple target information corresponding to each information category; Inputting the pre-designed third prompt and the corresponding multiple target information into the third language model, so as to obtain summary information of the multiple target information of the same information category through the third language model; Among them, the third prompt includes summary rules and / or summary examples, and the third prompt is used to guide the third language model: according to the summary rules and / or summary examples, generate summary information for multiple target information of the same information category.
6. The method according to claim 2, characterized in that Based on the plurality of target information, obtaining the target information includes: Inputting the pre-designed fourth prompt and the plurality of target information into a fourth language model, so as to generate summary information of each target information through the fourth language model; The fourth prompt includes summary rules and / or summary examples, and the fourth prompt is used to guide the fourth language model: generating summary information of each target information according to the summary rules and / or summary examples.
7. The method according to any one of claims 1 to 6, characterized in that: Also includes: Determine the type of target object; Based on the type of the target object, a target push template is obtained from a plurality of pre-configured push templates, wherein the target push template includes a target layout rule; Generate a target information report based on the target information and the target typesetting rules; Push the target information report to the target object.
8. The method according to any one of claims 1 to 6, characterized in that: Get the initial information set, including: Based on the preset information acquisition rules, multiple pieces of initial information corresponding to the target industry are collected by web crawler technology to form the initial information set; The information acquisition rules include the frequency of collection, time period, information source and / or information type.
9. An information processing device, characterized in that: The device comprises: A first acquisition module is used to acquire an initial information set, wherein the initial information set includes a plurality of initial information corresponding to a target industry; A second acquisition module, configured to acquire a target information set and target information corresponding to the target information set based on a pre-designed prompt, the initial information set and a large language model; Among them, the target information set includes multiple target information, and the target information includes initial information whose matching degree with the target industry is greater than a preset threshold; the target information information includes: the information category of each target information, the summary information of each target information, and the summary information of multiple target information for the same information category; the pre-designed prompts are used to guide the large language model to perform industry matching, classification, summary and summary on the multiple initial information.
10. A computer device, characterized in that: include: at least one processor; and a memory communicatively coupled to the at least one processor; wherein: The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.
12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.