External information recommendation method and system based on electric power construction industry

CN122594481APending Publication Date: 2026-08-18四川电力设计咨询有限责任公司
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202610947727.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-29
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

(1)缺乏对电力建设行业业务场景的深度适配,无法精准关联企业内部各业务板块需求,推送资讯多存在泛化问题,有效信息占比低

Benefits of technology

通过构建面向电力建设行业的RAG知识库,实现外部资讯与内部业务数据的深度关联,解决了内外资讯孤立的问题;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122594481A_ABST
    Figure CN122594481A_ABST
Patent Text Reader

Abstract

The application provides an external information recommendation method and system based on the power construction industry, which can realize deep correlation, intelligent interpretation and accurate targeted push of external information and internal business, relates to the information technology field of the power construction industry, and comprises the following steps: establishing a multi-dimensional structured index to obtain an RAG knowledge base; establishing a multi-level label system; obtaining external information from an industry information source to obtain a standardized external information core content package; and constructing an internal and external information correlation set; combining the business logic of the power construction industry to generate a deep comment for the external information; automatically generating a corresponding first-level label, second-level label and third-level label, and attaching the generated label to the deep comment; according to the generated label, matching the corresponding business board and post personnel, and pushing the external information original text and the deep comment to the information port of the target user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information technology in the power construction industry, specifically to a method and system for recommending external information in the power construction industry. Background Technology

[0002] As a vital component of the national economy, the power construction industry encompasses multiple business segments, including thermal power, hydropower, wind power, photovoltaic power, transmission lines, and substation construction. It is characterized by complex business scenarios, high levels of specialization, and extremely high requirements for the timeliness, accuracy, and relevance of external information. The development of large language models and Retrieval Enhanced Generation (RAG) technologies has provided technical support for achieving accurate matching, deep integration, and intelligent delivery of information. However, currently, power construction companies largely rely on general information platforms, industry websites, or manual screening to obtain external information, which has many shortcomings.

[0003] In the fields of semantic understanding and intelligent matching, several related technical solutions have been developed. For example, Chinese patent CN114238568A discloses a method for acquiring teacher resources, which constructs a demand synonym tree and a resource synonym tree, and calculates the total matching degree based on the matching degree and preset weights to achieve resource recommendation; Chinese patent CN118468881A proposes a semantic retrieval method for automatically extracting keywords, which preprocesses and segments the text, calculates a comprehensive score by combining word frequency and inverse text frequency values, and expands synonyms to improve retrieval accuracy; Chinese patent CN119938902B discloses an automatic recommendation method for standard terms, which uses a knowledge graph and gated attention fusion mechanism to process user input and searches for candidate recommended terms through a graph sampling algorithm.

[0004] However, existing technologies still have the following technical shortcomings: (1) Lack of deep adaptation to the business scenarios of the power construction industry, unable to accurately connect the needs of various business segments within the enterprise, and the information pushed has a generalization problem with a low proportion of effective information.

[0005] (2) External information is isolated from internal resources such as project documents, technical standards, and business data. Business personnel need to manually cross-reference them, making it difficult to quickly uncover the impact of external information on internal business.

[0006] (3) Lack of in-depth interpretation and targeted push functions. Most of the external information is raw content without targeted comments that are combined with the company's own business. It is impossible to achieve accurate targeted push according to business division and job requirements, which makes it difficult to give full play to the value of information.

[0007] (4) The existing tag system is mostly general type and lacks a special tag system that fits the business characteristics of the power construction industry, which makes it impossible to achieve refined classification and push of information.

[0008] Therefore, there is an urgent need to provide a method and system for recommending external information based on the power construction industry, which can achieve deep correlation between external information and internal business, intelligent interpretation, and precise targeted delivery. Summary of the Invention

[0009] The technical problem to be solved by the present invention is to provide a method and system for recommending external information based on the power construction industry, which can achieve deep correlation, intelligent interpretation and precise targeted push of external information and internal business.

[0010] The technical solution adopted by this invention to solve its technical problem is: an external information recommendation method based on the power construction industry, comprising the following steps: Collect internal business data from power construction companies to construct a standardized internal data source; use an embedding model adapted to power construction industry terminology to convert the standardized internal data source into vectors, store them in a vector database, and establish a multi-dimensional structured index to obtain the RAG knowledge base; A multi-level tagging system is established, which includes first-level tags, second-level tags, and third-level tags. The first-level tags correspond to core business segments, the second-level tags are the business sub-scenarios under the corresponding first-level tags, and the third-level tags correspond to technology types, policy regions, equipment models, or job types. The tagging system is configured with tag adaptation rules for automatic semantic analysis. External information is obtained from industry information sources, and the external information is standardized in format, redundant information is removed, and core content is extracted to obtain a standardized external information core content package. The standardized external information core content package is input into the large language model, which generates semantic retrieval vectors and / or keyword retrieval lists. The retrieval interface of the RAG knowledge base is called to perform a retrieval based on the semantic retrieval vectors and / or keyword retrieval lists, obtain a preset number of internal business data that are highly relevant to the external information, and construct an association set between the external information and the internal business data. The association set of external information and internal business data is input into the large language model, and combined with the business logic of the power construction industry, an in-depth commentary on the external information is generated. The commentary includes at least the interpretation of the core points of the external information, the correlation analysis with internal business, the potential impact, and reference suggestions. Based on the tag adaptation rules, the external information and the in-depth comments are semantically analyzed by the large language model to automatically generate corresponding first-level tags, second-level tags and third-level tags, and the generated tags are attached to the in-depth comments; Based on the generated tags, the corresponding business segments and personnel are matched, and the original external information and the in-depth comments are pushed to the target user's information portal.

[0011] Furthermore, the collection of internal business data from power construction enterprises and the construction of standardized internal data sources include: The data is deduplicated, denoised, and format-standardized by combining rule matching and machine learning to obtain cleaned data. The cleaned data was segmented using a general word segmentation tool combined with a dictionary specific to the power construction industry to obtain word vector data after segmentation. An information extraction method based on a large language model is adopted, and key information fields are preset to transform the unstructured text in the word vector data after word segmentation into structured data containing the key information fields, thereby obtaining a standardized internal data source.

[0012] Furthermore, the multi-dimensional structured index includes: using business segments as primary indexes, using at least one of core technology points, related positions, and equipment models as secondary indexes, and establishing an index association table that records the correspondence between index items and data IDs and vector IDs.

[0013] Furthermore, the tag adaptation rules include: The rules for mapping tags to business segments are used to establish a unique mapping relationship between first-level tags and core business segments, and between second- and third-level tags and the business segments corresponding to their respective first-level tags. The mapping rules between tags and job scope are used to establish the correspondence between three-level tags and specific job types in order to determine the target job personnel for information push. The criteria for determining information attributes include the core keywords and semantic features preset for each tag, as well as the determination logic for calling the RAG knowledge base for calibration when using fuzzy tags.

[0014] Furthermore, the standardized external information core content package is input into the large language model, which generates semantic retrieval vectors and keyword retrieval lists. The retrieval interface of the RAG knowledge base is called to perform vector similarity retrieval based on the semantic retrieval vectors and exact matching retrieval based on the keyword retrieval lists. The two retrieval results are merged and sorted to obtain a preset number of internal business data that are highly relevant to the external information, thus constructing an association set between external information and internal business data.

[0015] Furthermore, before constructing the association set of external information and internal business data, the method further includes: using the large language model to perform semantic verification on the top N internal business data obtained, and using the large language model to determine whether each internal business data has business relevance to external information, and removing irrelevant data that is determined to have no business relevance.

[0016] Furthermore, the method for generating the in-depth comments includes: Input prompt words into the large language model to clarify the comment structure and professional requirements; Three types of comment templates are preset: policy, technology, and supply chain. The large language model selects the corresponding template based on the type of external information and generates in-depth comments by combining the correlation set of external information and internal business data. The large language model verifies the terminology accuracy and logical coherence of the generated in-depth reviews. If errors are found, it calls the internal technical standards in the RAG knowledge base to make corrections and outputs the corrected in-depth reviews.

[0017] Furthermore, when automatically generating the corresponding first-level, second-level, and third-level tags, if the output of the large language model cannot clearly match the fuzzy tag of the third-level tag, the RAG knowledge base is called to retrieve internal data associated with the external information, and the fuzzy tag is calibrated in combination with internal business attributes until an accurate third-level tag is obtained.

[0018] Furthermore, it also includes iterative optimization steps: recording user viewing, clicking, commenting and feedback data on pushed information, optimizing the parameters of the large language model, tag matching accuracy and the retrieval algorithm of the RAG knowledge base based on the data, and dynamically updating the tag system and internal knowledge base.

[0019] An external information recommendation system based on the power construction industry includes: The RAG knowledge base construction module is used to collect internal business data of power construction enterprises and build a standardized internal data source. The standardized internal data source is converted into vectors using an embedding model adapted to the terminology of the power construction industry, stored in a vector database, and a multi-dimensional structured index is established to obtain the RAG knowledge base. The tag system management module is used to establish a multi-level tag system, which includes first-level tags, second-level tags, and third-level tags. The first-level tags correspond to core business segments, the second-level tags are the business sub-scenarios under the corresponding first-level tags, and the third-level tags correspond to technology types, policy regions, equipment models, or job types. The tag system is configured with tag adaptation rules for automatic semantic analysis. The external information preprocessing module is used to obtain external information from industry information sources, and to perform format standardization, redundant information removal and core content extraction on the external information to obtain a standardized external information core content package. The hybrid retrieval module is used to input the standardized external information core content package into the large language model, and the large language model generates semantic retrieval vectors and / or keyword retrieval lists; it calls the retrieval interface of the RAG knowledge base, performs retrieval based on the semantic retrieval vectors and / or keyword retrieval lists, obtains a preset number of internal business data that are highly relevant to the external information, and constructs an association set between external information and internal business data; The in-depth commentary generation module is used to input the association set of external information and internal business data into the large language model, and combine it with the business logic of the power construction industry to generate in-depth comments on the external information. The comments include at least the interpretation of the core points of the external information, the correlation analysis with internal business, potential impact, and reference suggestions. An automatic tag generation module is used to perform semantic analysis on the external information and the in-depth comments based on the tag adaptation rules and the large language model, automatically generate corresponding first-level tags, second-level tags and third-level tags, and attach the generated tags to the in-depth comments; The precise push module is used to match the generated tags with the corresponding business segments and personnel, and push the original external information and the in-depth comments to the information portal of the target user.

[0020] The beneficial effects of the present invention are as follows: Compared with the prior art, the present invention has the following beneficial effects: By building a RAG knowledge base for the power construction industry, a deep connection between external information and internal business data is achieved, solving the problem of isolated internal and external information. By generating hybrid retrieval strategies (semantic retrieval vectors and / or keyword retrieval lists) through a large language model, the recall and precision of information retrieval are improved. By combining large language models with industry business logic to generate in-depth commentary, external information is transformed into knowledge that can directly serve business decisions. By establishing a multi-level tagging system and automatically generating tags, we have achieved refined classification and precise targeted delivery of information. Attached Figure Description

[0021] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0022] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0023] like Figure 1 As shown, the external information recommendation method based on the power construction industry of the present invention includes the following steps: S1. Collect internal business data of power construction enterprises and construct a standardized internal data source; use an embedding model adapted to the terminology of the power construction industry to convert the standardized internal data source into vectors, store them in a vector database, and establish a multi-dimensional structured index to obtain the RAG knowledge base.

[0024] S2. Establish a multi-level tag system, which includes first-level tags, second-level tags, and third-level tags. The first-level tags correspond to core business segments, the second-level tags are the business sub-scenarios under the corresponding first-level tags, and the third-level tags correspond to technology types, policy regions, equipment models, or job types. The tag system is configured with tag adaptation rules for automatic semantic analysis. S3. Obtain external information from industry information sources, and standardize the format of the external information, remove redundant information and extract the core content to obtain a standardized external information core content package. S4. Input the standardized external information core content package into the large language model, and generate semantic retrieval vectors and / or keyword retrieval lists from the large language model; call the retrieval interface of the RAG knowledge base, perform retrieval based on the semantic retrieval vectors and / or keyword retrieval lists, obtain a preset number of internal business data that are highly correlated with the external information, and construct an association set between the external information and the internal business data. S5. Input the association set of the external information and internal business data into the large language model, and combine it with the business logic of the power construction industry to generate an in-depth commentary on the external information. The commentary shall at least include the interpretation of the core points of the external information, the correlation analysis with the internal business, the potential impact, and reference suggestions. S6. Based on the tag adaptation rules, semantic analysis is performed on the external information and the in-depth comments using the large language model to automatically generate corresponding first-level tags, second-level tags and third-level tags, and the generated tags are attached to the in-depth comments. S7. Based on the generated tags, match the corresponding business segments and personnel, and push the original external information and the in-depth comments to the target user's information portal.

[0025] In this invention, internal business data includes, but is not limited to: project construction documents (such as construction plans, acceptance reports, progress reports, safety briefing documents, etc.), technical standards (industry specifications, internal technical regulations, equipment operation manuals, etc.), bidding documents (tender announcements, tender documents, notices of award, evaluation reports, etc.), supply chain data (equipment supplier information, equipment model parameters, purchase contracts, delivery schedules, equipment operation and maintenance records, etc.), historical business reviews (project debriefings, technical discussion comments, business approval opinions, etc.), job descriptions (responsibilities, job content, professional skill requirements, etc. for each business position), and business segment division documents (enterprise business segment definition standards, division of responsibilities for each segment, business process specifications, etc.). This data can originate from various core systems within the enterprise, such as OA systems, project management platforms, document servers, supply chain management systems, human resource management systems, and bidding management systems.

[0026] In some embodiments of the present invention, the embedding model used for vectorization transformation is constructed based on the BERT (Bidirectional Encoder Representations from Transformers) model; in other embodiments, other pre-trained language models such as RoBERTa and ERNIE can also be used as the basic architecture. To make the embedding model more suitable for power construction industry terminology, preferably, the embedding models in the above embodiments can be further fine-tuned using a corpus specific to the power construction industry for domain-adaptive learning.

[0027] This invention constructs a RAG knowledge base that integrates a terminology adaptation and embedding model for the power construction industry with a multi-dimensional structured index. By combining a multi-level tagging system and a hybrid retrieval strategy, it achieves deep semantic association between external information and internal business data, effectively solving the problem of information separation between internal and external sources in traditional solutions. Simultaneously, it generates in-depth comments containing core point interpretations, correlation analysis, and reference suggestions based on a large language model, and automatically completes tag generation and job matching according to tag adaptation rules. This transforms generalized external information into precise decision-making knowledge that can directly serve personnel in different business segments and positions, significantly improving the accuracy, professionalism, and business value of information delivery, overcoming the shortcomings of existing technologies such as information generalization, lack of in-depth interpretation, and inability to target specific users.

[0028] In this embodiment of the invention, step S1: collecting internal business data of power construction enterprises and constructing a standardized internal data source specifically includes: The data is deduplicated, denoised, and format-standardized by combining rule matching and machine learning to obtain the cleaned data. Specifically, for deduplication, a method based on data hash value comparison is used to eliminate completely duplicate data. For similar data (such as different versions of the same solution), the optimal version is retained through text similarity algorithms (such as cosine similarity). Denoising includes eliminating invalid data (such as blank documents, garbled texts, meaningless placeholder content), eliminating special symbols and irrelevant redundant characters through regular expression matching, and filling in missing fields (such as device models, project numbers) based on the statistical rules of the same type of data (such as filling in the device model through the device supplier association). Format standardization includes uniformly converting documents in different formats (PDF, Word, Excel, TXT) into TXT format encoded in UTF-8, and uniformly converting structured data (such as Excel tables) into JSON format to ensure unified data formats.

[0029] The cleaned data is segmented by using a general segmentation tool combined with a dictionary exclusive to the electric power construction industry to obtain segmented word vector data. Based on the jieba segmentation framework, the dictionary exclusive to the electric power construction industry is imported. This dictionary contains professional terms, device names, and business scenario vocabulary in the electric power construction industry to avoid incorrect splitting of professional terms. At the same time, a stop word list is set to eliminate meaningless words such as "de", "le", "he" to ensure the accuracy of segmentation. After segmentation, word vectors containing professional terms are generated, providing support for subsequent structured processing and vector transformation.

[0030] An information extraction method based on a large language model is used. Key information fields are preset, and the unstructured text in the segmented word vector data is converted into structured data containing the key information fields to obtain a standardized internal data source. The large language model uses ChatGLM-4, and through prompt engineering, the large language model is guided to extract the corresponding fields, converting the unstructured text into structured data containing the above key information fields. For semi-structured data (such as formatted Word documents), by parsing the document format (such as title levels, tables), key information is extracted and filled in the corresponding fields. Finally, a standardized internal data source with unified fields and standardized content is formed, ensuring that each internal data contains core attributes (such as business segments, technical points, associated positions) that can be used for association and matching.

[0031] In some embodiments, the large language model employs a pre-trained language model with LLaMA 3 or equivalent performance. Optionally, the large language model can be further fine-tuned using power construction industry corpora (including technical standards, construction documents, policy documents, etc.) to enhance its semantic understanding of power construction industry terminology and business logic. It is understood that the technical solution of this application primarily guides the large language model to perform specific tasks through cue word engineering; the selection and fine-tuning of the model do not constitute a limitation on the scope of protection of this invention.

[0032] The above approach employs a layered processing strategy: First, data cleaning is performed using a combination of rule matching and machine learning to achieve efficient deduplication, noise reduction, and format standardization; then, a specialized dictionary for the power construction industry guides word segmentation to ensure the semantic integrity of technical terms; finally, a large language model-driven information extraction method enables intelligent structuring of unstructured data. This strategy generates high-quality, interconnected, and standardized internal data sources, laying a solid data foundation for subsequent RAG knowledge base construction and accurate matching of external information with internal business needs.

[0033] In some embodiments, the multi-dimensional structured index employs different dimensional divisions or hierarchical structures. For example, it can use time as the index to classify internal data by project stage or document creation time; or it can use geographic dimension as the index to classify internal data by project location or policy application area; or it can use document type as the index to classify internal data by data type such as construction plans, technical standards, and bidding documents. Furthermore, combinations of the above multiple dimensions can be used to form hybrid indexes to meet the refined retrieval needs of different business scenarios.

[0034] In this invention, preferably, the multi-dimensional structured index includes: using business segments as the primary index, and at least one of core technology points, related positions, and equipment models as secondary indexes, and establishing an index association table that records the correspondence between index items and data IDs and vector IDs. Specifically, this includes: establishing a primary index: using "business segments" as the primary index, classifying internal data according to business segments such as thermal power construction, hydropower construction, and transmission line engineering, to facilitate quick data filtering by business segment; establishing a secondary index: using at least one of "core technology points," "related positions," and "equipment models" as secondary indexes, further associating internal data under the same business segment with specific technologies, positions, or equipment. For example, associating construction documents and technical standards related to "photovoltaic bracket installation technology" with the primary index "new energy construction," the secondary index "photovoltaic power station construction," and "photovoltaic construction positions"; and establishing an index association table: recording the correspondence between each index item and internal data IDs and vector IDs, forming a three-dimensional association system of "index-data-vector," ensuring that relevant data can be quickly located through the index during retrieval, thus improving retrieval efficiency.

[0035] The above approach achieves a balance between retrieval efficiency and semantic accuracy through a two-level tree-structured index that combines "coarse screening by business segments with fine matching of core attributes," and it also works well with the tag system and job recommendation mechanism.

[0036] In some embodiments, the tag adaptation rules may be implemented in the following ways, such as pure keyword matching, tag recommendation based on collaborative filtering, tag prediction based on classification models, tag reasoning based on knowledge graphs, manual annotation by users, or a combination of the above methods.

[0037] In this invention, the following method is preferred, and the tag adaptation rules specifically include: 1. The rules for mapping tags to business segments are used to establish a unique mapping relationship between primary tags and core business segments, and between secondary and tertiary tags and the corresponding business segments of their respective primary tags. For example, the tag "Thermal Power Plant Construction - Boiler Installation - Water-Cooled Wall Installation Technology" clearly corresponds to the "Thermal Power Plant Construction" business segment, and is also associated with the core technical specifications of this business segment, ensuring a precise match between tags and business segments.

[0038] 2. The mapping rules between tags and job scope are used to establish the correspondence between the three-level tags and specific job types to determine the target personnel for information push. For example, the tag "Photovoltaic Construction - Bracket Installation Technology" corresponds to positions such as photovoltaic construction technician, photovoltaic construction team leader, and photovoltaic technology manager; the tag "Supply Chain Management - Equipment Procurement" corresponds to positions such as procurement specialist, procurement supervisor, and supply chain manager, ensuring that tags can accurately match target business personnel.

[0039] 3. Information attribute determination criteria, used to determine the tags that external information should match through keyword matching and semantic similarity calculation, and to call the RAG knowledge base for calibration when using fuzzy tags. Specifically, this includes: Keyword Determination: Pre-defined core keywords are used for each tag. For example, the core keywords for the tag "bracket installation technology" are "bracket installation," "photovoltaic bracket," and "bracket construction process." If the information contains ≥2 core keywords that are semantically related, the tag is considered a match. The threshold for the number of core keywords can be adjusted according to the specific granularity of the tag. For first-level tags with a broad coverage, the threshold can be appropriately increased to more than 3 to avoid over-matching.

[0040] Semantic Feature Determination: A large language model trained on a corpus of the power construction industry is used to pre-define the semantic features of each tag. For example, the semantic features of "policy compliance - new energy subsidy policy" are "subsidy", "policy", "new energy", and "financial support". Semantic analysis is performed through the large language model. If the similarity between the core semantics of the information and the semantic features of the tag is greater than or equal to a preset value (e.g., ≥0.8), then the tag is determined to match.

[0041] Fuzzy label calibration: If the large language model generates fuzzy labels (e.g., it cannot clearly distinguish between "photovoltaic bracket installation" and "photovoltaic inverter commissioning"), the RAG knowledge base is called to retrieve internal data associated with the information (such as relevant construction documents and job information), and the labels are calibrated in combination with internal business attributes until accurate third-level labels are obtained, ensuring the accuracy of label generation.

[0042] The above method adopts a three-layer tag adaptation rule, which is superior to other alternatives such as pure keyword matching, collaborative filtering, classification models, knowledge graph reasoning or user manual annotation in terms of semantic understanding ability, business adaptability, job-level push support, fuzzy tag calibration, interpretability and cold start friendliness.

[0043] In this invention, external information is obtained from various authoritative sources through API interfaces and / or web crawlers. This information includes industry policies, technological developments, supply chain information, industry exhibition / seminar information, peer case studies, and technical journal articles (research on core technologies related to power construction), etc. These authoritative sources include, but are not limited to: authoritative industry websites, policy platforms, supply chain platforms, technical journals, industry forums, and new media platforms.

[0044] The web crawler is built on the Scrapy framework and configured with crawler rules specific to the power construction industry. It is used to crawl information from websites without API interfaces. The crawler parameter configuration includes: setting the crawling frequency to once a day, dynamically adjusting according to information update needs; setting the request interval and simulating browser access behavior to circumvent website anti-crawling mechanisms; and adhering to the target website's robots.txt protocol to ensure the legality and stability of information acquisition.

[0045] This invention employs a standardized process consistent with internal data preprocessing to standardize the format of the acquired external information, uniformly converting it into TXT plain text format, removing redundant metadata such as layout, font, and color, and ensuring seamless association between internal and external information.

[0046] Then, redundant information is removed and core content is extracted from the standardized external information. In one embodiment of the invention, a text summarization algorithm is used to remove redundant content unrelated to power construction business, retaining substantive information; then, the title, abstract, key technical points, relevant fields, publication time, and publication source are extracted through the large language model to form a standardized external information core content package. The key technical points include, but are not limited to, new construction technologies and equipment models; the relevant fields include, but are not limited to, new energy, thermal power, and transmission lines.

[0047] To improve retrieval accuracy, this invention preferably employs a hybrid retrieval strategy combining semantic retrieval vectors and keyword retrieval lists. Specifically: the standardized external information core content package is input into a large language model, which generates semantic retrieval vectors and keyword retrieval lists; the retrieval interface of the RAG knowledge base is called, performing vector similarity retrieval based on the semantic retrieval vectors and exact matching retrieval based on the keyword retrieval lists; the two retrieval results are merged and sorted to obtain a preset number of internal business data items with the highest relevance to the external information, thus constructing an association set between external information and internal business data. The preset number can be any integer greater than or equal to 1.

[0048] To further improve retrieval accuracy, before constructing the association set of external information and internal business data, a large language model can be used to perform semantic verification on the top N internal business data obtained. The large language model determines whether each internal business data has business relevance to external information and removes irrelevant data that is determined to have no business relevance.

[0049] In one embodiment of the present invention, specifically: semantic retrieval is based on the semantic retrieval vector, and the cosine similarity with the internal business data vector is calculated in the vector database to filter out internal data with a similarity of not less than 0.75; keyword retrieval is based on the keyword retrieval list, and the keyword index of the internal data is matched to filter out internal data with a matching degree of not less than 80%; then a weighted fusion algorithm is used to fuse the dual retrieval results, wherein the weight of the semantic retrieval result accounts for 60% and the weight of the keyword retrieval result accounts for 40%, and the comprehensive relevance of each internal business data is calculated to filter out the top 3 internal business data with the highest comprehensive relevance; then the large language model performs semantic verification on the filtered internal business data to confirm its business relevance to external information (for example, when the external information is "new technology for photovoltaic bracket installation", the internal business data should be technical standards or project cases related to photovoltaic construction, rather than thermal power construction content), and irrelevant data with insufficient relevance is eliminated; finally, a set of associations between external information and internal business data is formed, which includes the core content of external information and the original text and key information of its associated internal business data.

[0050] In some embodiments, the semantic retrieval similarity is set to be no less than 0.80, the keyword retrieval matching degree is set to be no less than 85%, the weight of semantic retrieval results is 50%, and the weight of keyword retrieval results is 50%. These parameters can be dynamically adjusted according to business scenario requirements, data characteristics, and retrieval accuracy requirements. Those skilled in the art can set specific values ​​through experiments or experience; variations of these values ​​all fall within the protection scope of this invention.

[0051] To ensure the professionalism and practicality of the generated in-depth reviews, this invention sets the following four aspects as the core criteria for review generation: Key takeaways from external information: This includes crucial information such as policy content, technical details, and supply chain dynamics. Internal and external information relevance: This includes the points of convergence and divergence between external information and internal business data (such as technical standards, project cases, and business processes); Business logic of the power construction industry: including the industry's inherent logic such as construction process, technical specifications, policy requirements, and job responsibilities; Internal business needs of enterprises: including pain points in project construction, technology upgrade needs, policy compliance needs, and other practical concerns of enterprises.

[0052] The generation of in-depth reviews is based on the above four aspects of information, ensuring that the review content not only aligns with the substantive content of external information but also responds to the user's actual business scenario.

[0053] In one embodiment of the present invention, the method for generating in-depth comments specifically includes: 1. Input preset prompts into the large language model to clarify the comment structure and professional requirements. The comment structure includes: interpretation of core points, internal and external correlation analysis, potential impact, and reference suggestions; the professional requirements include: incorporating power construction industry terminology and business logic.

[0054] 2. Three types of comment templates are preset: policy-related, technology-related, and supply chain-related. The technology-related template explicitly includes the following: core technology interpretation, comparison with internal technical standards, application feasibility analysis, and job suitability suggestions. The policy-related and supply chain-related templates have corresponding differentiated structures. The large language model selects the appropriate template based on the type of external information and generates in-depth comments based on the association set between the external information and internal business data.

[0055] 3. The large language model performs quality checks on the generated in-depth reviews, specifically including terminology accuracy checks and logical coherence checks. If terminology errors or logical problems are found, the internal technical standards in the RAG knowledge base are invoked for correction, and the corrected in-depth review is output.

[0056] Furthermore, in order to improve the accuracy of tag generation, when automatically generating the corresponding first-level tags, second-level tags, and third-level tags, if the output of the large language model cannot clearly match the fuzzy tag of the third-level tag, the RAG knowledge base is called to retrieve internal data associated with the external information, and the fuzzy tag is calibrated in combination with internal business attributes until an accurate third-level tag is obtained.

[0057] After tags are generated, they are matched with corresponding business segments and personnel. Specifically: information is pushed to the information section of the corresponding business segment based on the first-level tags; corresponding personnel are matched based on the third-level tags; and a user preference model is built by combining user historical behavior data to prioritize pushing information with a high degree of matching with user preferences.

[0058] Furthermore, to improve the accuracy of matching pushed information with user needs, this method also includes an iterative optimization step: recording user viewing, clicking, commenting and feedback data on pushed information, optimizing the parameters of the large language model, tag matching accuracy and the retrieval algorithm of the RAG knowledge base based on the data, and dynamically updating the tag system and internal knowledge base.

[0059] This invention also provides an external information recommendation system based on the power construction industry, comprising: The RAG knowledge base construction module is used to perform step S1; The tag system management module is used to execute step S2; An external information preprocessing module is used to perform step S3; A hybrid retrieval module is used to perform step S4; A deep comment generation module is used to perform step S5; An automatic label generation module is used to perform step S6; and, The precise push module is used to perform step S7.

Claims

1. A method for recommending external information based on the power construction industry, characterized in that: Includes the following steps: Collect internal business data from power construction companies to construct a standardized internal data source; use an embedding model adapted to power construction industry terminology to convert the standardized internal data source into vectors, store them in a vector database, and establish a multi-dimensional structured index to obtain the RAG knowledge base; A multi-level tagging system is established, which includes first-level tags, second-level tags, and third-level tags. The first-level tags correspond to core business segments, the second-level tags are the business sub-scenarios under the corresponding first-level tags, and the third-level tags correspond to technology types, policy regions, equipment models, or job types. The tagging system is configured with tag adaptation rules for automatic semantic analysis. External information is obtained from industry information sources, and the external information is standardized in format, redundant information is removed, and core content is extracted to obtain a standardized external information core content package. The standardized external information core content package is input into the large language model, which generates semantic retrieval vectors and / or keyword retrieval lists. Call the retrieval interface of the RAG knowledge base, perform a retrieval based on the semantic retrieval vector and / or keyword retrieval list, obtain a preset number of internal business data that are highly relevant to the external information, and construct an association set between the external information and the internal business data; The association set of external information and internal business data is input into the large language model, and combined with the business logic of the power construction industry, an in-depth commentary on the external information is generated. The commentary includes at least the interpretation of the core points of the external information, the correlation analysis with internal business, the potential impact, and reference suggestions. Based on the tag adaptation rules, the external information and the in-depth comments are semantically analyzed by the large language model to automatically generate corresponding first-level tags, second-level tags and third-level tags, and the generated tags are attached to the in-depth comments; Based on the generated tags, the corresponding business segments and personnel are matched, and the original external information and the in-depth comments are pushed to the target user's information portal.

2. The method for recommending external information based on the power construction industry as described in claim 1, characterized in that, The collection of internal business data from power construction companies and the construction of standardized internal data sources include: The data is deduplicated, denoised, and format-standardized by combining rule matching and machine learning to obtain cleaned data. The cleaned data was segmented using a general word segmentation tool combined with a dictionary specific to the power construction industry to obtain word vector data after segmentation. An information extraction method based on a large language model is adopted, and key information fields are preset to transform the unstructured text in the word vector data after word segmentation into structured data containing the key information fields, thereby obtaining a standardized internal data source.

3. The method for recommending external information based on the power construction industry as described in claim 1, characterized in that, The multi-dimensional structured index includes: a business segment as the first-level index, at least one of core technology points, related positions, and equipment models as the second-level index, and an index association table that records the correspondence between index items and data IDs and vector IDs.

4. The method for recommending external information based on the power construction industry as described in claim 1, characterized in that, The tag adaptation rules include: The rules for mapping tags to business segments are used to establish a unique mapping relationship between first-level tags and core business segments, and between second- and third-level tags and the business segments corresponding to their respective first-level tags. The mapping rules between tags and job scope are used to establish the correspondence between three-level tags and specific job types in order to determine the target job personnel for information push. The criteria for determining information attributes include the core keywords and semantic features preset for each tag, as well as the determination logic for calling the RAG knowledge base for calibration when using fuzzy tags.

5. The method for recommending external information based on the power construction industry as described in claim 1, characterized in that, The standardized external information core content package is input into a large language model, which generates semantic retrieval vectors and keyword retrieval lists. The retrieval interface of the RAG knowledge base is called to perform vector similarity retrieval based on the semantic retrieval vectors and exact matching retrieval based on the keyword retrieval lists. The two retrieval results are merged and sorted to obtain a preset number of internal business data that are highly relevant to the external information, and to construct an association set between external information and internal business data.

6. The method for recommending external information based on the power construction industry as described in claim 1 or 5, characterized in that, Before constructing the association set of external information and internal business data, the process also includes: using the large language model to perform semantic verification on the top N internal business data obtained, and using the large language model to determine whether each internal business data has business relevance to external information, and removing irrelevant data that is determined to have no business relevance.

7. The method for recommending external information based on the power construction industry as described in claim 1, characterized in that, The method for generating in-depth comments includes: Input prompt words into the large language model to clarify the comment structure and professional requirements; Three types of comment templates are preset: policy, technology, and supply chain. The large language model selects the corresponding template based on the type of external information and generates in-depth comments by combining the correlation set of external information and internal business data. The large language model verifies the terminology accuracy and logical coherence of the generated in-depth reviews. If errors are found, it calls the internal technical standards in the RAG knowledge base to make corrections and outputs the corrected in-depth reviews.

8. The method for recommending external information based on the power construction industry as described in claim 1, characterized in that, When automatically generating the corresponding first-level, second-level, and third-level tags, if the output of the large language model cannot clearly match the fuzzy tag of the third-level tag, the RAG knowledge base is called to retrieve internal data associated with the external information, and the fuzzy tag is calibrated in combination with internal business attributes until an accurate third-level tag is obtained.

9. The method for recommending external information based on the power construction industry as described in claim 1, characterized in that, It also includes iterative optimization steps: recording user viewing, clicking, commenting and feedback data on pushed information, optimizing the parameters of the large language model, tag matching accuracy and the retrieval algorithm of the RAG knowledge base based on the data, and dynamically updating the tag system and internal knowledge base.

10. An external information recommendation system based on the power construction industry, characterized in that: include: The RAG knowledge base building module is used to collect internal business data from power construction companies and build standardized internal data sources. The standardized internal data source is transformed into vectors using an embedding model adapted to the terminology of the power construction industry, stored in a vector database, and a multi-dimensional structured index is established to obtain the RAG knowledge base. The tag system management module is used to establish a multi-level tag system, which includes first-level tags, second-level tags, and third-level tags. The first-level tags correspond to core business segments, the second-level tags are the business sub-scenarios under the corresponding first-level tags, and the third-level tags correspond to technology types, policy regions, equipment models, or job types. The tag system is configured with tag adaptation rules for automatic semantic analysis. The external information preprocessing module is used to obtain external information from industry information sources, and to perform format standardization, redundant information removal and core content extraction on the external information to obtain a standardized external information core content package. The hybrid retrieval module is used to input the standardized external information core content package into the large language model, and the large language model generates semantic retrieval vectors and / or keyword retrieval lists. Call the retrieval interface of the RAG knowledge base, perform a retrieval based on the semantic retrieval vector and / or keyword retrieval list, obtain a preset number of internal business data that are highly relevant to the external information, and construct an association set between the external information and the internal business data; The in-depth commentary generation module is used to input the association set of external information and internal business data into the large language model, and combine it with the business logic of the power construction industry to generate in-depth comments on the external information. The comments include at least the interpretation of the core points of the external information, the correlation analysis with internal business, potential impact, and reference suggestions. An automatic tag generation module is used to perform semantic analysis on the external information and the in-depth comments based on the tag adaptation rules and the large language model, automatically generate corresponding first-level tags, second-level tags and third-level tags, and attach the generated tags to the in-depth comments; The precise push module is used to match the generated tags with the corresponding business segments and personnel, and push the original external information and the in-depth comments to the information portal of the target user.

Citation Information

Patent Citations

  • Teacher resource acquisition method and system and terminal equipment

    CN114238568A

  • Semantic retrieval method and system for automatically extracting keywords

    CN118468881A

  • A method for automatically recommending standard terms

    CN119938902B