Intelligent standard knowledge retrieval system based on AI large model

The intelligent normative knowledge retrieval system based on AI large models solves the problems of fragmentation, weak semantic understanding, and lagging updates in traditional normative knowledge acquisition methods. It achieves accurate positioning and efficient retrieval of normative knowledge, supports cross-industry expansion, and meets the diverse needs of enterprises.

CN120950673AInactive Publication Date: 2025-11-14JIANGSU INSPIRE INTERNET OF THINGS TECH CO LTD +1

Patent Information

Application Number
CN202511470337.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-15
Publication Date
2025-11-14
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional methods of acquiring normative knowledge suffer from problems such as knowledge fragmentation, weak semantic understanding, insufficient usability, and lagging updates, making it difficult to meet the industry's demand for accurate and efficient acquisition of normative knowledge.

Method used

The intelligent normative knowledge retrieval system based on AI big data models achieves accurate positioning and application assistance of normative knowledge through multi-source dynamic collection, semantic understanding and indexing, intelligent retrieval, knowledge enhancement and user interaction feedback modules.

Benefits of technology

It enables precise positioning and efficient retrieval of standardized knowledge, lowers the barrier to entry, supports cross-industry expansion, improves retrieval efficiency and accuracy, and meets the diverse needs of different enterprises.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950673A_ABST
    Figure CN120950673A_ABST
Patent Text Reader

Abstract

The invention relates to the field of artificial intelligence, in particular to an intelligent standard knowledge retrieval system based on an AI large model, which comprises a standard knowledge management database, a multi-source acquisition preprocessing module, a semantic understanding and indexing module, an intelligent retrieval engine module, a knowledge enhancement module, a user interaction feedback module, a security control module and a deployment expansion module. The standard knowledge management database is used for storing standard knowledge design data, real-time retrieval data and feedback data and constructing a dynamically updated standard knowledge resource pool; through the multi-source standard knowledge acquisition and preprocessing module, multi-channel objective standard data can be integrated and dynamically updated, and the problem of knowledge fragmentation is solved; and through the semantic understanding and indexing module, industry professional semantic accurate matching is realized, the limitation of traditional keyword retrieval is broken through, and the accuracy and efficiency of standard knowledge retrieval are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence, specifically to an intelligent standardized knowledge retrieval system based on a large AI model. Background Technology

[0002] In compliance management and business operations in industries such as construction, healthcare, and manufacturing, the demand for accurate and efficient acquisition of regulatory knowledge is becoming increasingly stringent. As the standardization process in these industries accelerates, the number of various national / industry regulations and internal corporate standards has surged and is being updated frequently. At the same time, cross-domain regulations are becoming increasingly interconnected. Practitioners need to quickly locate the appropriate regulatory knowledge to ensure business compliance and efficiency, but traditional methods of acquiring regulatory knowledge are no longer sufficient to meet these needs.

[0003] Traditional solutions mainly include three categories: manual review, simple keyword search tools, and general search engines. Manual review requires flipping through paper specifications or locally stored documents and comparing chapters and clauses one by one. Simple keyword search tools rely on basic databases and return results through precise keyword matching. General search engines crawl relevant content from the entire Internet and rely on algorithms to sort and display it.

[0004] However, traditional solutions have significant drawbacks: First, knowledge is fragmented, with multiple sources of regulations stored in a scattered manner, making cross-channel integration difficult and prone to missing related clauses. Second, semantic understanding is weak, and keyword matching cannot capture deeper needs; for example, "fire protection in high-rise residential buildings" is difficult to accurately match with corresponding national regulations. Third, usability is insufficient; search results only list the text of the clauses, lacking popular explanations and application examples, making it difficult for users to understand and apply them. Fourth, updates are lagging, making it difficult to synchronize regulations revisions and repeal information in a timely manner, easily leading to the misuse of outdated regulations. Therefore, developing an intelligent regulations knowledge retrieval system that can integrate regulations knowledge, perform accurate semantic retrieval, and assist in understanding and application is crucial. Summary of the Invention

[0005] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide an intelligent normative knowledge retrieval system based on an AI large model to address the problems raised in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution: an intelligent standardized knowledge retrieval system based on an AI large-scale model, comprising: Standardized Knowledge Management Database: Construct a standardized data architecture, create a standardized knowledge resource pool that supports dynamic updates, and store objective standardized knowledge data, real-time retrieval data, and feedback data.

[0007] Multi-source data acquisition and preprocessing module: Utilizing a multi-source dynamic adaptive acquisition mechanism, this module obtains objective data with clear sources and verifiability from official authoritative sources, internal company data, and dynamically updated channels. After standardization of format, normalization of terminology, and structured annotation of attributes, the data is transmitted to the management database.

[0008] Semantic Understanding and Indexing Module: Constructs a domain-adaptive AI model, achieving precise industry semantic matching through lightweight fine-tuning. It transforms preprocessed data into high-dimensional vectors containing specialized semantic features, builds three types of indexes—semantic, structured attribute, and hierarchical association—and establishes a mapping relationship with the database.

[0009] The intelligent search engine module analyzes user search intent, corrects input errors, and enables text and multimodal search based on indexes. It utilizes multi-factor ranking, deduplication and merging, and cross-normative association recommendations to output preliminary search results.

[0010] Knowledge Enhancement Module: Intelligently parses professional and standard texts, associates them with practical application cases, identifies standard conflicts, traces changes, and generates enhanced search results.

[0011] User interaction feedback module: Provides a conversational interface, supporting multi-turn dialogues and contextual association. Collects user feedback to optimize the indexing and iterative model.

[0012] Security management module: Builds a multi-role dynamic permission architecture, defines access and operation permissions, encrypts sensitive data, and combines operation logs to meet compliance auditing requirements.

[0013] Deployment extension module: Provides a multi-scenario deployment architecture, enabling cross-industry expansion and external system integration through industry-adaptive plugins and open interfaces.

[0014] The technical effects and advantages of this invention are as follows: 1. This invention integrates multi-channel data through a multi-source normative knowledge acquisition and preprocessing module, and combines it with a large model-driven semantic understanding and index building module. This not only solves the problems of fragmented normative knowledge and confusing terminology, but also breaks through the limitations of traditional keyword retrieval, achieving precise positioning of normative knowledge and greatly improving retrieval efficiency and accuracy.

[0015] 2. This invention transforms professional standards into easy-to-understand explanations and connects them with real-world cases through a knowledge enhancement and explanation generation module. It can also identify standard conflicts and changes, solving the pain point of "finding the standards but not knowing how to use them," lowering the threshold for using standard knowledge, and helping users to correctly apply the standards.

[0016] 3. This invention provides cloud and local dual deployment modes through deployment and adaptation of extended modules. Combined with industry plug-in design, it can be adapted across industries without refactoring the core module, meeting the needs of different enterprises and greatly reducing the cost of implementation and expansion. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of the core process.

[0019] Figure 2 This is a schematic diagram of the feedback closed-loop process.

[0020] Figure 3 The flowchart supports the core link framework. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.

[0022] Please see Figure 1 As shown, this invention provides an intelligent standardized knowledge retrieval system based on an AI large model, including a standardized knowledge management database, a multi-source acquisition and preprocessing module, a semantic understanding and indexing module, an intelligent retrieval engine module, a knowledge enhancement module, and a user interaction feedback module.

[0023] The multi-source acquisition and preprocessing module is connected to the standardized knowledge management database, which in turn is connected to the semantic understanding and indexing module. The semantic understanding and indexing module is connected to the intelligent retrieval engine module, which is connected to the knowledge enhancement module. Finally, the knowledge enhancement module is connected to the user interaction feedback module.

[0024] Please see Figure 2 As shown, Figure 2 This is a schematic diagram of the feedback closed-loop process, including a schematic diagram of the feedback closed-loop process, a standardized knowledge management database, and a semantic understanding and indexing module.

[0025] The user interaction feedback module feeds data to the standardized knowledge management database and provides optimization requirements to the semantic understanding and indexing module.

[0026] Please see Figure 3 As shown, Figure 3 The core link framework supports a flowchart, including a security control module, a deployment and extension module, a standardized knowledge management database, a semantic understanding and indexing module, an intelligent search engine module, and a user interaction feedback module.

[0027] The security control module manages the standardized knowledge management database, semantic understanding and indexing module, intelligent search engine module, and user interaction feedback module; the deployment extension module provides support for the standardized knowledge management database, semantic understanding and indexing module, and user interaction feedback module.

[0028] Standardized Knowledge Management Database: Construct a standardized data architecture, create a standardized knowledge resource pool that supports dynamic updates, and store objective standardized knowledge data, real-time retrieval data, and feedback data.

[0029] This embodiment employs a hybrid architecture combining relational and vector databases. The relational database stores structured data, while the vector database stores high-dimensional semantic vector data. Specific design data required includes, but is not limited to, industry-specific metadata, domain terminology databases, and index configuration parameters. Data generated in real-time during system operation includes, but is not limited to, user search data, feedback data, and specification update data.

[0030] Multi-source data acquisition and preprocessing module: Utilizing a multi-source dynamic adaptive acquisition mechanism, this module obtains objective data with clear sources and verifiability from official authoritative sources, internal company data, and dynamically updated channels. After format standardization, terminology normalization, and attribute structure annotation, the data is transmitted to the management database; this includes the following steps:

[0031] S1.1: Through source feature identification logic, automatically adapt to the objective data formats of different sources and collect the following data with clear sources and verifiability: official authoritative sources such as national / industry standard texts, legal provisions, official revision announcements and repeal notices; enterprise internal sources such as enterprise technical specification documents, internal process standards, and product compliance parameter tables; and dynamically updated sources such as official API push specification update information, industry association compliance guidelines, and publicly available specification implementation case data. S1.2: Based on the domain terminology association graph, the equivalence relationship between cross-document terms is determined through term semantic similarity calculation, thereby achieving unified term mapping. The term semantic similarity calculation follows the formula below: ; in, , These are the semantic vectors of the terms to be matched. These are the semantic vector weight coefficients. The length of the longest common subsequence of the two terms. , These are the character lengths of the two terms, respectively.

[0032] S1.3: Through the dual logic of source-end update signal capture and user feedback verification, the revision, repeal, and addition status of the collected objective standard data are automatically identified, the old version of the standard data is marked as invalid, and the standard knowledge management database is updated synchronously.

[0033] The hardware required for this embodiment includes a data acquisition server and a network adapter; the data acquisition channels specifically include the standard open download interface on the official website of the State Administration for Market Regulation, the standard subscription interface on the official website of international standards organizations, the technical specification upload path of the enterprise's internal office system, and the compliance guide push interface of industry associations; the multi-format data processing tools that support parsing include PDF text parsing libraries, document processing libraries, and table reading libraries. When standardizing the format, the standard document is uniformly converted into a chapter structure of "General Principles-Terminology-Technical Requirements-Appendices", and the chapter titles adopt the format of "First-level heading-Second-level heading".

[0034] Based on a domain terminology database, cross-document term alignment is achieved using a semantic similarity formula, which is: ; Selecting terms , ,set up ;in , The actual dimension is 384, but it is simplified to 3 dimensions here; the semantic similarity part is calculated according to the formula as follows: The similar parts of the text were calculated as follows: The final similarity is If the value is greater than 0.8, it is determined to be an equivalent term, and the association is realized.

[0035] Attribute annotations extract the effective date and applicable fields from the specification using regular expressions. Specification update monitoring adopts a combination of "scheduled crawling + API push". It crawls official update announcements every 2 hours and receives real-time revision notifications pushed by the official API. If a specification is detected as obsolete, it automatically marks the corresponding old version in the database as "invalid" and adds an obsolescence timestamp.

[0036] Semantic Understanding and Indexing Module: Constructs a domain-adaptive AI model, achieving precise industry-specific semantic matching through lightweight fine-tuning. It transforms preprocessed data into high-dimensional vectors containing specialized semantic features, builds three types of indexes: semantic, structured attribute, and hierarchical association, and establishes a mapping relationship with the database. This includes the following steps: S2.1: Input industry-specific corpus and use a lightweight fine-tuning mechanism with local parameter updates to adjust model parameters to achieve accurate understanding of industry-specific semantics; S2.2: Input the preprocessed objective normative knowledge data into the domain-adaptive AI model to generate a high-dimensional vector containing domain-specific semantic features. The generation process follows the formula below: ; in, To standardize the semantic vectors for domain-specific knowledge, The general semantic vector output by the basic model. This is a weighted matrix for industry sectors. Semantic vectors specific to the domain terminology; S2.3: Construct semantic vector index, structured attribute index, and hierarchical association index. The semantic vector index realizes the semantic association and positioning of standard knowledge through a high-dimensional vector storage architecture. The structured attribute index realizes the combined condition filtering based on the standard's effective date, applicable field, and priority. The hierarchical association index realizes context tracing based on the chapter-clause-subclause hierarchical relationship. And establish an association mapping between each index and the standard knowledge management database.

[0037] It should be noted that the indexes constructed in the semantic understanding and indexing module satisfy the following: the semantic vector index uses a high-dimensional vector storage and retrieval architecture to quickly calculate the semantic similarity of objective normative knowledge; the structured attribute index uses a multi-dimensional attribute mapping structure to quickly filter objective normative attributes; the hierarchical association index uses a parent-child node association logic to construct a tree structure to achieve hierarchical tracing of objective normative knowledge; and the association mapping between the indexes is achieved through a unified identifier.

[0038] This embodiment selects a basic large model that supports semantic understanding and uses LoRA lightweight fine-tuning technology (learning rate 5e-5, training epochs 3, batch size 8). The fine-tuning corpus consists of industry standard texts and explanations of industry professional terms, enabling the model to accurately understand industry professional semantics.

[0039] Taking the standard clause "the width of fire lanes should not be less than 4 meters" as an example: Basic Vector The weight matrix for the construction industry is as follows: Emphasize the weighting of dimensions related to "width" and "fire protection"; The field term "fire lane" Further calculations yielded: This vector is more in line with the semantics of the construction industry and is stored in a vector database.

[0040] Semantic vector generation employs a semantic encoding model, transforming standardized chapters, clauses, and terms into 384-dimensional semantic vectors. The semantic vector index uses a clustering index type from a vector database, with 1024 cluster centers and a retrieval parameter of 32 during querying. The structured attribute index is built on a relational database, setting "Effective Date," "Applicable Fields," and "Priority" as index fields, and supporting combined condition filtering. The hierarchical association index uses a tree structure, linking parent node IDs with child node IDs to achieve contextual tracing from clauses to chapters.

[0041] The intelligent search engine module analyzes user search intent, corrects input errors, and enables text and multimodal search based on indexes. Utilizing multi-factor ranking, deduplication and merging, and cross-normative association recommendations, it outputs preliminary search results, including the following steps: S3.1: Perform semantic analysis on user search input to identify the domain attributes, content scope, and accuracy requirements of the search needs; generate a candidate intent set for fuzzy input for user confirmation to complete the accurate correction of search intent; S3.2: If the user input contains image information, the image-text semantic fusion algorithm is used to convert the image information into a semantic feature vector, which is then fused and matched with the text retrieval vector. The fusion process follows the following formula: where is the multimodal fusion feature vector, is the image feature weight coefficient, is the image semantic feature vector, and is the text retrieval feature vector; S3.3: Through semantic association retrieval logic, it achieves accurate matching of keywords, natural language questions, normative identifiers and objective normative knowledge data to obtain text retrieval results.

[0042] Furthermore, the search result processing includes the following steps: S4.1: By weighting and fusing four core factors—semantic similarity, normative priority, timeliness, and user feedback—the comprehensive score of the search results is calculated and a ranking sequence is generated. The score calculation follows the formula below: ,in, To calculate the overall score for the search results, Semantic similarity Priority of standards Timeliness User feedback weighting coefficients and ; S4.2: Perform an objective and standardized semantic comparison of the sorted search results, identify duplicate or related standardized clauses, automatically merge duplicate results, and mark the differences in applicable scenarios of different clauses; S4.3: Based on the normative knowledge association graph, calculate the semantic relevance of the target clause to other objective normative clauses, recommend supplementary normative knowledge related to the target clause, and form preliminary search results.

[0043] The search interface in this embodiment is developed using a front-end development framework and supports four input methods: keywords, natural language questions, images, and standard numbers. Search intent parsing is based on a finely tuned AI model, using prompt word templates to extract the user's search domain, core needs, and standard type; for fuzzy input, a candidate intent set is generated for user confirmation.

[0044] This example addresses the user's search requirement of uploading a "fire hydrant installation site diagram" and inputting "whether it complies with specifications." Set image feature weight coefficients Generate image vectors The search query "fire hydrant installation specifications" generates a text search feature vector. ;Calculations show that: Text retrieval is achieved through "semantic vector matching + structured attribute filtering". For example, when retrieving a specification number, the corresponding specification is first located by matching the specification number with the structured index, and then the specific chapter is located by matching the semantic vector.

[0045] The comprehensive score of the search results is further calculated using the following formula: The weights are set as follows: Selecting a search result: the semantic similarity and the vector similarity between the user input "fire lane width" are... The standard priority is assigned as P=1 for national standards (priority: national > industry > enterprise); the timeliness is defined as effective in 2014 and revised in 2018, with "timeliness within 5 years after revision = 1.0, decreasing by 0.05 for each year beyond that", and T=0.9 in 2025; the historical usefulness rate of user feedback is 80%, assigned a value of 0.8; the overall score is calculated as 0.92; a score ≥0.8 is considered a high-quality result and will be displayed with priority.

[0046] The deduplication and merging results were based on a semantic similarity greater than 0.95, with differences in applicable scenarios noted after merging. Cross-normative association recommendation was based on normative knowledge association graphs, with recommendations made when the association was greater than 0.8, to help users obtain complete normative knowledge.

[0047] Knowledge Enhancement Module: Intelligently parses professional and standard texts, associates them with practical application cases, identifies standard conflicts, traces changes, and generates enhanced search results.

[0048] The intelligent parsing mechanism for regulatory clauses uses a domain-adaptive AI model to transform objective professional regulatory texts into accessible explanations, semantically annotates technical terms, and establishes associations with a domain terminology database, enabling rapid querying of term definitions. The scenario-case association algorithm matches the semantic features of cases with the features of regulatory clauses, filtering publicly available regulatory implementation cases (such as officially reported compliance cases) that fit the clauses. Case sources include publicly available industry scenario data and user-uploaded compliance cases. The regulatory conflict identification logic compares multiple objective regulatory contents to identify content conflicts between different regulations and marks conflict resolution principles. The regulatory change tracking mechanism records version change information for objective regulatory clauses, displaying the differences between old and new versions and their effective dates.

[0049] This embodiment's specification parsing is based on a finely tuned AI model. It uses prompt word templates to transform specification clauses into plain explanations, including technical requirements, operational points, and common errors. Semantic annotation of technical terms is performed and linked to a domain terminology database, enabling rapid querying of term definitions. The scenario case association algorithm matches the similarity between case semantic vectors and specification clause vectors (threshold > 0.75). The case database stores compliant and non-compliant cases (sourced from publicly available industry reports and internal company compliance records). Specification conflict identification compares the content of clauses from different specifications, identifies conflict points, and provides conflict resolution principles. Specification change tracking records version change information for specification clauses, displaying the differences between old and new versions and their effective dates.

[0050] User interaction feedback module: Provides a conversational interface, supporting multi-turn dialogues and contextual association. Collects user feedback and optimizes the indexing and iterative model, including the following steps: S5.1: Provides a conversational search interface, which retains the user's historical search records and interaction content through a contextual semantic memory mechanism, enabling coherent understanding of search intent in multi-turn dialogues; when the user asks follow-up questions, the search needs are accurately located based on the historical context; S5.2: Provides lightweight feedback options and in-depth feedback entry. Lightweight feedback allows users to rate search results as "useful / useless" and select reasons for invalidity, including: irrelevant results, unclear explanations, and outdated terms. In-depth feedback allows users to upload error correction suggestions and optimization requirements. S5.3: Transform user feedback data into model optimization parameters and index adjustment criteria, and update domain-adaptive AI model parameters and multi-dimensional indexes. The model parameter update follows the formula below: ; in, For the updated model parameters, These are the model parameters before the update. For learning rate, This is the gradient of the loss function based on the feedback data F.

[0051] The dialogue interface in this embodiment uses a real-time communication protocol for interaction. Contextual memory retains the content of the last 5 rounds of dialogue, allowing users to accurately locate their needs when asking follow-up questions based on historical context. User feedback is collected through front-end buttons. Lightweight feedback has "useful" and "useless" options; clicking "useless" requires selecting a reason. In-depth feedback provides a text input box, supporting users to upload error correction suggestions and new examples.

[0052] This embodiment optimizes the intent recognition parameters for the search term "fire hydrant installation": (Model parameters before update) Learning rate Based on 100 search results for "fire hydrant installation", calculations were performed. Feedback data shows intent recognition bias; a positive gradient indicates the need to increase parameters. The model parameter update formula is used to calculate... The updated parameters improve the accuracy of intent recognition for the search term "fire hydrant installation".

[0053] Feedback data processing adopts a "daily summary + weekly iteration" mechanism. Feedback data is summarized daily, and the model and index parameters are fine-tuned weekly using the feedback data. If experts confirm that the parsing is incorrect, the parsing content of the clauses is updated, and the incorrect cases are added to the fine-tuning corpus.

[0054] Security management module: Establishes a multi-role dynamic permission architecture, defines access and operation permissions, encrypts sensitive data, and combines operation logs to meet compliance auditing requirements, including: The multi-role dynamic permission architecture dynamically allocates access and operation permissions for objective and standardized knowledge based on the user's domain, responsibilities, and enterprise needs, thereby achieving fine-grained control of permissions. The data encryption mechanism employs dual protection logic of transmission encryption and storage encryption to encrypt sensitive, objective, and standardized knowledge, ensuring the security of data transmission and storage. The operation log traceability mechanism records users' standardized access, retrieval operations, feedback submissions, and other behavioral information, forming a traceable operation log to meet compliance audit requirements.

[0055] This embodiment employs a role-based access control model, defining four roles: administrator, domain expert, enterprise user, and general user. Administrators have system configuration and specification review permissions, and their accounts require two-factor authentication. Domain experts have feedback review and terminology database maintenance permissions, and can only access specifications within their assigned domain. Enterprise users can access internal and publicly available specifications, and can only modify specifications uploaded by their own company. General users can only access publicly available specifications and have no modification permissions. Regarding data security, secure transmission protocols are used for encryption during transmission, and advanced encryption standard algorithms are used for encryption during storage. Operation logs record user operation information, are stored using a log storage architecture, are retained for six months, support multi-dimensional queries, and undergo quarterly compliance checks against relevant laws and regulations.

[0056] Deployment Extension Module: Provides a multi-scenario deployment architecture, enabling cross-industry expansion and external system integration through industry-adaptive plugins and open interfaces, including: The multi-scenario deployment architecture provides two modes: cloud service deployment and local deployment. Cloud deployment supports multiple users to share access to objective and standardized knowledge, while local deployment supports the local storage and management of objective and standardized knowledge within the enterprise. The industry-adaptive plugin mechanism uses modular design to encapsulate industry-related model parameters, terminology libraries, and objective normative knowledge resources into independent plugins. When adding a new industry, only the corresponding plugin needs to be replaced to achieve system adaptation, without the need to refactor the core module. The open interface architecture provides standardized interfaces that support integration with existing enterprise systems. It can also embed a standard retrieval function entry point into external systems to enable rapid access to objective standard knowledge.

[0057] This embodiment's cloud deployment utilizes cloud server instances, providing web-based services and API services. Small and medium-sized enterprise (SME) users can log in via their accounts to conduct web-based searches or connect to their own systems through the API. Local deployment provides a deployment package, supporting deployment on internal enterprise servers. After deployment, enterprise specifications are stored in a local database, meeting data localization requirements. Industry-adaptive plugins adopt a modular design; each plugin includes industry-related model weights, a domain terminology library, and initialization data for the specification knowledge base. When adding a new industry, only the corresponding plugin needs to be uploaded and the initialization script executed. Open APIs adopt API documentation standards, supporting integration with existing enterprise systems and allowing embedding of specification search functionality into external systems.

[0058] Secondly: The accompanying drawings of the embodiments disclosed in this invention only involve the structures involved in the embodiments disclosed in this invention. Other structures can refer to the general design. In the absence of conflict, the same embodiment and different embodiments of this invention can be combined with each other. In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. An intelligent standardized knowledge retrieval system based on AI large-scale models, characterized in that, include: Standardized knowledge management database: Construct a standardized data architecture, create a standardized knowledge resource pool that supports dynamic updates, and store objective standardized knowledge data, real-time retrieval data, and feedback data; Multi-source data acquisition and preprocessing module: Utilizes a multi-source dynamic adaptive acquisition mechanism to obtain objective data with clear sources and verifiability from official authoritative sources, internal enterprise data, and dynamically updated channels; after format standardization, terminology normalization, and attribute structure annotation, it is transmitted to the management database; Semantic Understanding and Indexing Module: Constructs a domain-adaptive AI model, achieving precise industry semantic matching through lightweight fine-tuning; transforms preprocessed data into high-dimensional vectors containing professional semantic features, constructs three types of indexes: semantic, structured attributes, and hierarchical associations, and establishes an association mapping with the database; The intelligent search engine module analyzes user search intent, corrects input errors, and performs text and multimodal search based on indexes; it also outputs preliminary search results using multi-factor ranking, deduplication and merging, and cross-normative association recommendations. Knowledge Enhancement Module: Intelligently parses professional and standard texts, associates them with practical application cases, identifies standard conflicts, traces changes, and generates enhanced search results; User interaction feedback module: Provides a conversational interactive interface that supports multi-turn dialogue and contextual association; Collect user feedback to optimize the indexing and iteration model; Security management module: Build a multi-role dynamic permission architecture, define access and operation permissions, encrypt sensitive data, and combine operation logs to meet compliance audit requirements; Deployment extension module: Provides a multi-scenario deployment architecture, enabling cross-industry expansion and external system integration through industry-adaptive plugins and open interfaces.

2. The intelligent standardized knowledge retrieval system based on AI large model according to claim 1, characterized in that: The multi-source acquisition preprocessing module includes the following steps: S1.1: Through source feature identification logic, automatically adapt to the objective data formats of different sources and collect the following data with clear sources and verifiability: official authoritative sources such as national / industry standard texts, legal provisions, official revision announcements and repeal notices; enterprise internal sources such as enterprise technical specification documents, internal process standards, and product compliance parameter tables; and dynamically updated sources such as official API push specification update information, industry association compliance guidelines, and publicly available specification implementation case data. S1.2: Based on the domain terminology association graph, the equivalence relationship between cross-document terms is determined through term semantic similarity calculation, thereby achieving unified term mapping. The term semantic similarity calculation follows the formula below: ; in, , These are the semantic vectors of the terms to be matched. These are the semantic vector weight coefficients. The length of the longest common subsequence of the two terms. , These are the character lengths of the two terms, respectively. S1.3: Through the dual logic of source-end update signal capture and user feedback verification, the revision, repeal, and addition status of the collected objective standard data are automatically identified, the old version of the standard data is marked as invalid, and the standard knowledge management database is updated synchronously.

3. The intelligent standardized knowledge retrieval system based on AI large model according to claim 1, characterized in that: The semantic understanding and indexing module includes the following steps: S2.1: Input industry-specific corpus and use a lightweight fine-tuning mechanism with local parameter updates to adjust model parameters to achieve accurate understanding of industry-specific semantics; S2.2: Input the preprocessed objective normative knowledge data into the domain-adaptive AI model to generate a high-dimensional vector containing domain-specific semantic features. The generation process follows the formula below: ; in, To standardize the semantic vectors for domain-specific knowledge, The general semantic vector output by the basic model. This is a weighted matrix for industry sectors. Semantic vectors specific to the domain terminology; S2.3: Construct semantic vector index, structured attribute index, and hierarchical association index. The semantic vector index realizes the semantic association and positioning of standard knowledge through a high-dimensional vector storage architecture. The structured attribute index realizes the combined condition filtering based on the standard's effective date, applicable field, and priority. The hierarchical association index realizes context tracing based on the chapter-clause-subclause hierarchical relationship. And establish an association mapping between each index and the standard knowledge management database.

4. The intelligent standardized knowledge retrieval system based on AI large model according to claim 1, characterized in that: The intelligent search engine module includes the following steps for parsing and retrieving search input: S3.1: Perform semantic analysis on user search input to identify the domain attributes, content scope, and accuracy requirements of the search needs; generate a candidate intent set for fuzzy input for user confirmation to complete the accurate correction of search intent; S3.2: If the user input contains image information, an image-text semantic fusion algorithm is used to convert the image information into a semantic feature vector, which is then fused and matched with the text retrieval vector. The fusion process follows the formula below: ,in This is a multimodal fusion feature vector. These are the image feature weight coefficients. For image semantic feature vectors, This is a text retrieval feature vector; S3.3: Through semantic association retrieval logic, it achieves accurate matching of keywords, natural language questions, normative identifiers and objective normative knowledge data to obtain text retrieval results.

5. The intelligent standardized knowledge retrieval system based on AI large model according to claim 1, characterized in that: The intelligent search engine module processes search results including the following steps: S4.1: By weighting and fusing four core factors—semantic similarity, normative priority, timeliness, and user feedback—the comprehensive score of the search results is calculated and a ranking sequence is generated. The score calculation follows the formula below: ,in, To calculate the overall score for the search results, Semantic similarity Priority of standards Timeliness User feedback weighting coefficients and ; S4.2: Perform an objective and standardized semantic comparison of the sorted search results, identify duplicate or related standardized clauses, automatically merge duplicate results, and mark the differences in applicable scenarios of different clauses; S4.3: Based on the normative knowledge association graph, calculate the semantic relevance of the target clause to other objective normative clauses, recommend supplementary normative knowledge related to the target clause, and form preliminary search results.

6. The intelligent standardized knowledge retrieval system based on AI large model according to claim 1, characterized in that: The knowledge enhancement module includes: a smart parsing mechanism for normative clauses that uses a domain-adapted AI model to transform objective professional normative texts into easy-to-understand explanations, semantically annotates professional terms and establishes associations with a domain terminology database to enable rapid querying of terminology definitions; a scenario case association algorithm that uses the matching calculation of case semantic features and normative clause features to filter publicly available normative implementation cases that are compatible with the clauses, with case sources including publicly available industry scenario data and user-uploaded compliance cases; a normative conflict identification logic that compares multiple objective normative contents to identify content conflict points between different norms and annotates conflict resolution principles; and a normative change tracking mechanism that records version change information of objective normative clauses, displaying the differences between the old and new versions and their effective dates.

7. The intelligent standardized knowledge retrieval system based on AI large model according to claim 1, characterized in that: The user interaction feedback module includes the following steps: S5.1: Provides a conversational search interface, which retains the user's historical search records and interaction content through a contextual semantic memory mechanism, enabling coherent understanding of search intent in multi-turn dialogues; when the user asks follow-up questions, the search needs are accurately located based on the historical context; S5.2: Provides lightweight feedback options and in-depth feedback entry. Lightweight feedback allows users to rate search results as "useful / useless" and select reasons for invalidity, including: irrelevant results, unclear explanations, and outdated terms. In-depth feedback allows users to upload error correction suggestions and optimization requirements. S5.3: Transform user feedback data into model optimization parameters and index adjustment criteria, and update domain-adaptive AI model parameters and multi-dimensional indexes. The model parameter update follows the formula below: ; in, For the updated model parameters, These are the model parameters before the update. For learning rate, This is the gradient of the loss function based on the feedback data F.

8. The intelligent standardized knowledge retrieval system based on AI large model according to claim 1, characterized in that: The security management module includes: The multi-role dynamic permission architecture dynamically allocates access and operation permissions for objective and standardized knowledge based on the user's domain, responsibilities, and enterprise needs. The data encryption mechanism employs dual protection logic of transmission encryption and storage encryption to encrypt sensitive, objective, and normative knowledge. The operation log traceability mechanism records users' standardized access, retrieval operations, feedback submissions, and other behavioral information, forming a traceable operation log.

9. The intelligent standardized knowledge retrieval system based on AI large model according to claim 1, characterized in that: The deployment extension module includes: The multi-scenario deployment architecture provides two modes: cloud service deployment and local deployment. Cloud deployment supports multiple users to share access to objective and standardized knowledge, while local deployment supports the local storage and management of objective and standardized knowledge within the enterprise. The industry-adaptive plugin mechanism uses modular design to encapsulate industry-related model parameters, terminology libraries, and objective normative knowledge resources into independent plugins. When adding a new industry, only the corresponding plugin needs to be replaced to achieve system adaptation, without the need to refactor the core module. The open interface architecture provides standardized interfaces that support integration with existing enterprise systems and allow for the embedding of standardized search function entry points into external systems.

Citation Information

Patent Citations

  • Power data security policy large model question-answering system and method based on relation pooling

    CN119646160A

  • Large language model knowledge base question answering system based on multi-path fusion recall retrieval algorithm

    CN120144773A

  • Dynamic knowledge retrieval enhancement method based on large language model

    CN120407570A

  • Knowledge question and answer rapid processing method and system based on artificial intelligence

    CN120596639A

  • Multi-modal heterogeneous knowledge fusion construction and semantic enhancement retrieval system based on large model

    CN120705362A

Cited By

  • Large model question answering system and method for business hall

    CN121413776A

  • Demand specification generation method and system based on multi-modal understanding

    CN121525656A

  • A method and system for generating requirements specifications based on multimodal understanding

    CN121525656B

  • AI auxiliary interaction and virtual social contact system and method based on local area network private cloud

    CN121615734A

  • Intelligent real estate management system development method based on large model

    CN121903802A