A method and system for constructing and demand mining of an industry knowledge graph based on human-computer collaboration

By using a human-machine collaborative approach, multi-source text data is acquired, and industry knowledge graphs are constructed and optimized using natural language processing and graph embedding technologies. Combined with expert review and user profiling, the shortcomings of industry knowledge graphs in multi-source data fusion and tacit knowledge mining are solved, enabling precise demand mining and decision support.

CN122432151APending Publication Date: 2026-07-21GUANGZHOU YUNZHIDACHUANG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGZHOU YUNZHIDACHUANG TECH CO LTD
Filing Date
2026-04-30
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing industry knowledge graph construction technologies do not fully integrate multi-source data processing, making it difficult to guarantee the accuracy of knowledge graphs. Furthermore, graph embedding technology cannot effectively identify implicit knowledge and lacks deep mining capabilities.

Method used

By using a human-computer collaboration approach, multi-source text data is acquired, a preliminary knowledge graph is constructed using natural language processing technology, potential entity relationships are discovered by combining graph embedding technology, and domain experts review and correct the data to optimize the knowledge graph. Finally, user profiles are used to mine user needs.

Benefits of technology

It enhances the professionalism and accuracy of knowledge graphs, enables precise mining of industry needs, and provides reliable data support for industry decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122432151A_ABST
    Figure CN122432151A_ABST
Patent Text Reader

Abstract

The application discloses a kind of industry knowledge graph construction and demand mining method and system based on man-machine cooperation.The method comprises the following steps: obtaining the text data of industry literature, report, social media and forum and carrying out cleaning pretreatment;Entity recognition and relationship extraction are carried out using natural language processing technology, and a preliminary industry knowledge graph is constructed;The entities and relationships of the graph are mapped into low-dimensional vectors by graph embedding technology, and the similarity of the vectors is calculated to determine the potential correlation between entities;The graph and potential correlation are manually reviewed and corrected in combination with expert knowledge to obtain an optimized knowledge graph;Combine user portrait, and based on the optimized knowledge graph, demand mining is carried out and demand information is output.The application improves the accuracy of the knowledge graph through multi-source text fusion, graph embedding mining and expert verification closed loop, and realizes the accurate mining of industry demand.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, specifically to a method and system for constructing industry knowledge graphs and mining needs based on human-machine collaboration. Background Technology

[0002] In the process of deepening digital transformation, industry knowledge graphs serve as the core technological foundation supporting intelligent decision-making and accurate mining of user needs. The quality of their construction directly affects the effectiveness of application scenarios such as business intelligence and market analysis.

[0003] However, existing industry knowledge graph construction technologies face multiple challenges in practical applications. In the multi-source data processing stage, general methods rely excessively on modality diversity strategies, incorporating non-textual data such as images and sensor data into the fusion process. This results in insufficient multi-source data fusion and makes it difficult to guarantee the accuracy of the knowledge graph.

[0004] In knowledge representation and association mining, graph embedding technology is often used as a standalone tool, only achieving basic low-dimensional vector mapping. It cannot effectively identify implicit knowledge such as technological evolution paths, product iteration patterns, or potential demand relationships, causing the graph to remain at the level of explicit relationships and lacking in-depth mining capabilities. These problems collectively limit the effectiveness of industry knowledge graphs in demand mining. Summary of the Invention

[0005] The main objective of this application is to provide a method and system for constructing industry knowledge graphs and mining requirements based on human-machine collaboration, aiming to solve at least one of the above-mentioned technical problems.

[0006] The first aspect of this application provides a method for constructing industry knowledge graphs and mining requirements based on human-computer collaboration, including the following steps: S1. Obtain text data from industry literature, industry reports, social media and industry forums, and clean and preprocess the text data; S2. Utilize natural language processing technology to perform entity recognition and relation extraction on the cleaned and preprocessed text data in order to construct a preliminary industry knowledge graph. S3. Using graph embedding technology, the entities and relationships in the preliminary industry knowledge graph are mapped into low-dimensional vectors, and the similarity between entity vectors is calculated. Based on the similarity, potential inter-entity associations are determined. S4. The preliminary industry knowledge graph and the potential inter-entity relationships are manually reviewed by experts, and the knowledge graph is corrected according to the review results to obtain an optimized industry knowledge graph. S5. Combining user profiles, perform demand mining based on the optimized industry knowledge graph, and output demand information.

[0007] In some embodiments of this application, step S1 includes cleaning and preprocessing the text data by removing noise and duplicate information.

[0008] In some embodiments of this application, in step S1, the text data includes structured text data and unstructured text data.

[0009] In some embodiments of this application, step S2, which involves using natural language processing technology to perform entity recognition and relation extraction, includes: extracting entities using a named entity recognition model and determining the relationships between entities using a relation classification model or a joint extraction model.

[0010] In some embodiments of this application, step S3, which involves discovering potential inter-entity associations based on the similarity, includes: setting a similarity threshold; comparing the similarity between entity vectors with the similarity threshold; and identifying entity pairs with similarity higher than the threshold and without direct relationship edges in the preliminary industry knowledge graph as potential inter-entity associations.

[0011] In some embodiments of this application, step S4, which involves manual review incorporating expert knowledge, includes: having domain experts review the accuracy of the entities and relationships extracted from the preliminary industry knowledge graph, and review the rationality of the potential inter-entity associations determined in step S3.

[0012] In some embodiments of this application, step S4, which involves modifying the knowledge graph based on the review results, includes adding, deleting, or modifying entities, relationships, or associations in the knowledge graph based on expert review opinions.

[0013] In some embodiments of this application, step S5, which involves combining the user profile with the optimized industry knowledge graph to perform demand mining, includes: parsing the user profile to locate related entities in the optimized industry knowledge graph; and using the related entities as starting points, performing association queries or path exploration in the optimized industry knowledge graph to mine demand information related to the user profile.

[0014] In some embodiments of this application, after step S4, step S4a is also included: storing the optimized industry knowledge graph in a knowledge base and providing query services for demand mining in step S5.

[0015] Another aspect of this application provides an industry knowledge graph construction and demand mining system based on human-machine collaboration, including: The data acquisition and processing module is used to acquire text data from industry literature, industry reports, social media and industry forums, and to clean and preprocess the text data. The knowledge graph construction module is used to perform entity recognition and relation extraction on the cleaned and pre-processed text data using natural language processing technology in order to construct a preliminary industry knowledge graph. The association discovery module is used to map entities and relationships in the preliminary industry knowledge graph into low-dimensional vectors using graph embedding technology, calculate the similarity between entity vectors, and discover potential associations between entities based on the similarity. The manual verification and correction module provides a human-machine collaborative verification interface, receives expert knowledge for manual review of the preliminary industry knowledge graph and the potential inter-entity relationships, and corrects the knowledge graph based on the review results to obtain an optimized industry knowledge graph. The demand mining module is used to combine user profiles with the optimized industry knowledge graph to mine demand information and output demand information.

[0016] The embodiments of this application include at least the following beneficial effects: This application constructs a preliminary industry knowledge graph through multi-source text data fusion and natural language processing technology, and discovers implicit entity relationships using graph embedding technology. Furthermore, the introduction of expert knowledge for manual review and correction effectively improves the professionalism and accuracy of the knowledge graph, overcoming the problems of insufficient multi-source data fusion, inadequate implicit knowledge mining, and difficulty in guaranteeing the accuracy of knowledge graphs in traditional methods. Therefore, based on the optimized industry knowledge graph and user profiles, it is possible to accurately mine industry needs and provide data support for industry decision-making.

[0017] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0018] To more clearly illustrate the technical solutions of the embodiments of this application, the relevant drawings of the embodiments of this application are described below. It should be understood that the drawings described below are only for the convenience of clearly describing some embodiments of the technical solutions of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating the steps of an industry knowledge graph construction and demand mining method based on human-machine collaboration provided in an embodiment of this application. Figure 2 This is a schematic diagram of a module of an industry knowledge graph construction and demand mining system based on human-machine collaboration provided in an embodiment of this application. Detailed Implementation

[0020] The technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0021] To facilitate understanding of this invention, the following explanations are provided for key terms: Industry knowledge graph: An industry knowledge graph is a knowledge base that organizes and stores industry-specific knowledge in a graph structure, where nodes represent entities within the industry and edges represent relationships between entities. Its function is to structure and connect discrete industry information, facilitating machine understanding and reasoning.

[0022] Human-machine collaboration: Human-machine collaboration refers to a working model in which human experts and computer systems cooperate and complement each other's strengths in a specific task. In this method, the computer system is responsible for automatically processing large amounts of data and initially constructing knowledge, while human experts use their domain knowledge to review, correct, and make decisions to improve the quality and reliability of the final result.

[0023] Demand mining: Demand mining refers to the process of identifying, extracting, and analyzing potential user or market demands from massive amounts of data. Its goal is to discover problems, preferences, or expectations that users haven't explicitly expressed but that actually exist, providing a basis for product development, service optimization, and market strategy formulation.

[0024] Graph embedding technology is a method that maps nodes and edges in a graph structure to a low-dimensional continuous vector space. Through this mapping, the structural and semantic information of the graph is encoded into vectors, making it easy to perform operations such as similarity calculation, clustering, and classification in the vector space, thereby discovering hidden patterns and associations in the graph.

[0025] Natural Language Processing (NLP) is a branch of artificial intelligence that focuses on enabling computers to understand, interpret, generate, and process human natural language. In this approach, it is primarily used to automatically identify entities and extract relationships between them from textual data.

[0026] User personas: A user persona is a user model formed by abstracting and summarizing information such as the characteristics, behaviors, and preferences of a target user. By integrating multi-dimensional data, it depicts a three-dimensional image of the user, enabling the system to understand user needs and provide personalized services.

[0027] The first aspect of this application provides a method for constructing industry knowledge graphs and mining requirements based on human-computer collaboration. See also... Figure 1 The method includes the following steps: S1. Obtain text data from industry literature, industry reports, social media and industry forums, and clean and preprocess the text data; S2. Utilize natural language processing technology to perform entity recognition and relation extraction on the cleaned and preprocessed text data in order to construct a preliminary industry knowledge graph. S3. Using graph embedding technology, the entities and relationships in the preliminary industry knowledge graph are mapped into low-dimensional vectors, and the similarity between entity vectors is calculated. Based on this similarity, potential inter-entity associations are determined. S4. Combine expert knowledge to conduct manual review of the preliminary industry knowledge graph and the potential inter-entity relationships, and revise the knowledge graph according to the review results to obtain the optimized industry knowledge graph. S5. Based on the optimized industry knowledge graph, combine user profiles to mine demand information and output demand information.

[0028] The specific implementation method is as follows: In step S1, text data from industry literature, industry reports, social media, and industry forums is first acquired, and then cleaned and preprocessed. Text data can be acquired in various ways; for example, web crawling techniques can be used to scrape relevant text content from public websites, or data can be obtained from partner platforms via API interfaces. Cleaning and preprocessing may include removing non-text content such as HTML tags, special symbols, and emoticons from the text, as well as performing basic linguistic processing such as word segmentation and part-of-speech tagging.

[0029] In step S2, natural language processing (NLP) techniques are used to perform entity recognition and relation extraction on the cleaned and preprocessed text data to construct a preliminary industry knowledge graph. Entity recognition can employ rule-based methods, such as predefined keyword lists or regular expressions, to identify industry-specific entities. Relation extraction can utilize syntactic analysis-based methods, identifying predicate relationships between entities by analyzing sentence structure. Thus, a preliminary knowledge graph skeleton containing entities and relations can be constructed.

[0030] In step S3, the entities and relationships in the preliminary industry knowledge graph are mapped into low-dimensional vectors using graph embedding technology, and the similarity between entity vectors is calculated. Based on this similarity, potential inter-entity associations are determined. Graph embedding technology can represent each entity and relationship in the graph as a numerical vector. For example, entity sequences can be generated using a random walk method, and then these sequences can be trained using a word vector model to obtain low-dimensional vector representations of the entities. The similarity between entity vectors can be measured by calculating the cosine similarity of the vectors. When two entity vectors have a high similarity, they can be considered to have a potential association even if there is no direct connection in the preliminary knowledge graph.

[0031] In step S4, the preliminary industry knowledge graph and the potential inter-entity relationships are manually reviewed using expert knowledge. Based on the review results, the knowledge graph is revised to obtain an optimized industry knowledge graph. The manual review can be conducted through an interactive interface where domain experts review each entity and relationship automatically extracted by the system to determine its accuracy. Simultaneously, experts can also review the potential inter-entity relationships discovered in step S3 to determine whether they conform to industry logic and actual conditions. Based on the expert review comments, new entities or relationships can be manually added, incorrect entities or relationships can be deleted, or the attributes of existing entities or relationships can be modified.

[0032] In step S5, based on the optimized industry knowledge graph and combined with the user profile, demand mining is performed, and demand information is output. The user profile can include the user's basic information, browsing history, purchasing behavior, etc. During demand mining, relevant entities can first be searched in the optimized industry knowledge graph based on keywords or tags in the user profile. For example, if the user profile shows that the user is interested in "new energy vehicles," the entity "new energy vehicles" and its related entities can be located in the knowledge graph. Then, by performing simple neighbor queries or multi-hop path queries in the knowledge graph, potential demand points related to the user's interests can be discovered, such as "new energy vehicle battery technology" and "charging pile construction." Finally, this mined demand information can be presented in the form of a text list.

[0033] This application constructs a preliminary industry knowledge graph through multi-source text data fusion and natural language processing techniques, and utilizes graph embedding technology to discover implicit entity relationships. Furthermore, expert knowledge is introduced for manual review and correction, effectively improving the professionalism and accuracy of the knowledge graph and overcoming the problems of insufficient multi-source data fusion, inadequate implicit knowledge mining, and difficulty in guaranteeing the accuracy of knowledge graphs in traditional methods. Therefore, based on the optimized industry knowledge graph and user profiles, it is possible to accurately mine industry needs and provide data support for industry decision-making.

[0034] In some embodiments of this application, the cleaning and preprocessing in step S1 includes: removing noise and duplicate information from the text data.

[0035] Specifically, text cleaning and preprocessing are crucial initial steps in natural language processing tasks, aiming to improve data quality and make it more suitable for subsequent analysis and model training. Its main purpose is to eliminate inconsistencies, errors, noise, and redundancy in the data, thereby ensuring the accuracy and efficiency of subsequent steps (such as entity recognition and relation extraction). This typically includes, but is not limited to, operations such as removing special characters, HTML tags, stop words, punctuation standardization, case conversion, lemmatization, or stemming. These operations help transform raw, heterogeneous text data into a unified, standardized format.

[0036] Noise removal from text data refers to identifying and removing data fragments that are irrelevant to the analysis objective, interfere with data quality, or introduce errors. In industry knowledge graph construction, noise may include advertising content, irrelevant comments, garbled text, generic terms outside the industry, grammatical errors, or spelling errors. Noise removal ensures that subsequent processing focuses on valuable semantic content. This can be achieved through various techniques, such as rule-based filtering, which uses regular expressions or keyword lists to identify and remove advertisements, garbled text, and non-textual content with specific formats; or statistical methods, using statistical indicators such as term frequency-inverse document frequency (TF-IDF) to identify and filter out overly common or overly rare meaningless words; or machine learning methods, training classification models to distinguish between valid and noisy information, for example, classifying text as "industry-related" or "noise."

[0037] Meanwhile, removing duplicate information from text data refers to identifying and eliminating identical or highly similar text content that appears multiple times in the dataset. Duplicate information is very common when acquiring data from multiple sources; for example, the same report may be reprinted by multiple media outlets, or the same comment may be published multiple times. Retaining duplicate information not only increases storage and computational costs but also leads to redundancy in entities and relationships within the knowledge graph, affecting the accuracy and authority of the knowledge. This can be achieved through precise matching, hashing the text (e.g., MD5, SHA-256), and quickly identifying identical text by comparing hash values; or by using fuzzy matching / similarity calculation. For texts that are not identical but semantically highly similar, metrics such as Jaccard similarity, cosine similarity, and edit distance can be used to measure the similarity between texts. When the similarity exceeds a preset threshold, it is considered duplicate information and deduplicated; alternatively, fingerprinting technology can be used to extract the characteristic fingerprints of the text (e.g., MinHash), and comparing these fingerprints can efficiently detect near-duplicates.

[0038] Through the above technical solution, the acquired text data is cleaned and preprocessed in step S1. Specifically, removing noise and duplicate information significantly improves the quality of the original data. Removing noise ensures that subsequent entity recognition and relation extraction processes can focus on meaningful industry content, avoiding interference from irrelevant information in model training and knowledge extraction, thereby improving the accuracy of the extraction results. Simultaneously, removing duplicate information avoids redundant entities and relations in the knowledge graph, reducing storage space and computing resource consumption, and ensuring the simplicity and authority of the knowledge graph. This high-quality input data lays a solid foundation for the preliminary construction of the industry knowledge graph in subsequent step S2, making the constructed knowledge graph more accurate and efficient, and ultimately improving the effectiveness and reliability of demand mining based on this knowledge graph.

[0039] In some embodiments of this application, the text data in step S1 includes structured text data and unstructured text data.

[0040] Specifically, structured text data refers to text information with a well-defined data model or pattern, typically existing in the form of tables, database records, XML files, or JSON files. This type of data is highly organized and easy for machines to parse and process; examples include statistical tables in industry reports, financial data in corporate financial statements, and product specification tables. By processing structured text data, precise entities and relationships can be efficiently extracted, providing high-quality initial information for knowledge graphs.

[0041] Unstructured text data refers to textual information that does not possess a predefined data model or pattern. It typically exists in the form of natural language, such as the full text of industry documents, user comments on social media, and discussion posts in industry forums. This type of data contains rich semantic information and tacit knowledge, but its complexity and diversity make it difficult for machines to process, requiring in-depth analysis using natural language processing techniques. By processing unstructured text data, a wider range of entities and relationships can be discovered, capturing deeper information such as industry dynamics and user sentiment.

[0042] In step S1, this application clarifies that text data includes not only unstructured text data but also structured text data. This comprehensive data acquisition strategy ensures the full extraction of information from multi-source, heterogeneous industry data, avoiding information omissions due to the limited data format. The high precision and ease of processing of structured text data, combined with the breadth and depth of unstructured text data, makes the constructed preliminary industry knowledge graph more complete, accurate, and rich. This not only provides a more solid foundation for subsequent entity recognition and relationship extraction but also greatly improves the overall quality of the knowledge graph, thus providing more comprehensive and reliable data support for subsequent demand mining based on the optimized industry knowledge graph.

[0043] In some embodiments of this application, the entity recognition and relation extraction using natural language processing technology in step S2 includes: extracting entities using a named entity recognition model and determining the relationships between entities using a relation classification model or a joint extraction model.

[0044] Specifically, entity extraction using named entity recognition models refers to identifying and extracting entities with specific meanings from text data using specially designed models, such as names of people, places, organizations, product names, and industry terms. Named entity recognition models can be implemented using various techniques. For example, rule-based methods match predefined patterns and dictionaries; statistical machine learning methods, such as Hidden Markov Models (HMMs) or Conditional Random Fields (CRFs), learn entity boundaries and types through training; or deep learning methods, such as Recurrent Neural Networks (RNNs), Convolutional Neural Networks (CNNs), or Transformer models, combined with pre-trained language models (such as BERT and RoBERTa) for entity recognition. These deep learning models can automatically learn text features and exhibit higher accuracy and robustness in complex contexts.

[0045] After entity extraction, a relation classification model or a joint extraction model is used to determine the relationships between entities. Relationship classification models are typically performed after entity recognition. They determine whether a predefined relationship exists between identified entity pairs and classify the relationship. For example, lexical, syntactic, and semantic features between entity pairs can be extracted, and then trained using traditional classifiers such as Support Vector Machines (SVM) or Logistic Regression, or deep learning models (such as BERT-based classifiers) can be used to encode and classify entity pairs and their context. Joint extraction models aim to simultaneously complete entity recognition and relation extraction tasks, avoiding the error accumulation problems that may exist in staged methods and better utilizing the interdependencies between entities and relations. Joint extraction models can be implemented using shared encoders, a combination of sequence labeling and graph structures, or end-to-end multi-task learning, directly outputting entity and relation triples from the original text.

[0046] By explicitly employing a named entity recognition model for entity extraction and combining it with a relation classification model or a joint extraction model for relation determination, this application can significantly improve the accuracy and efficiency of identifying key entities and extracting relationships between entities from massive amounts of industry text data.

[0047] In some embodiments of this application, the step of discovering potential inter-entity associations based on similarity includes: setting a similarity threshold; comparing the similarity between entity vectors with the similarity threshold; and identifying entity pairs with similarity higher than the threshold and no direct relationship edge in the preliminary industry knowledge graph as potential inter-entity associations.

[0048] Specifically, the similarity threshold is a preset value that defines the minimum standard of similarity between entity pairs. Only entity pairs that meet or exceed this standard are considered potentially related. This threshold can be set based on the experience and knowledge of domain experts, through statistical analysis of historical data, or by training and optimizing machine learning models. For example, a floating-point value between 0 and 1 can be set to represent the cosine similarity or the reciprocal of the Euclidean distance between entity vectors. In practice, this threshold can be finely adjusted through iterative experiments and expert feedback according to the needs of specific application scenarios to achieve the best balance between recall and precision. After generating entity vectors and calculating entity pair similarities, the system compares the similarity between entity vectors with the similarity threshold. The purpose is to perform preliminary screening of all calculated entity pair similarities to identify those entity pairs that exhibit sufficiently strong semantic or structural correlation. Specifically, the system iterates through the similarity scores of all entity pairs and compares them one by one with the preset similarity threshold. An entity pair will only be considered as a potential candidate for association if its similarity score meets (e.g., is greater than or equal to) the threshold.

[0049] Subsequently, entity pairs with similarity exceeding a threshold and lacking direct relation edges in the initial industry knowledge graph are identified as potential inter-entity associations. This step is crucial for discovering novel potential associations. After initially screening entity pairs with similarity exceeding the threshold, the system further examines these candidate entity pairs. Specifically, for each candidate entity pair, the system queries the currently constructed initial industry knowledge graph to confirm whether any form of direct relation edge exists between the two entities. If the query results indicate that there is currently no explicit relation edge between the two entities, then the entity pair is ultimately identified as a "potential inter-entity association." This mechanism ensures that the discovered associations are novel and not captured by the existing knowledge graph, thereby avoiding redundant processing of known information and allowing subsequent manual review to focus more on discovering and verifying new knowledge.

[0050] Through the aforementioned technical solution, this application can effectively refine the screening of initially discovered entity relationships. First, by setting and comparing similarity thresholds, it ensures that only semantically highly related entity pairs are considered, thus filtering out a large amount of low-value noise. Second, and more importantly, by excluding entity pairs with existing direct relationships in the initial industry knowledge graph, this application can focus on identifying novel relationships that have potential value but have not yet been explicitly captured. This significantly reduces the number of entity pairs requiring manual review, avoids redundant review of known relationships, and greatly improves the efficiency and accuracy of expert review. This allows the human-machine collaborative correction of the knowledge graph to focus more on discovering and verifying new knowledge, thereby enhancing the completeness, accuracy, and innovativeness of the optimized industry knowledge graph.

[0051] In some embodiments of this application, the manual review in step S4, which incorporates expert knowledge, includes: having domain experts review the accuracy of the entities and relationships extracted from the preliminary industry knowledge graph, and review the rationality of the potential inter-entity associations determined in step S3.

[0052] Specifically, the accuracy of entities and relationships extracted from the initial industry knowledge graph is reviewed by domain experts. This involves manual verification by professionals with deep knowledge and extensive experience in specific industries or professional fields after the initial industry knowledge graph has been constructed using automated technology. These domain experts have a thorough understanding of industry terminology, concepts, entity types, and their interrelationships, and can meticulously review the entities (e.g., company names, products, technologies, events, etc.) and relationships (e.g., "produces," "includes," "belongs to," "affects," etc.) extracted by automated tools. The review covers the correctness of entity boundaries, the accuracy of entity types, the appropriateness of relationship types, the logic of relationship directions, and whether there are any omissions or errors in extraction. For example, experts can use a visual interface to correct the types of identified entities or confirm or reject specific extracted relationships.

[0053] Simultaneously, reviewing the rationality of the potential inter-entity relationships identified in step S3 involves domain experts using their professional judgment to assess the authenticity, logic, and practical value of these relationships within the industry context. These potential relationships are inferred based on data patterns and semantic similarity and may include some relationships that are unreasonable or unimportant in the real world. The expert's role at this stage is to identify these relationships; for example, to determine whether the potential relationship found between "Technology A" and "Product B" exists in the actual industry, whether it is a causal relationship, a parallel relationship, or another type of relationship, and whether this relationship is meaningful for industry analysis. Experts can use an interactive interface to "confirm," "reject," or "mark as pending" these potential relationships and can selectively add annotations to guide subsequent knowledge graph revisions.

[0054] By introducing domain experts to manually review entities, relationships, and potential connections in the initially constructed industry knowledge graph, the limitations of automated technology in understanding complex industry semantics and identifying subtle connections can be effectively compensated for. Leveraging their extensive industry experience and expertise, domain experts can accurately assess the accuracy of automated extraction results and the rationality of potential connections, promptly correcting errors, supplementing omissions, and filtering out false connections that do not conform to industry realities. This significantly improves the quality and reliability of the industry knowledge graph, ensuring that it more accurately and comprehensively reflects the industry's knowledge system. Based on this, demand mining will utilize a higher-quality knowledge graph, thereby outputting more accurate and valuable demand information, avoiding biases in demand mining caused by errors in the underlying knowledge graph, and providing users with more reliable decision support.

[0055] In some embodiments of this application, modifying the knowledge graph based on the review results includes adding, deleting, or modifying entities, relationships, or associations in the knowledge graph based on expert review opinions.

[0056] The expert review opinions are feedback from domain experts who review the accuracy of the entities and relationships extracted from the preliminary industry knowledge graph, as well as the rationality of the potential inter-entity relationships identified in step S3. These opinions typically include confirmation of existing knowledge, identification of errors, suggestions for supplementing missing information, and judgments on the validity of potential relationships. Expert review opinions serve as the basis for revising the knowledge graph, ensuring the professionalism and authority of the revisions. When expert review reveals unidentified or unextracted key industry concepts, terms, figures, organizations, products, or other entities in the preliminary industry knowledge graph, the system provides an interface allowing experts to manually input or identify and add these new entities using auxiliary tools. For example, an expert can point out an emerging technology term or an important industry player that is not included and add it as a new entity to the knowledge graph. If expert review finds redundant, misidentified, or non-compliant entities in the preliminary industry knowledge graph, the system allows experts to remove them from the knowledge graph. For example, a common term that is misidentified as an entity, or an outdated entity that no longer has industry significance, can be deleted. When expert review reveals inaccurate or incomplete entity information in the preliminary industry knowledge graph, the system provides functionality allowing experts to modify entity attributes, names, types, etc. For example, correcting spelling errors, updating category, or supplementing key attribute information. When expert review identifies unextracted but necessary relationships between entities in the preliminary industry knowledge graph, the system allows experts to manually establish these relationships. For example, experts can identify an "application" relationship between a product and a technology, or a "founding" relationship between a company and a founder, and add them to the knowledge graph. If expert review finds incorrectly extracted or unreasonable entity relationships in the preliminary industry knowledge graph, the system allows experts to remove them from the knowledge graph. For example, two entities may be incorrectly identified as having a certain relationship, or a relationship may not hold true in a specific context. When expert review finds inaccurate types, directions, or strengths of existing relationships or potential associations identified in step S3 in the preliminary industry knowledge graph, the system provides functionality allowing experts to modify them. For example, a "containment" relationship can be revised to a more precise "composition" relationship, or the confidence level of a certain association can be adjusted.

[0057] Through the above technical solution, this application effectively solves the problem of relying solely on manual review and lacking an effective correction mechanism. Expert review opinions are no longer merely for identifying problems, but directly drive the optimization process of the knowledge graph. By allowing experts to directly add, delete, or modify entities, relationships, or associations in the knowledge graph, this method can seamlessly integrate the deep knowledge and experience of domain experts into the construction of the knowledge graph. This not only corrects potential recognition errors or omissions in natural language processing technology, but also supplements implicit knowledge and common sense that are difficult for machines to discover, thereby significantly improving the accuracy, completeness, and authority of the industry knowledge graph. The resulting optimized industry knowledge graph will be closer to industry reality, providing a more solid and reliable knowledge foundation for the subsequent demand mining step S5, making the results of demand mining more accurate and valuable. This human-machine collaborative correction mechanism fully leverages the efficiency of machines in processing large-scale data and the accuracy of human expert judgment, achieving continuous iteration and improvement of the knowledge graph quality.

[0058] In some embodiments of this application, step S5, which combines user profiles with the optimized industry knowledge graph for demand mining, includes: parsing the user profile to locate related entities in the optimized industry knowledge graph; and using the related entities as starting points, performing association queries or path exploration in the optimized industry knowledge graph to mine demand information related to the user profile.

[0059] Specifically, the first step is to analyze user profiles to extract key information that corresponds to or is related to entities in the industry knowledge graph. User profiles typically include basic user attributes (such as industry, job title, company size), behavioral data (such as browsing history, search keywords, and purchase records), and preference tags (such as areas of interest in technology and product types). The analysis process may involve natural language processing of unstructured text information in the user profile. For example, techniques such as named entity recognition, keyword extraction, and topic modeling can be used to identify specific entities (such as company names, product models, and technical concepts) or abstract concepts mentioned in the user profile. For structured attributes, predefined mapping rules can be used to directly associate them with specific types of entities or entity attributes in the knowledge graph. Through this analysis process, the diverse information in the user profile can be transformed into identifiable and actionable entity nodes in the knowledge graph, thus providing a clear starting point for subsequent demand mining.

[0060] Secondly, once entities related to the user profile are located in the optimized industry knowledge graph, these entities serve as starting points for association queries or path exploration within the knowledge graph. Association queries refer to starting from the initial entity and, based on the relationships between entities in the knowledge graph, querying directly connected entities (one-hop queries) or entities connected through multiple layers of relationships (multi-hop queries). For example, if the starting entity is "artificial intelligence technology," one can query entities related to this technology such as "application scenarios," "core algorithms," and "related products." Path exploration is more complex, utilizing graph algorithms (such as breadth-first search, depth-first search, shortest path algorithms, or random walk-based algorithms) to discover potential, indirect association paths between entities. For example, starting from a "problem entity" of interest to the user, one can explore which "technology entities" or "solution entities" might solve the problem, thereby revealing the user's potential, deeper needs. Through these association queries and path explorations, the knowledge graph can be systematically traversed to discover new entities, relationships, or patterns related to the user profile that may represent user needs.

[0061] The above technical solution first involves a refined analysis of user profiles, transforming key information into identifiable entities in a knowledge graph, thus providing a clear starting point for subsequent demand mining. Then, based on these identified related entities, structured association queries or path exploration are performed within the optimized industry knowledge graph, enabling efficient and accurate discovery of potential demand information highly matched to the user profiles. This method avoids blind searching or generic matching, significantly improving the targeting and effectiveness of demand mining.

[0062] In some embodiments of this application, the graph embedding technique used in step S3 above includes at least one of the TransE, TransH, Node2Vec, or RotatE algorithms. Specifically, the TransE (Translating Embeddings) algorithm models relations as translation operations of entities in a low-dimensional vector space. That is, for a fact triple (head entity, relation, tail entity), its vector representation satisfies that the head entity vector plus the relation vector is approximately equal to the tail entity vector, thereby effectively capturing the translational relations between entities. The TransH (Translating Embeddings on Hyperplanes) algorithm improves upon TransE by introducing a relation-specific hyperplane, allowing entities to have different projected representations under different relations, thereby better handling complex relations (such as one-to-many, many-to-one, and many-to-many relations) and avoiding the limitations caused by entities sharing the same representation in different relations. The Node2Vec algorithm is a graph embedding method based on random walks. It generates a sequence of node contexts by controlling the random walk strategy (leaning towards breadth-first search or depth-first search), and then uses a Skip-gram model to learn the low-dimensional vector representation of the nodes, thus flexibly capturing local and global structural information of the graph. The RotatE (Rotation Embedding) algorithm, on the other hand, models relations in complex vector space as rotation operations from the head entity to the tail entity. Through this rotation transformation, it can effectively capture rich relational patterns in knowledge graphs, including symmetric / antisymmetric, inversion, and combinatorial relations, thereby generating more expressive entity and relation embeddings.

[0063] Through the aforementioned technical solutions, advanced graph embedding algorithms such as TransE, TransH, Node2Vec, or RotatE can be used in the initial construction of industry knowledge graphs to more effectively capture the complex semantic and structural information of entities and relationships within the knowledge graph. These algorithms can generate high-quality low-dimensional vector representations, enabling entity vectors to more accurately reflect their true semantic and relational characteristics within the knowledge graph. Therefore, the entity similarity calculated based on these high-quality vectors will be more accurate, thus more reliably discovering potential relationships between entities, significantly improving the quality of the initial industry knowledge graph construction and the efficiency of subsequent manual review.

[0064] In some embodiments of this application, the output form of the requirement information in step S5 may include at least one of the following: structured requirement entries, requirement association graphs, or requirement summaries described in natural language.

[0065] Specifically, structured requirement entries refer to organizing and presenting the mined requirement information in a predefined data format. For example, these requirement entries can be represented as key-value pairs, JSON objects, XML documents, or database records, containing fields such as the requirement's name, description, related entities, priority, and source. This structured format facilitates automated processing, storage, retrieval, and integration with other business systems by computer systems, thereby achieving efficient data exchange and utilization.

[0066] A demand relationship graph is a graphical representation of the mined demand information and its interrelationships. This graph consists of nodes and edges, where nodes represent different demands or entities related to those demands, and edges represent various relationships between these demands or entities, such as dependency, inclusion, similarity, or causation. Through this graph format, users can intuitively understand complex demand networks, discover potential demand clusters, conflicts, or points of collaboration, thereby assisting decision-makers in macro-level analysis and planning.

[0067] Natural Language Processing (NLP) descriptions of requirements summaries are concise texts that extract and summarize complex requirements information using NLP techniques. These summaries are typically presented in easily understandable sentences or paragraphs, quickly conveying core requirements, key features, and important background information. Their generation can be based on text summarization algorithms, extracting and integrating key information from raw requirement data and knowledge graphs, aiming to provide convenience for non-technical users or those needing a quick overview of requirements.

[0068] Through the above technical solutions, this application can provide diversified output formats of demand information to adapt to the specific needs of different user groups and application scenarios. Structured demand items facilitate automated processing and integration by the system, improving data utilization efficiency and the accuracy of subsequent analysis; demand correlation graphs intuitively display the complex relationships between demands, helping users quickly understand and discover potential connections, thereby gaining deeper insights; while natural language descriptions of demand summaries provide a concise and easy-to-understand overview, allowing non-technical personnel to quickly grasp the core requirements. This multi-format output significantly enhances the practicality, operability, and user experience of the demand mining results.

[0069] In some embodiments of this application, the user profile in step S5 is constructed based on the user's historical behavior data, attribute information, and preference tags.

[0070] Specifically, user profiles are sets of user characteristics formed by abstracting and modeling multi-dimensional information about users, aiming to comprehensively and accurately depict users' characteristics, behaviors, preferences, and needs. Historical user behavior data refers to the operation records left by users on specific platforms, systems, or scenarios, such as browsing history, search history, purchase history, click history, and interaction history. This data can directly reflect users' interests, usage habits, and potential needs. For example, in the application scenario of industry knowledge graphs, historical user behavior data can include users' reading time, saved articles, search keywords, and participation in discussions on industry report reading platforms. By analyzing this behavioral data, users' explicit and implicit interests can be discovered. Attribute information refers to users' basic static characteristics, such as age, gender, region, occupation, educational background, industry, and company size. This information helps in user group segmentation and fine-grained analysis, providing a macro-context and constraints for demand mining. For example, a financial analyst and a manufacturing engineer, even with similar historical behaviors, will have different needs for industry knowledge due to their occupational attributes. Attribute information is usually obtained through user registration information, questionnaires, or third-party data interfaces. Preference tags are keywords or topics that reflect a user's interests and personalized needs, derived through analysis, mining, or manual annotation of user behavioral data and attribute information. For example, a user might be tagged with preference tags such as "artificial intelligence," "big data," or "new energy vehicles." These tags can be automatically generated by the system or actively set by the user. Preference tags abstract complex user behavior and attribute information into semantic units that are easier to understand and match, thereby enabling more efficient location of relevant entities and relationships within a knowledge graph.

[0071] By constructing user profiles based on users' historical behavioral data, attribute information, and preference tags, it is possible to comprehensively and multidimensionally depict users' characteristics and needs. Historical behavioral data provides dynamic, real-time clues to users' interests, attribute information provides stable, contextual group characteristics of users, and preference tags further abstract and summarize users' personalized interest tendencies. This refined, multi-source integrated user profile enables the system to more accurately understand users' true intentions and potential needs in subsequent demand mining steps. When performing association queries or path exploration in the optimized industry knowledge graph, it can locate related entities based on more accurate user profiles, effectively filtering out knowledge irrelevant to users, thereby significantly improving the matching degree, accuracy, and personalization of demand information.

[0072] In some embodiments of this application, after step S4, step S4a is also included: storing the optimized industry knowledge graph in a knowledge base and providing query services for demand mining in step S5.

[0073] Specifically, the optimized industry knowledge graph is stored in a knowledge base to achieve persistent, structured management, and efficient access to knowledge. The knowledge base can take various forms. For example, it can be a storage system based on graph databases (such as Neo4j or JanusGraph), where entities are represented as nodes, relations as edges, and entity and relation attributes are stored as node or edge attributes. Alternatively, it can be a system based on RDF triples, storing information from the knowledge graph in a subject-verb-object format. After step S4 completes the correction of the knowledge graph and obtains the optimized industry knowledge graph, the system triggers a storage operation to serialize all entities, relations, and their attributes in the graph and write them to the preset knowledge base, thereby ensuring the integrity and traceability of the knowledge.

[0074] Simultaneously, a query service is provided for the demand mining step S5, aiming to enable the demand mining module to efficiently and flexibly retrieve the required knowledge from the knowledge base and support complex query logic. The query service can be implemented by providing standardized API interfaces (such as RESTful APIs) or specialized graph query language interfaces (such as SPARQL and Cypher). The demand mining module can send query requests to the knowledge base through these interfaces, for example, querying entities related to a specific user profile, exploring multi-hop paths between entities, or performing complex pattern matching queries. The query service can internally include a query optimizer and indexing mechanism to ensure efficient retrieval and response on large-scale knowledge graphs.

[0075] Through the above technical solutions, the optimized industry knowledge graph can be persistently stored, avoiding the need to repeatedly build or load the knowledge graph each time requirements are mined, significantly improving knowledge reusability and system operating efficiency. The query service provided by the knowledge base enables the requirements mining module to efficiently and conveniently access and retrieve information in the knowledge graph, thereby improving the accuracy and response speed of requirements mining.

[0076] A second aspect of this application provides a system for constructing industry knowledge graphs and mining requirements based on human-machine collaboration. See also... Figure 2 The system includes: The data acquisition and processing module is used to acquire text data from industry literature, industry reports, social media and industry forums, and to clean and preprocess the text data. The knowledge graph construction module is used to perform entity recognition and relation extraction on the cleaned and preprocessed text data using natural language processing technology in order to construct a preliminary industry knowledge graph. The association discovery module is used to map entities and relationships in the preliminary industry knowledge graph into low-dimensional vectors through graph embedding technology, calculate the similarity between entity vectors, and discover potential associations between entities based on the similarity. The manual verification and correction module provides a human-machine collaborative verification interface, which receives expert knowledge to manually review the preliminary industry knowledge graph and the potential relationships between entities, and corrects the knowledge graph based on the review results to obtain an optimized industry knowledge graph. The demand mining module is used to combine user profiles with the optimized industry knowledge graph to mine demand information and output demand information.

[0077] In some of the embodiments described above in this application, steps such as multi-source data fusion, graph embedding technology application, and knowledge graph update optimization are proposed. However, in the implementation process, existing technologies introduce non-textual data with low relevance to industry knowledge, resulting in noise redundancy. Graph embedding technology is not deeply coupled with the construction process and therefore cannot uncover implicit relationships. Furthermore, the purely algorithm-driven update mechanism lacks expert verification, making it difficult to guarantee the professionalism and accuracy of the knowledge graph. To address this, this application further proposes a systematic technical solution, specifically including the coordinated operation of the following modules: The data acquisition and processing module is configured to focus solely on text data from industry literature, industry reports, social media, and industry forums, avoiding the introduction of non-textual modal data such as images or sensor data, thereby effectively reducing noise and data redundancy. This module removes noise and repetitive information from the text through pre-processing cleaning, ensuring the purity of the input data and providing a high-quality text foundation for subsequent industry knowledge extraction.

[0078] The knowledge graph construction module is designed to utilize natural language processing (NLP) technology to perform entity recognition and relation extraction on the cleaned and preprocessed text data. Entity recognition employs a named entity recognition model, while relation extraction utilizes a relation classification model or a joint extraction model to construct a preliminary industry knowledge graph. This module is strictly limited to processing both structured and unstructured text data, ensuring that the knowledge graph construction closely aligns with industry expertise and real-time market dynamics.

[0079] The association discovery module is designed to map entities and relationships in the initial industry knowledge graph into low-dimensional vectors using graph embedding technology and calculate the similarity between entity vectors. Specifically, this module sets a similarity threshold, compares the similarity of entity vectors with the threshold, and identifies entity pairs with similarity higher than the threshold and no direct relationship edges in the initial knowledge graph as potential inter-entity associations. The graph embedding technology employs at least one of the TransE, TransH, Node2Vec, or RotatE algorithms, deeply coupling vector mapping with the industry knowledge graph construction process, thereby customizing the discovery of implicit industry patterns such as technology associations, product associations, and demand associations.

[0080] The manual verification and correction module provides a human-machine collaborative verification interface, receiving feedback from domain experts on the accuracy of the extracted entities and relationships in the preliminary industry knowledge graph, as well as the reasonableness of potential relationships between entities. Based on expert feedback, this module performs correction operations on the knowledge graph, including adding, deleting, or modifying entities, relationships, or associations, ensuring that the optimized industry knowledge graph meets the core requirements of accuracy and authority. Finally, the module stores the optimized industry knowledge graph in a knowledge base, providing query services for demand mining.

[0081] The demand mining module is configured to combine user profiles built based on historical user behavior data, attribute information, and preference tags, and then perform demand mining based on the optimized industry knowledge graph. This module first parses the user profile to locate related entities in the knowledge graph, then uses these related entities as starting points to perform association queries or path exploration, ultimately outputting structured demand entries, demand association graphs, or demand summaries described in natural language.

[0082] Through the above technical solutions, this application achieves accurate fusion of multi-source text data, avoiding noise interference introduced by non-text data; graph embedding technology is deeply coupled with the entire knowledge graph construction process, effectively uncovering implicit industry connections; and the deep involvement of expert knowledge ensures the professionalism and accuracy of the knowledge graph. Overall, this system overcomes the shortcomings of existing technologies in industry knowledge graph construction, such as data redundancy, insufficient implicit knowledge mining, and lack of professionalism, providing reliable technical support for in-depth mining of industry needs.

[0083] The following example will provide a more detailed explanation of the above technical solution: Suppose that in the field of new energy vehicles, it is necessary to build an industry knowledge graph and mine user needs.

[0084] First, in the data acquisition and processing phase, the system obtains text data from multiple channels. For example, it acquires structured and unstructured industry literature data from academic journals, technical patents, and industry analysis reports related to new energy vehicles; industry report data from market research reports and corporate financial reports; user comments and news updates from social media platforms such as Weibo, WeChat official accounts, and Autohome; and user discussions and technical Q&A data from professional technical forums and car enthusiast communities. This raw text data may contain a large amount of noise, such as advertisements, irrelevant comments, and duplicate content. The system cleans and preprocesses this data to remove noise and duplicate information, ensuring the quality of data in subsequent processing. In this way, this method focuses on fusing professional text data highly relevant to industry knowledge, avoiding the noise and redundancy introduced by fusing non-text data such as images and sensors in existing technologies, thus more effectively integrating industry expertise with real-time market dynamics.

[0085] Next, in the knowledge graph construction phase, the system utilizes natural language processing (NLP) technology to perform entity recognition and relation extraction on the cleaned and preprocessed text data. For example, a pre-trained named entity recognition model (such as a BERT-based model) is used to identify entities such as "Tesla," "CATL," "lithium iron phosphate battery," "charging pile," and "driving range" from the text. Simultaneously, a relation classification model or joint extraction model is used to identify the relationships between these entities; for example, a "supplier" relationship exists between "Tesla" and "CATL," and an "influence" relationship exists between "lithium iron phosphate battery" and "driving range." Through this step, an initial industry knowledge graph containing entities and relationships is constructed.

[0086] Subsequently, in the association discovery phase, the system employs graph embedding techniques, such as TransE or Node2Vec algorithms, to map entities and relationships in the initial industry knowledge graph into low-dimensional vectors. For example, the vectors for "Tesla" and "electric vehicle" will be relatively close in the vector space, while the vectors for "Tesla" and "gasoline vehicle" will be relatively far apart. The system calculates the similarity between these entity vectors and sets a similarity threshold. By comparing the similarity between entity vectors with this threshold, the system can discover entity pairs with similarities higher than the threshold but for which no direct relationship edge has yet been established in the initial knowledge graph. For example, the system might find a high similarity between "solid-state battery" and "energy density," but no explicit direct relationship has been established in the initial graph; in this case, it is identified as a potential entity association. This process deeply couples graph embedding techniques with the construction of the industry knowledge graph, enabling the discovery of hidden technical, product, and demand associations within the industry that are currently difficult to find using existing techniques, thus compensating for the shortcomings of only capturing general semantic associations.

[0087] During the manual verification and correction phase, the system provides a human-machine collaborative verification interface, inviting experts in the new energy vehicle field to review the preliminary industry knowledge graph. Experts will verify the accuracy of extracted entities (e.g., whether "lithium iron phosphate battery" is accurate) and relationships (e.g., whether the "supplier" relationship between "Tesla" and "CATL" is correct). Simultaneously, experts will also review the rationality of potential entity associations identified in step S3 (e.g., the association between "solid-state battery" and "energy density"), judging whether they conform to industry realities. Based on the experts' review opinions, the system will correct the knowledge graph, for example, by adding new entities or relationships, deleting inaccurate entities or relationships, or modifying the attributes of existing entities or relationships. For example, experts may point out that a certain entity is incorrectly identified, or that a certain relationship is not fully extracted; the system will adjust accordingly based on this feedback. The corrected knowledge graph, which is the optimized industry knowledge graph, will then be stored in the knowledge base to provide query services for subsequent demand mining. This human review mechanism, which combines expert knowledge, effectively solves the problem of insufficient professionalism and accuracy of knowledge graphs caused by purely algorithm-driven approaches in existing technologies, ensuring the "accuracy and authority" of knowledge graphs.

[0088] Finally, in the demand mining phase, the system combines user profiles with an optimized industry knowledge graph to mine demands. For example, for user A whose user profile shows a preference for "high performance," "long battery life," and interest in "intelligent driving," the system will analyze the user profile and locate related entities such as "high-performance motor," "high-energy-density battery," and "autonomous driving chip" in the optimized industry knowledge graph. Starting from these entities, the system performs association queries or path exploration in the knowledge graph, for example, exploring the association between "high-performance motor" and "heat dissipation technology," or between "high-energy-density battery" and "charging speed." In this way, the system can uncover potential demands from user A, such as a demand for "efficient heat dissipation system," a demand for "fast charging solution," or a demand for "Level 4 autonomous driving function." Ultimately, the system outputs this demand information in the form of structured demand entries, demand association graphs, or demand summaries described in natural language. The construction of user profiles is based on users' historical behavioral data (such as browsing history and purchase history), attribute information (such as occupation and region), and preference tags (such as enthusiasm for specific brands and technologies). This in-depth approach can accurately identify users' specific needs in the field of new energy vehicles, overcoming the limitations of existing technologies that are difficult to adapt to actual industry application scenarios.

[0089] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A method for constructing industry knowledge graphs and mining requirements based on human-machine collaboration, characterized in that, Includes the following steps: S1. Obtain text data from industry literature, industry reports, social media, and industry forums, and clean and preprocess the text data. S2. Utilize natural language processing technology to perform entity recognition and relation extraction on the cleaned and preprocessed text data in order to construct a preliminary industry knowledge graph. S3. Using graph embedding technology, the entities and relationships in the preliminary industry knowledge graph are mapped into low-dimensional vectors, and the similarity between entity vectors is calculated. Based on the similarity, potential inter-entity associations are determined. S4. The preliminary industry knowledge graph and the potential inter-entity relationships are manually reviewed by experts, and the knowledge graph is corrected according to the review results to obtain an optimized industry knowledge graph. S5. Combining user profiles, perform demand mining based on the optimized industry knowledge graph, and output demand information.

2. The method according to claim 1, characterized in that, In step S1, the cleaning and preprocessing includes: removing noise and duplicate information from the text data.

3. The method according to claim 1, characterized in that, In step S1, the text data includes structured text data and unstructured text data.

4. The method according to claim 1, characterized in that, In step S2, the entity recognition and relation extraction using natural language processing technology includes: extracting entities using a named entity recognition model and determining the relationships between entities using a relation classification model or a joint extraction model.

5. The method according to claim 1, characterized in that, In step S3, discovering potential inter-entity associations based on the similarity includes: Set a similarity threshold; The similarity between entity vectors is compared with the similarity threshold. Entity pairs with similarity higher than the threshold and no direct relationship edge in the preliminary industry knowledge graph are identified as potential inter-entity associations.

6. The method according to claim 1, characterized in that, In step S4, the manual review combined with expert knowledge includes: domain experts reviewing the accuracy of the entities and relationships extracted from the preliminary industry knowledge graph, and reviewing the rationality of the potential inter-entity associations determined in step S3.

7. The method according to claim 6, characterized in that, In step S4, the step of correcting the knowledge graph based on the review results includes: adding, deleting, or modifying entities, relationships, or associations in the knowledge graph based on expert review opinions.

8. The method according to claim 1, characterized in that, In step S5, the step of combining user profiles and performing demand mining based on the optimized industry knowledge graph includes: The user profile is analyzed to locate related entities in the optimized industry knowledge graph; Starting with the associated entities, perform association queries or path explorations in the optimized industry knowledge graph to uncover demand information related to the user profile.

9. The method according to claim 1, characterized in that, After step S4, the process also includes step S4a: storing the optimized industry knowledge graph in a knowledge base and providing query services for demand mining in step S5.

10. A system for constructing industry knowledge graphs and mining requirements based on human-machine collaboration, characterized in that, include: The data acquisition and processing module is used to acquire text data from industry literature, industry reports, social media and industry forums, and to clean and preprocess the text data. The knowledge graph construction module is used to perform entity recognition and relation extraction on the cleaned and pre-processed text data using natural language processing technology in order to construct a preliminary industry knowledge graph. The association discovery module is used to map entities and relationships in the preliminary industry knowledge graph into low-dimensional vectors using graph embedding technology, calculate the similarity between entity vectors, and discover potential associations between entities based on the similarity. The manual verification and correction module provides a human-machine collaborative verification interface, receives expert knowledge for manual review of the preliminary industry knowledge graph and the potential inter-entity relationships, and corrects the knowledge graph based on the review results to obtain an optimized industry knowledge graph. The demand mining module is used to combine user profiles with the optimized industry knowledge graph to mine demand information and output demand information.