Multi-dimensional data association optimization method and medium

By building an industrial chain map knowledge base and thinking chain reasoning mechanism, and using large language models to gradually judge the industrial field classification of multidimensional data, the problems of data association accuracy and inefficiency in traditional methods are solved, and efficient and accurate multidimensional data association are achieved.

CN120448462APending Publication Date: 2025-08-08SUZHOU CASMINO INFORMATION TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510539469.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

Traditional multidimensional data association methods cannot effectively utilize the deep relationship between data for reasoning and judgment, resulting in insufficient accuracy and practicality of correlation results. In addition, large language models are prone to hallucinations when dealing with multidimensional data associations, and it is difficult to deal with complex problems efficiently.

Method used

By building an industrial chain map as a knowledge base, using the large language model's tips from less to more and thinking chain reasoning mechanism, we gradually judge the industrial field classification and link nodes of multi-dimensional data, and accurately match it with the upper and lower levels of the industrial chain map to form the final correlation link.

Benefits of technology

It significantly improves the accuracy and efficiency of multi-dimensional data associations, alleviates the illusion of large language models, and avoids the problems of inefficiency and high error rates in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448462A_ABST
    Figure CN120448462A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-dimensional data association optimization method and a medium, and the method comprises the steps: determining a core dimension of multi-dimensional data, and obtaining a first data set and a second data set according to the core dimension; the second data set constructs an industrial chain atlas through an intermediary mechanism, and the industrial chain atlas relation serves as a knowledge base. And taking an upper and lower layer relationship of the industrial chain graph as a reasoning link of the thinking chain, and performing industrial field classification and node judgment on the data in the first data set by utilizing a small-to-large prompt technology of a large language model so as to output a corresponding label. The labels are accurately matched with the names of all links of the industrial chain atlas, complete links are supplemented to form final association links, and the method effectively improves the accuracy and efficiency of data association.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data processing, and in particular to a multidimensional data association optimization method and medium. Background Art

[0002] With the rapid development of science and technology and the continuous upgrading of industries, the amount of data is exploding, and the correlation analysis of multidimensional data has become crucial in many fields. Traditional data correlation methods often rely on simple keyword matching or rule-based systems. These methods have many limitations when dealing with complex multidimensional data and cannot effectively utilize the deep relationships between data for reasoning and judgment. In addition, as the complexity of industrial chains continues to increase, traditional correlation methods have difficulty accurately locating data at specific locations within the industry chain, which greatly reduces the accuracy and practicality of the correlation results.

[0003] In recent years, large language models (LLMs) have made significant progress in natural language processing, demonstrating powerful language understanding and generation capabilities. However, these models also present challenges when processing multidimensional data associations. For example, they are prone to "hallucinations," where they generate content that is inconsistent with the facts, resulting in reduced accuracy in association results. Furthermore, when processing complex problems, large language models lack effective reasoning mechanisms and knowledge guidance, making them difficult to efficiently handle multi-level, multi-dimensional data association tasks.

[0004] Therefore, how to efficiently correlate and match these multidimensional data to explore potential conversion opportunities and optimize resource allocation is an urgent problem to be solved. Summary of the Invention

[0005] To overcome the above shortcomings, the present invention aims to provide a multi-dimensional data association optimization method and medium to effectively improve the accuracy and efficiency of data association.

[0006] In order to achieve the above objectives, the present invention adopts a technical solution: a multidimensional data association optimization method, comprising:

[0007] S100, determining a core dimension of multidimensional data, and acquiring a first data set and a second data set according to the core dimension;

[0008] S200: The second data set constructs an industrial chain map through an intermediary mechanism, and uses the industrial chain map relationship as a knowledge base;

[0009] S300: Using the upper and lower layer relationships of the industrial chain map as the reasoning links of the thinking chain, using the small-to-many prompting technology of the large language model to classify the data in the first dataset into industrial fields and judge each link point to output corresponding labels;

[0010] S400: Accurately match the label with the name of each link in the industrial chain map, and complete the link to form a final associated link.

[0011] The beneficial effects of the present invention are:

[0012] By building an industrial chain knowledge graph as a knowledge base, we provide professional knowledge support for large language models, effectively alleviate the "hallucination" phenomenon of large language models, and significantly improve the accuracy of data association.

[0013] By utilizing the reasoning mechanism of the thinking chain and the prompting technology from small to large, complex problems can be decomposed into multiple sub-problems and solved step by step, avoiding the low efficiency and high error rate caused by the one-time processing of complex data in traditional methods.

[0014] By combining thought chains, small-to-many prompting technology, and industry chain knowledge graphs, the accuracy and efficiency of data association can be effectively improved.

[0015] Specifically, S200, the second data set constructs an industrial chain map through an intermediary mechanism, and uses the industrial chain map relationship as a knowledge base, specifically including:

[0016] S21. Define the key links at the front level of the industry chain map;

[0017] S22: The large language model provides clear input and constraints, and the large language model is expanded and updated at the industry chain level according to the second data set to form a graph database;

[0018] S23. Summarize the triplets in the graph database and organize them into a complete industrial chain map.

[0019] Specifically, S300 uses the upper and lower layer relationships of the industrial chain graph as the reasoning link of the thinking chain, and utilizes the small-to-many prompting technology of the large language model to classify the data in the first data set into industrial fields and judge each link point to output corresponding labels.

[0020] S31, inputting the data in the first data set into the large language model for domain judgment;

[0021] S32: The large language model matches the field output in S31 with the first level of the industrial chain map, and inputs the second level link corresponding to the industrial chain map to perform secondary link judgment on the data content in S31;

[0022] S33: The large language model matches the secondary link output in S32 with the second level of the industrial chain map, and inputs the third level link corresponding to the industrial chain map, and performs a third level link judgment on the data content in S31;

[0023] S34. Repeat S33 to judge the next level link until the final link of the data in S31 in the industrial chain map is judged and output as a label.

[0024] Furthermore, the second data set is updated in real time or periodically, and the nodes and edge relationships of the industrial chain graph are automatically updated according to the updated second data set to form a new knowledge base.

[0025] The knowledge base supports dynamic updates to ensure the timeliness and accuracy of the knowledge base and adapt to the dynamic changes in the industrial chain.

[0026] Furthermore, the data in the first data set only includes the first source data. At this time, multiple large language models are set to classify the first source data into industry fields and judge each link point. The label is the intersection of the judgment results of multiple large language models.

[0027] There is a certain degree of divergence when using data from a single source to classify industrial fields and judge each link point. Therefore, the sources of selected data are diversified, and judgments are made on data from different sources. The intersection of different judgment methods is selected, and finally the intersection of the two is taken to improve the accuracy of judgment.

[0028] Furthermore, the data in the first data set is the sum of the first source data and the second source data. At this time, the first source data and the second source data are respectively classified into industry fields and judged at each link point, and the label is the intersection of the judgment results of the first source data and the second source data.

[0029] Some data can only have one source, but by using different large language models to classify industry fields and judge each link point, the intersection is finally taken to improve accuracy.

[0030] Specifically, the first source data is internal data, and the second source data is network data retrieved from the network using a large language model.

[0031] Furthermore, before proceeding to step S300, a large language model is used to extract core data from some of the data in the first dataset. This extracted core data is then categorized by industry and used to determine each link. This extracted core data is then further processed to improve the accuracy of the hierarchical link determination.

[0032] Furthermore, the second data set is industry and industry research report data related to a specific industry.

[0033] The present invention also discloses a storable medium, wherein a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the multidimensional data association optimization method are implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 This is a process of an embodiment of the present invention Figure 1 ;

[0035] Figure 2 This is a process of an embodiment of the present invention Figure 2 ;

[0036] Figure 3 This is a process of an embodiment of the present invention Figure 3 ;

[0037] Figure 4 A schematic diagram of an industrial chain map in one embodiment of the present invention;

[0038] Figure 5 Schematic diagram of the decomposition of thought chain problems in one embodiment of the present invention. DETAILED DESCRIPTION

[0039] The preferred embodiments of the present invention are described in detail below with reference to the accompanying drawings so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making a clearer and more precise definition of the protection scope of the present invention.

[0040] See attached Figure 1 As shown, a multidimensional data association optimization method of the present invention specifically includes:

[0041] S100: Determine a core dimension of multidimensional data, and obtain a first data set and a second data set according to the core dimension.

[0042] For example, in the field of scientific and technological achievement transformation, the core dimensions mainly revolve around data related to scientific and technological achievement transformation. The first dataset includes enterprise and institutional data, talent data, patent data, paper data, and investment and financing data. Talent data includes research directions, professional backgrounds, and academic achievements; patent data includes patent abstracts and technical fields; paper data includes paper abstracts and research topics; and investment and financing data includes investment and financing projects. The second dataset consists of industry and industrial research reports related to specific industries, typically unstructured data such as industry analysis reports and industry classification standards.

[0043] The core dimensions are artificially defined and may be adjusted in different fields. The data in the first and second datasets can be obtained from the Internet or are internal company materials.

[0044] S200. The second data set constructs an industrial chain map through an intermediary mechanism, and uses the industrial chain map relationship as a knowledge base.

[0045] The industrial chain map presents a hierarchical structure, including multiple levels from top to bottom, clearly showing the position of the nodes of each link in the chain, and storing the nodes and edge relationships of the industrial chain map to form a knowledge base. This knowledge base provides professional knowledge support for the subsequent large language model to make judgments in fields and links, significantly improving the accuracy and efficiency of judgments.

[0046] In one embodiment, the knowledge base supports dynamic updates, that is, the second data set can be updated in real time or periodically, and the nodes and edge relationships of the industrial chain graph are automatically updated for the updated second data set to form a new knowledge base, thereby ensuring the timeliness and accuracy of the knowledge base.

[0047] S300. Use the upper and lower layer relationships of the industrial chain map as the reasoning link of the thinking chain, and use the large language model's from-few-to-many prompting technology to classify the data in the first data set into industrial fields and judge each link point to output the corresponding label.

[0048] According to the dimensional characteristics of different data in the first data set, an exclusive prompt is set. According to the hierarchical relationship between the entities in the industrial chain map, first determine which field of the first level of the industrial chain map the input data belongs to. Then, according to the secondary sub-links included in the first-level links of the industrial chain map, determine the corresponding second-level sub-links. Use the reasoning bridging method of the thinking chain to obtain the link and level to which the data belongs.

[0049] When making hierarchical judgments, the most relevant fragments or data for the question are first retrieved from the knowledge base. This retrieval method primarily uses vector search, combining the retrieved information with the input question and feeding it into the Big Language Model to enhance the model's understanding and ability to answer questions. Specifically, Big Language prioritizes reference to the knowledge base. If no reference is available, Big Language will provide a divergent answer, labeling different data dimensions.

[0050] For example, the data in the first input data set is: Taicang Xinxin Photovoltaic Technology Co., Ltd.'s products are ['high-purity polycrystalline silicon wafers', 'monocrystalline silicon wafers', 'solar-grade and electronic-grade polycrystalline silicon wafers', 'solar cells and modules', 'polycrystalline silicon solar panels', 'solar photovoltaic modules', 'monocrystalline silicon solar panels', 'solar photovoltaic power generation systems', 'photovoltaic power station operation and maintenance services']. The knowledge base search obtains one of the results: photovoltaic_upstream_raw materials and equipment_raw materials_silicon materials_monocrystalline silicon, etc., and then the link is input into the large language model, so that the field and link to which it belongs are judged, which reduces model illusions and makes the judgment more accurate.

[0051] This approach, based on thought chains and industry chain diagrams, can effectively alleviate the illusions of large language models and improve judgment accuracy. It also directly uses the output labels of large models as association criteria, improving the efficiency of multidimensional data association.

[0052] S400: Accurately match the labels with the names of each link in the industrial chain map, and complete the link to form the final associated link.

[0053] In this example, by constructing an industry chain knowledge graph as a knowledge base, we provide professional knowledge support for large language models, effectively alleviating the "hallucination" phenomenon of large language models and significantly improving the accuracy of data association. By leveraging the reasoning mechanism of thought chains and the prompting technology from small to large, complex problems are broken down into multiple sub-problems and solved step by step, avoiding the inefficiency and high error rate caused by traditional methods of processing complex data all at once.

[0054] See attached Figure 2 As shown, step S200 specifically includes:

[0055] S21. Define the key links at the front level of the industrial chain map.

[0056] The key links at the front level are manually defined. A specific industry is selected and its core links are manually sorted out. For example, the photovoltaic industry is selected to set the photovoltaic industry chain, including upstream raw materials and equipment, midstream battery cell and component manufacturing, and downstream power generation systems and applications.

[0057] S22. The large language model provides clear input and constraints, and the large language model expands and updates the industrial chain level according to the second data set to form a graph database.

[0058] For example, the prompt question of the large language model is set as: What are the raw materials and equipment required for the upstream of the photovoltaic industry chain?

[0059] The second dataset was fed into the large language model, which produced the following output: In the photovoltaic industry chain, raw materials primarily include silicon (such as monocrystalline silicon, polycrystalline silicon, and amorphous silicon, with solar-grade polycrystalline silicon having a purity between 6N and 9N), silver and aluminum pastes (used for conductivity and connection of solar cells), PET base films and fluorine films (used for photovoltaic module encapsulation to protect cells and improve stability and durability), and other auxiliary materials (such as copper, indium, selenium, perovskite, gallium arsenide, and cadmium telluride). Regarding equipment, silicon material equipment (such as reduction furnaces and pulling furnaces) is used to produce high-purity silicon materials; silicon wafer equipment (such as slicers, grinders, and polishers) is used for cutting, grinding, and polishing silicon wafers.

[0060] The large language model outputs a paragraph based on the second dataset. The large language model then extracts entity names and relationships from this paragraph and stores them in a graph database. The graph database is divided into multiple triples. For example, triples include photovoltaics_upstream_raw materials and equipment; raw materials and equipment_raw materials_silicon materials; raw materials_silicon materials_monocrystalline silicon; and so on.

[0061] S23. Summarize the triplets in the graph database and organize them into a complete industrial chain map.

[0062] For example, please refer to the attached industrial chain map. Figure 4 As shown, Photovoltaic_Upstream_Raw Materials and Equipment_Raw Materials_Silicon Materials_Monocrystalline Silicon, this is a complete link rule.

[0063] In this embodiment, the complete industrial chain map is input into the knowledge base for vectorization. In the subsequent step S300, the large language model gives priority to selecting key information of the input text that exists in the industrial chain map when generating labels. If the similarity is relatively high, the industrial chain link is input into the large language model, and an answer is given after consideration.

[0064] See attached Figure 3 As shown, step 300 specifically includes:

[0065] S31. Input the data in the first data set into the large language model for domain judgment.

[0066] Fields are selected from a set range of fields, and optional fields are manually set. For example, fields are set according to the industry categories of strategic emerging industry standards and national industry standards, including energy storage, photovoltaics, power batteries, hydrogen energy, biomass energy, nuclear energy, wind energy, new display, consumer electronics, semiconductors and integrated circuits, photonics, biomedicine, big health, medical devices, robotics, engineering machinery, energy conservation and environmental protection, industrial mother machines and integrated equipment, aviation, aerospace, elevators, online new economy, computing power economy, artificial intelligence, industrial Internet, new energy vehicles, auto parts, intelligent vehicle networking, information technology application innovation, industrial software, new chemical materials, advanced metal materials, new nanomaterials, chemical fibers, clothing and home textiles, green home appliances, food and beverages, sensors, metaverse, quantum technology, marine engineering, shipbuilding, rail transit, instrumentation, carbon fiber, graphene, superconducting materials, metamaterials, intelligent bionic materials, 3D printing materials, green building materials, smart grid, modern agriculture and food, big data and cloud computing, and communications.

[0067] S32. The large language model matches the fields output in S31 with the first level of the industrial chain map, and inputs the second-level links corresponding to the industrial chain map to perform secondary link judgment on the data content in S31.

[0068] S33. The large language model matches the secondary links output in S32 with the second level of the industrial chain map, and inputs the third-level links corresponding to the industrial chain map to make a third-level link judgment on the data content in S31.

[0069] S34. Repeat S33 to judge the next level link until the final link of the data in S31 in the industrial chain map is judged and output as a label.

[0070] Steps S31-S34 use this prompting technology from small to large to make judgments layer by layer, and ultimately accurately locate the specific position of the data in the industrial chain map.

[0071] In the process of steps S31-S34, the original problem is decomposed into sub-problems and gradually solved in the process of thinking chain prompts, so as to achieve a progressive judgment from top to bottom. Figure 5As shown, the original question is: Based on the company's products, determine which links in the industrial chain they belong to. This original question is broken down into multiple sub-questions at each level. Sub-question 1: Based on the company's products, determine which fields are involved. Optional fields include fields 1, 2, 3, and 4... → Sub-question 2: Based on the company's products, determine which upstream, midstream, and downstream links are related to fields 1 and 2. Field 1 includes links 1, 2, and 3, and field 2 includes links 1, 2, and 3. → Sub-question 3: Based on the company's products, determine which sub-links of link 1 are related to fields 1 and 2. Field 1 includes: Link 1 includes: sub-links 1 and 2, and field 2: Link 1 includes: sub-links 1 and 2. This layer-by-layer decomposition of the questions is then fed into the large language model, which then provides an answer to each question, determining the specific position of the data in the industrial chain map step by step.

[0072] In one embodiment, the data in the first data set have different sources depending on the dimensions to which they belong. Some dimensions have only first source data, and the first source data is existing internal data. Some dimensions have data that are the sum of first source data and second source data, and the second source data is network data retrieved from the network using a large language model. When the data is the sum of existing internal data and network data retrieved from the network using a large language model, the internal data and network data are respectively classified into industry fields and judged at each link point through steps S31-S34. The final field and link point in each step are the intersection of the judgment results of the internal data and the network data. The data is only existing internal data. In order to improve the accuracy, multiple large language models can be added to classify the data into industry fields and judge the various link points. The final field and link point are the intersection of the judgment results of multiple language models.

[0073] For example, enterprise organization data includes existing internal data and network data detected on the network using a large language model.

[0074] The first step is to make domain judgments on existing internal data and network data respectively.

[0075] Internal data:

[0076] Input model "prompt": "XXX Co., Ltd.'s business scope includes: production and research and development of solar-grade and electronic-grade polycrystalline silicon wafers, silicon single crystals with a diameter of 200 mm or more, and polished wafers; research and development of solar cells and modules; sales of self-produced products; and provision of related technical services. Photovoltaic equipment and component sales; information consulting services (excluding licensed information consulting services)." Based on the main products listed in the business scope, determine the company's areas of involvement. [Important] areas must be identified from the field list below. Unable to identify these areas will be output as other areas. The field list is fixed as: Energy Storage, Photovoltaics, Power Batteries, and Hydrogen Energy. The output format is fixed as: ["Field 1", "Field 2"].

[0077] Knowledge base search results: Photovoltaic_Upstream_Raw Materials and Equipment_Silicon Materials_Polysilicon, Photovoltaic_Upstream_Raw Materials and Equipment_Silicon Materials_Monocrystalline Silicon, etc.

[0078] Model output: ["photovoltaic", "semiconductor integrated circuit"].

[0079] Network data:

[0080] Input model "prompt": "Search online for XXX Co., Ltd.'s main products. Based on the product information, determine the company's areas of involvement. [Important] areas must be identified from the field list below. If no area can be identified, the output will be other areas." The field list is fixed as: energy storage, photovoltaics, power batteries, and hydrogen energy. The output format is fixed as: ["Field 1", "Field 2"].

[0081] Knowledge base search results: Photovoltaic_Upstream_Raw Materials and Equipment_Silicon Materials_Polysilicon, Photovoltaic_Upstream_Raw Materials and Equipment_Silicon Materials_Monocrystalline Silicon, etc.

[0082] Output result: ["photovoltaic"].

[0083] We can see the difference between the two judgment methods mentioned above. The first judgment based on internal data has a certain degree of divergence, and the second online search will also have the same problem. Therefore, multiple methods are needed to make judgments, select the intersection of different judgment methods, and finally take the intersection of the two.

[0084] Field final result: ["photovoltaic"].

[0085] The second step is to determine the secondary links to which the existing internal data and network data belong.

[0086] Internal data:

[0087] Input model "prompt": "XXX Co., Ltd.'s business scope includes: production and development of solar-grade and electronic-grade polycrystalline silicon wafers, silicon single crystals with a diameter of 200mm or more, and polished wafers; development of solar cells and modules; sales of self-produced products and provision of related technical services; sales of photovoltaic equipment and components; information consulting services (excluding licensed information consulting services)." This company is involved in the photovoltaic field. Based on the above business scope, which links of the photovoltaic industry does this company participate in? The photovoltaic industry chain includes the following links: raw materials and equipment, modules, cell equipment, module equipment, power generation systems, and applications. The output format is fixed as: [Field]: ["Link 1", "Link 2"...]

[0088] Knowledge base search results: Photovoltaic_Upstream_Raw Materials and Equipment_Silicon Materials_Polysilicon, Photovoltaic_Upstream_Raw Materials and Equipment_Silicon Materials_Monocrystalline Silicon, Photovoltaic_Midstream_Photovoltaic Modules, etc.

[0089] Model output results: [Photovoltaic]: [“Raw materials and equipment”, “Cell equipment”, “Module equipment”, “Module”].

[0090] Network data:

[0091] Input model "prompt": "XXX Co., Ltd.'s main products. Based on product information, determine which links in the photovoltaic industry the company is involved in. The photovoltaic industry chain includes: raw materials and equipment, components, cell equipment, module equipment, power generation systems, and applications. The output format is fixed as: [Field]: ["Link 1", "Link 2"...]

[0092] Knowledge base search results: Photovoltaic_upstream_raw materials and equipment_silicon materials_polysilicon, Photovoltaic_upstream_raw materials and equipment_silicon materials_monocrystalline silicon, Photovoltaic_midstream_modules, etc.

[0093] Model output: [Photovoltaic]: [“Raw materials and equipment”, “Components”].

[0094] Similarly, we select the intersection of different judgment methods and finally take the intersection of the two.

[0095] The final result of the secondary link: [“Photovoltaic”]: [“Raw materials and equipment”, “Components”].

[0096] The third step is to determine the three-level links to which the existing internal data and network data belong.

[0097] The judgment of the third level is the same as the second step above.

[0098] Internal data:

[0099] Input model "prompt": "XXX Company's business scope includes: production and R&D of solar-grade and electronic-grade polycrystalline silicon wafers, silicon single crystals with a diameter of 200mm or greater, and polished wafers; R&D of solar cells and modules; sales of self-produced products and provision of related technical services. Sales of photovoltaic equipment and components; information consulting services (excluding licensed information consulting services)." This involves raw materials, equipment, and modules in the photovoltaic field. Based on this business scope, determine which sub-segments of the raw materials, equipment, and modules are involved in this company? 1. Raw materials and equipment include: silicon materials, silicon wafers, silver paste, aluminum paste, PET base film, fluorine film, other auxiliary materials, single crystal furnaces, polycrystalline furnaces, slicers, roller mills, polishers, etc.; 2. Modules include: solar cells, photovoltaic glass, EVA film, backsheets, frames, junction boxes, sealants, etc. The output format is fixed as: ["Step 1"]:[Sub-Step 1", "Sub-Step 2"...].

[0100] Knowledge base search results: Photovoltaic_Upstream_Raw Materials and Equipment_Silicon Materials_Polysilicon, Photovoltaic_Upstream_Raw Materials and Equipment_Silicon Materials_Monocrystalline Silicon, Photovoltaic_Midstream_Photovoltaic Modules, etc.

[0101] Model output results: [“Raw materials and equipment”]: [“Silicon material”, “Silicon wafer”, “Single crystal furnace”, “Slicer”, “Polishing machine”]; [“Component”]: [“Cell”, “Photovoltaic glass”, “EVA film”, “Back sheet”, “Frame”, “Junction box”, “Sealant”].

[0102] Network data:

[0103] Input model "prompt": "Search online for XXX Co., Ltd.'s main products. Based on the product information, determine which sub-segments of the photovoltaic raw materials, equipment, and components are involved. 1. Raw materials and equipment include: silicon materials, silicon wafers, silver paste, aluminum paste, PET base film, fluorine film, other auxiliary materials, single crystal furnaces, polycrystalline furnaces, slicers, roller mills, polishers, etc.; 2. Components include: solar cells, photovoltaic glass, EVA film, backplane, frame, junction box, sealant, etc. The output format is fixed as: ["Link 1"]: [Sub-Link 1", "Sub-Link 2"...].

[0104] Knowledge base search results: Photovoltaic_Upstream_Raw Materials and Equipment_Silicon Materials_Polysilicon, Photovoltaic_Upstream_Raw Materials and Equipment_Silicon Materials_Monocrystalline Silicon, Photovoltaic_Midstream_Photovoltaic Modules, etc.

[0105] Model output: [“Raw materials and equipment”]: [“Silicon material”, “Silicon wafer”]; [“Component”]: [“Solar cell”]

[0106] Similarly, we select the intersection of different judgment methods and finally take the intersection of the two.

[0107] The final result of the secondary link: [“Raw materials and equipment”]: [“Silicon materials”, “Silicon wafers”]; [“Components”]: [“Battery cells”].

[0108] If the secondary link is the final link, the labels output in step S300 for this data are: Photovoltaic: ["Raw Materials and Equipment"]: ["Silicon Materials", "Silicon Wafers"]; ["Modules"]: ["Solar Cells"]. In step S400, these labels are precisely matched with the names of each link in the industry chain map, completing the final associated links: Photovoltaic_Upstream_Raw Materials and Equipment_Silicon Materials; Photovoltaic_Upstream_Raw Materials and Equipment_Silicon Wafers, etc. Ultimately, the data dimension's industry sector, secondary link, sub-link labels, and corresponding associated links are obtained, enabling the association of multidimensional data.

[0109] For some data in the first data set, before performing step S300, it is necessary to extract core data through a large language model, and then perform subsequent data processing on the extracted core data, so as to improve the accuracy of hierarchical link judgment.

[0110] For example, the patent data are all existing internal data. Before the label extraction of the patent data, the patent title and abstract are first selected to extract the key data of the product and technology, and the key data of the product and technology are used to determine the field to which it belongs.

[0111] Input model "prompt": "Sample text: title and summary. Extract the name of a key application product from the above text, and extract a technical optimization name from the detailed technical description."

[0112] Output results: Product name: Single crystal silicon wafer slurry system; Technology name: Conductivity adjustment method.

[0113] Afterwards, the field to which it belongs is determined, and the product name: single crystal silicon wafer slurry system; technology name: conductivity adjustment method are input into the large language model.

[0114] The method in this embodiment constructs an industrial chain knowledge graph as a knowledge base, utilizes the reasoning mechanism of the thinking chain and the prompting technology from small to large, and gradually inputs multidimensional data into a large language model for association matching, thereby effectively improving the accuracy and efficiency of data association.

[0115] The present invention also discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the multidimensional data association optimization method are implemented.

[0116] Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing related hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of the above-mentioned various method embodiments. The computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device that can carry the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunications signal, and a software distribution medium.

[0117] The above embodiments are only for illustrating the technical concept and features of the present invention. Its purpose is to enable people familiar with this technology to understand the content of the present invention and implement it. It cannot be used to limit the scope of protection of the present invention. Any equivalent changes or modifications made according to the spirit of the present invention should be included in the scope of protection of the present invention.

Claims

1. A multidimensional data association optimization method, characterized by: include: S100, determining a core dimension of multidimensional data, and acquiring a first data set and a second data set according to the core dimension; S200: The second data set constructs an industrial chain map through an intermediary mechanism, and uses the industrial chain map relationship as a knowledge base; S300: Using the upper and lower layer relationships of the industrial chain map as the reasoning links of the thinking chain, using the small-to-many prompting technology of the large language model to classify the data in the first dataset into industrial fields and judge each link point to output corresponding labels; S400: Accurately match the label with the name of each link in the industrial chain map, and complete the link to form a final associated link.

2. The multidimensional data association optimization method according to claim 1, characterized in that: S200: The second data set constructs an industrial chain map through an intermediary mechanism, and uses the industrial chain map relationship as a knowledge base, specifically including: S21. Define the key links at the front level of the industry chain map; S22: The large language model provides clear input and constraints, and the large language model is expanded and updated at the industry chain level according to the second data set to form a graph database; S23. Summarize the triplets in the graph database and organize them into a complete industrial chain map.

3. The multidimensional data association optimization method according to claim 1, wherein: S300: Using the upper and lower layer relationships of the industrial chain map as the reasoning links of the thinking chain, and using the small-to-many prompting technology of the large language model to classify the data in the first dataset into industrial fields and judge each link point to output corresponding labels, specifically including: S31, inputting the data in the first data set into the large language model for domain judgment; S32: The large language model matches the field output in S31 with the first level of the industrial chain map, and inputs the second level link corresponding to the industrial chain map to perform secondary link judgment on the data content in S31; S33: The large language model matches the secondary link output in S32 with the second level of the industrial chain map, and inputs the third level link corresponding to the industrial chain map, and performs a third level link judgment on the data content in S31; S34. Repeat S33 to judge the next level link until the final link of the data in S31 in the industrial chain map is judged and output as a label.

4. The multidimensional data association optimization method according to any one of claims 1 to 3, characterized in that: The second data set is updated in real time or periodically, and the nodes and edge relationships of the industrial chain graph are automatically updated according to the updated second data set to form a new knowledge base.

5. The multidimensional data association optimization method according to claim 1, wherein: The data in the first data set only includes the first source data. At this time, multiple large language models are set to classify the first source data into industry fields and judge each link point. The label is the intersection of the judgment results of multiple large language models.

6. The multidimensional data association optimization method according to claim 1, characterized in that: The data in the first data set is the sum of the first source data and the second source data. At this time, the first source data and the second source data are respectively classified into industry fields and judged at each link point, and the label is the intersection of the judgment results of the first source data and the second source data.

7. The multidimensional data association optimization method according to claim 6, characterized in that: The first source data is internal data, and the second source data is network data retrieved from the network using a large language model.

8. The multidimensional data association optimization method according to claim 1, characterized in that: For some of the data in the first data set, before performing step S300, it is necessary to extract core data from the data through a large language model, and then classify the extracted core data into industry fields and judge each link point.

9. The multidimensional data association optimization method according to claim 1, characterized in that: The second data set is industry and industry research report data related to a specific industry.

10. A storable medium, characterized in that: A computer program is stored on a computer-readable storage medium, and when the computer program is executed by a processor, the steps of the multidimensional data association optimization method according to any one of claims 1 to 9 are implemented.