Method, device and equipment for recycling of biomaterials based on multi-modal large language model

By combining a multimodal large language model with a knowledge graph, the system automatically identifies images of recycled materials and generates recycling reports, solving the problem of low report generation efficiency caused by the dispersion of knowledge about recycled materials and achieving fast and accurate automatic generation of recycling reports.

CN120804344BActive Publication Date: 2026-02-06HANGZHOU TIANYAN ZHILIAN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511299644.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2026-02-06
Estimated Expiration
2045-09-12

AI Technical Summary

Technical Problem

In the current technology, knowledge about recycled materials is scattered across different platforms and documents, lacking a unified organization, which leads to low efficiency in generating recycled material recycling reports and an inability to quickly process recycling needs.

Method used

A method based on a multimodal large language model is adopted, which combines knowledge graphs in the form of node vectors and text to automatically identify images of recycled materials and generate material information. Structured and unstructured knowledge is acquired through parallel threads to automatically generate recycling reports.

Benefits of technology

It enables automatic recognition and rapid knowledge retrieval of images of recycled materials, generating recycling reports with practical guidance, reducing reliance on manual identification, and improving the efficiency and accuracy of report generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120804344B_ABST
    Figure CN120804344B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a method, device and equipment for recycling of recycled materials based on a multi-modal large language model. A specific embodiment of the method comprises: generating material information and query statement information corresponding to a target recycled material image by using a pre-trained multi-modal large language model; querying a knowledge vector block corresponding to the query statement information from a material knowledge base by using a first thread; querying first material knowledge corresponding to the knowledge vector block from a first knowledge graph; querying second material knowledge corresponding to the material information from a second knowledge graph by using a second thread; determining target material knowledge corresponding to an output of a target thread; and generating a material recycling report corresponding to the material information by using the multi-modal large language model according to the target material knowledge, so as to recycle the recycled materials according to the material recycling report. The embodiment automatically identifies the recycled material image, quickly acquires recycled material-related knowledge, and generates a recycling report.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present disclosure relate to the field of computer technology, and in particular, to a method, device and equipment for recycling of secondary materials based on a multi-modal large language model. BACKGROUND

[0002] At present, with the global promotion of sustainable development and circular economy, secondary materials (including waste metals, plastics, paper, electronic components, construction waste and various types) have become efficient recycling and value maximization resources. To recycle diversified secondary materials, detailed and accurate recycling reports are needed. For the generation of secondary material recycling reports, the commonly used method is to rely on manual review and writing.

[0003] However, when the above method is used to generate the recycling report of secondary materials, the following technical problems often exist:

[0004] Due to the dispersion of secondary material knowledge, the relevant knowledge content is scattered in different platforms, documents and industry specifications, and lacks unified organization and collection. The staff spends a long time in the search process and cannot quickly obtain the corresponding knowledge of the secondary materials. This leads to low efficiency of report generation and difficulty in timely processing of secondary material recycling needs.

[0005] The above information disclosed in the background section is only intended to enhance the understanding of the background of the present inventive concept, and therefore, it can include information that does not form the prior art known to those of ordinary skill in the art in the country. SUMMARY

[0006] The summary section is used to introduce the concepts in a brief form, which will be described in detail in the specific embodiments section. The summary section is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0007] Some embodiments of the present disclosure propose a method, device and equipment for recycling of secondary materials based on a multi-modal large language model to solve one or more of the technical problems mentioned in the background section.

[0008] In a first aspect, some embodiments of the present disclosure provide a method for recycling of recycled materials based on a multi-modal large language model, comprising: generating, by using a pre-trained multi-modal large language model, material information corresponding to a target recycled material image and query sentence information; performing, by using a first thread, the following material knowledge acquisition steps: querying a knowledge vector block corresponding to the query sentence information from a material knowledge base; querying first material knowledge corresponding to the knowledge vector block from a first knowledge graph corresponding to a recycled material field, wherein the first knowledge graph is a knowledge graph in the form of node vectors; querying, by using a second thread, second material knowledge corresponding to the material information from a second knowledge graph corresponding to the recycled material field, wherein the second knowledge graph is a knowledge graph in the form of text; determining target material knowledge corresponding to an output of a target thread, wherein the target thread is a thread that ends execution earliest among the first thread and the second thread, and the target material knowledge is knowledge that is generated earliest among the first material knowledge and the second material knowledge; and generating, by using the multi-modal large language model, a material recycling report corresponding to the material information according to the target material knowledge, so as to recycle the recycled material according to the material recycling report.

[0009] In a second aspect, some embodiments of the present disclosure provide an apparatus for recycling of recycled materials based on a multi-modal large language model, comprising: an acquisition unit configured to generate, by using a pre-trained multi-modal large language model, material information corresponding to a target recycled material image and query sentence information; a first thread unit configured to perform, by using a first thread, the following material knowledge acquisition steps: querying a knowledge vector block corresponding to the query sentence information from a material knowledge base; querying first material knowledge corresponding to the knowledge vector block from a first knowledge graph corresponding to a recycled material field, wherein the first knowledge graph is a knowledge graph in the form of node vectors; a second thread unit configured to query, by using a second thread, second material knowledge corresponding to the material information from a second knowledge graph corresponding to the recycled material field, wherein the second knowledge graph is a knowledge graph in the form of text; a knowledge generation unit configured to determine target material knowledge corresponding to an output of a target thread, wherein the target thread is a thread that ends execution earliest among the first thread and the second thread, and the target material knowledge is knowledge that is generated earliest among the first material knowledge and the second material knowledge; and a report generation unit configured to generate, by using the multi-modal large language model, a material recycling report corresponding to the material information according to the target material knowledge, so as to recycle the recycled material according to the material recycling report.

[0010] In a third aspect, some embodiments of the present disclosure provide an electronic device, comprising: one or more processors; a storage device having one or more programs stored thereon, when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any implementation manner of the first aspect.

[0011] In a fourth aspect, some embodiments of the present disclosure provide a computer readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method as described in any implementation manner of the first aspect.

[0012] The above various embodiments of the present disclosure have the following beneficial effects: through the multi-modal large language model-based recycled material recycling method of some embodiments of the present disclosure, the recycled material image can be automatically identified, the structured and unstructured knowledge graph is combined to realize fast knowledge retrieval, and a recycling report is automatically generated. In the report generation process, the traditional method often needs manual identification and classification of the material first, and then searches for processing specifications and industry information through multiple platforms, which takes a long time and cannot quickly generate a recycling report with practical guiding significance according to the material image.

[0013] Based on this, the multi-modal large language model-based recycled material recycling method of some embodiments of the present disclosure first generates the material information and query statement information corresponding to the target recycled material image by using the pre-trained multi-modal large language model. In this way, automatic perception and semantic understanding of the recycled material image can be achieved, reducing the dependence on manual identification and providing a structured input basis for subsequent knowledge retrieval. Second, using the first thread, the following material knowledge acquisition steps are performed: querying the knowledge vector block corresponding to the query statement information from the material knowledge base; querying the first material knowledge corresponding to the knowledge vector block from the first knowledge graph corresponding to the recycled material field, wherein the first knowledge graph is a knowledge graph in the form of node vectors. In this way, professional material knowledge related to the query statement can be efficiently obtained based on the structured query path, and the knowledge corresponding to the recycled material can be quickly obtained. Then, using the second thread, the second material knowledge corresponding to the material information is queried from the second knowledge graph corresponding to the recycled material field, wherein the second knowledge graph is a knowledge graph in the form of text. Then, the target material knowledge corresponding to the output of the target thread is determined, wherein the target thread is the thread that ends earliest among the first thread and the second thread, and the target material knowledge is the knowledge generated earliest among the first material knowledge and the second material knowledge. In this way, the knowledge corresponding to the recycled material can be quickly obtained under the premise of effective recycled material knowledge. Finally, according to the target material knowledge, the multi-modal large language model is used to generate the material recycling report corresponding to the material information, so as to recycle the recycled material according to the material recycling report. In this way, a recycling recommendation report with practical value and professional depth can be automatically generated. In summary, through automatic identification of the recycled material image, combined with knowledge graph parallel retrieval, the knowledge related to the recycled material is quickly obtained, the recycling report is generated, and the subsequent material sorting, processing and recycling of practitioners are guided. BRIEF DESCRIPTION OF DRAWINGS

[0014] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by referring to the following detailed description in conjunction with the accompanying drawings. In the drawings, like reference numerals refer to like elements or features throughout. It should be understood that the drawings are schematic and elements and features are not necessarily to scale.

[0015] Figure 1 is a flowchart of some embodiments of the multi-modal large language model-based recycled material recycling method according to the present disclosure;

[0016] Figure 2 is a structural schematic diagram of some embodiments of the multi-modal large language model-based recycled material recycling device according to the present disclosure;

[0017] Figure 3 is a structural schematic diagram of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation

[0018] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0019] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0020] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0021] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0022] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0023] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0024] refer to Figure 1 The diagram illustrates a flow 100 of some embodiments of a recycling method based on a multimodal large language model according to the present disclosure. This recycling method based on a multimodal large language model includes the following steps:

[0025] Step 101: Using a pre-trained multimodal large language model, generate material information and query statement information corresponding to the target recycled material image.

[0026] In some embodiments, the execution subject (e.g., an electronic device) of the above-mentioned method for recycling materials based on a multi-modal large language model can be a server cluster deployed locally or in the cloud. Among them, the above-mentioned pre-trained multi-modal large language model can support a large language model that can process multi-modal data. In practice, the input of the multi-modal large language model can be multi-modal data, and the corresponding output can also be multi-modal data. Multi-modal data can include, but is not limited to, at least one of the following: video modal corresponding modal data, audio modal corresponding modal data, text modal corresponding modal data, and image modal corresponding modal data. The multi-modal large language model can directly process images and generate a large language model corresponding to the text.

[0027] The above-mentioned large language model can be a model trained based on a large amount of image text (e.g., GPT-4V and Gemini). The image text pair is a binary tuple data composed of a recyclable material image and corresponding structured annotation information. For example, the image text pair can be a binary tuple data composed of a brass scrap image and the text "brass, water yield rate about 95%, copper content about 65%, containing a small amount of plastic impurities".

[0028] The above-mentioned target recyclable material image can be a real photo of the material to be recycled, which has identifiable visual feature information (e.g., material, morphology, and type). The above-mentioned material information can be structured description information (e.g., material type, physical state, component identification, and pollution level) extracted from the image. The above-mentioned query statement information can be a query instruction automatically generated by the multi-modal large language model for accurate matching of knowledge base retrieval.

[0029] As an example, first, input the recyclable material image into the multi-modal large language model, use the multi-modal joint inference of the large model to extract and infer the features contained in the image, and obtain the feature vector of the material information. Finally, use the decoding part of the multi-modal large language model to decode the feature vector of the material information, and construct the query statement information to obtain the structured description of the material information and the query statement.

[0030] In some optional implementations of some embodiments, the multi-modal large language model is a large language model that is supervised fine-tuned and domain-aligned based on a multi-modal data set; and the above-mentioned multi-modal large language model can be trained by the following steps:

[0031] In a first step, an open-source multi-modal large language model is obtained as an initial multi-modal large language model. The open-source multi-modal large language model can be a publicly available source code multi-modal large model capable of processing image and text inputs simultaneously (e.g., LLaVA and MiniGPT-4). The initial multi-modal large language model can be an untrained and fine-tuned multi-modal large model. In practice, the multi-modal large model can be obtained from an open-source code repository (e.g., GitHub).

[0032] In a second step, a pre-constructed sample dataset is obtained, where the data in the sample dataset is in the form of material image-text pairs for supervised fine-tuning. The sample dataset can be a training dataset for supervised fine-tuning. The material image-text pairs can be in the form of (image, text), e.g., {“image”: “photo of scrap metal parts”, “text”: [“material”: “steel”, “purity”: “95%”, “process”: “electrolysis”]}. The supervised fine-tuning can refer to a process of further optimizing the performance of a pre-trained model using labeled data.

[0033] In a third step, a pre-constructed preference dataset is obtained, where the preference dataset includes material image data and material image text data, and the data in the preference dataset is used for domain alignment training. The preference dataset can be a dataset for alignment training. The material image data can be images in the sample dataset. The material image text data can be expert-preferred high-quality answers (positive examples) and ordinary answers (negative examples). In practice, domain experts (e.g., experienced recycling practitioners and quality inspectors) can be used to evaluate the recognition results corresponding to the images to obtain the material image text data. As an example, the recognition results can be obtained using a PPO (Proximal Policy Optimization) algorithm.

[0034] In a fourth step, the following model supervised fine-tuning training steps are performed based on the sample dataset:

[0035] In a first sub-step, at least one sample data in the sample dataset is input into the initial multi-modal large language model to obtain corresponding material information for the at least one sample data. The material information can be structured text content generated by the model for the sample data.

[0036] A second sub-step is to determine whether the initial multi-modal large language model reaches a preset optimization target according to a first target loss function. The optimization target can be a preset model training stop condition. For example, the optimization target can be that the cross-entropy loss function value is reduced to below a preset value (for example, 0.1). The first target loss function can be a cross-entropy loss function for measuring the difference between the text content of the model output and the standard answer text.

[0037] A third sub-step is to use the initial multi-modal large language model as a supervised fine-tuned multi-modal large language model in response to the initial multi-modal large language model reaching the optimization target.

[0038] Optionally, the execution subject can further perform the following steps:

[0039] A first step is to adjust the training parameters of the initial multi-modal large language model in response to the initial multi-modal large language model not reaching the optimization target, and use the adjusted initial multi-modal large language model as the initial multi-modal large language model to perform the model supervised fine-tuning training step again. The training parameters can include learning rate, training batch size, and maximum training rounds. In practice, the parameters can be modified according to the state of model training. For example, if the model training appears gradient shock, the learning rate is reduced.

[0040] A second step is to determine the supervised fine-tuned multi-modal large language model as an initial domain-aligned multi-modal large language model. The initial domain-aligned multi-modal large language model can be an initial model that is fine-tuned and then trained for domain alignment.

[0041] A third step is to perform the following model domain alignment step according to the preference data set:

[0042] A first sub-step is to input at least one regenerated material image data in the preference data set into the initial domain-aligned multi-modal large language model to obtain corresponding material information of the at least one regenerated material image data.

[0043] A second sub-step is to compare the corresponding material information of the at least one regenerated material image data with the corresponding regenerated material image text data to obtain a second target loss function. The second target loss function can be a comparison loss function for the preference data. The comparison loss function can be a function for measuring the matching degree of the model generated result and the domain expert preference. For example, DPO (Direct Preference Optimization).

[0044] A third sub-step, in response to whether the initial field-aligned multi-modal large language model meets the preset optimization target according to the second target loss function.

[0045] A fourth sub-step, in response to the initial field-aligned multi-modal large language model meeting the optimization target, taking the initial field-aligned multi-modal large language model as the field-aligned multi-modal large language model.

[0046] A fourth step, in response to the initial field-aligned multi-modal large language model not meeting the optimization target, adjusting the training parameters of the initial field-aligned multi-modal large language model, and using the adjusted initial field-aligned multi-modal large language model as the initial field-aligned multi-modal large language model, and executing the model field alignment step again.

[0047] Step 102, using a first thread to execute the following material knowledge acquisition steps:

[0048] Step 1021, querying the knowledge vector block corresponding to the query statement information from the material knowledge base.

[0049] In some embodiments, the execution subject can query the knowledge vector block corresponding to the query statement information from the material knowledge base. The material knowledge base can be a database that stores structured or semi-structured data related to materials (e.g., raw materials, parts). The database can be related information of the material, including attributes of the material (such as material, specifications, supplier information), use scenarios, processing technology, industry standards and specifications, and professional literature and publications. The knowledge vector block can be a vector representation of a knowledge fragment (e.g., a sentence, a paragraph, or a set of attributes) stored in the material knowledge base. For example, the query statement information can be a query vector converted by the same embedding model as the knowledge vector block. By determining the similarity between the query vector and the knowledge vector block in the knowledge base, the knowledge vector block corresponding to the query statement information is determined. The first thread can be a retrieval process based on semantic understanding. The natural language information is converted into a vector form, and a semantic match is performed in the first knowledge graph in the form of node vectors to obtain the corresponding material knowledge.

[0050] As an example, first, the query statement information is preprocessed and converted into a query vector. Second, the query vector is used for similarity retrieval in the vector database to find the most similar knowledge vector block to the query vector. Finally, the text information of the original knowledge corresponding to the knowledge vector block is returned.

[0051] Step 1022, querying the first material knowledge corresponding to the knowledge vector block from the first knowledge graph corresponding to the material field.

[0052] In some embodiments, the execution subject can query the first material knowledge corresponding to the knowledge vector block from the first knowledge graph in the field of renewable material. The first knowledge graph is a knowledge graph in the form of node vectors. The first knowledge graph can be a structured knowledge graph constructed in the field of renewable materials. The first knowledge graph can include entities (such as material types and attributes), relationships (such as “is a”, “has a property”, and “related contact”), and attributes (such as purity and recycling method). The first material knowledge is the structured knowledge result (such as text information content) retrieved from the first knowledge graph.

[0053] As an example, first, a nearest neighbor search algorithm (such as HNSW, Hierarchical Navigable Small World) can be performed to determine entities semantically similar to the knowledge vector block in the knowledge graph as a candidate entity set. Finally, the complete data record information of the candidate entity nodes and their associated edges in the candidate entity set is called through the graph database query interface as the output result of the target material knowledge.

[0054] Optionally, the material knowledge base and the first knowledge graph are cooperated through static anchor points and dynamic entity alignment; and the material knowledge base and the first knowledge graph are cooperated through the following steps:

[0055] First, the standard fields of the material knowledge in the material knowledge base are extracted. The standard fields can be predefined key attribute fields (such as material name, material number, and international standard number) for cross-knowledge base alignment. In practice, the predefined fields of each material knowledge in the material database can be extracted and recorded and stored as a set or a list.

[0056] Second, according to the standard fields, the graph nodes in the first knowledge graph are matched to obtain a mapping relationship. The graph nodes can be entity nodes in the knowledge graph with unique identifiers and attributes (such as material name and material number). The mapping relationship can be the correspondence between the records in the material knowledge base and the nodes in the knowledge graph. In practice, the standard fields can be used as query conditions to match the nodes in the knowledge graph through queries. The queried nodes are associated with the material knowledge records to form a mapping dictionary.

[0057] Third, in response to the mapping relationship being not empty, the material knowledge is bound to the graph nodes as static anchor points. The static anchor points can be binding relationships used to avoid repeated matching.

[0058] As an example, first, a new relationship edge is created between the material knowledge entity node and the corresponding graph node in the knowledge graph. Second, the material knowledge record is located in the material knowledge base, and a unique identifier pointing to the above graph node is added in the metadata. Finally, the identification information of the record in the knowledge base is filled into the newly created relationship edge to obtain the static anchor point of the record bidirectional relationship.

[0059] Fourthly, in response to the above mapping relationship being empty, the following dynamic entity alignment operation steps are performed:

[0060] First sub-step, determine the knowledge vector of the above material knowledge. Wherein, the knowledge vector can be a form of representation of the material knowledge in the vector space. In practice, the material knowledge can be input into an embedding model (for example, a pre-trained BERT model) to generate a knowledge vector.

[0061] Second sub-step, determine the similar node of the knowledge vector in the entity vector space of the first knowledge graph. Wherein, the entity vector space can be a set composed of the vector form representation of all graph node entities in the knowledge graph. The similar node can be a graph node close to the knowledge vector in mathematical distance (for example, Euclidean distance) or angle (for example, cosine similarity). In practice, the similar node can be obtained according to the cosine similarity of the determined knowledge vector and the entity vector corresponding to the graph node.

[0062] Third sub-step, determine the confidence score of the similar node. Wherein, the confidence score can be a reliability score (for example, cosine similarity) representing the entity alignment result. In practice, the confidence score can be obtained by determining the cosine similarity of each similar node and normalizing the similarity.

[0063] Fourth sub-step, in response to the above confidence score being greater than a preset threshold, the similar node and the above material knowledge are bound as a static anchor point. Wherein, the preset threshold can be the minimum value of the preset entity alignment confidence (for example, the preset threshold can be 0.8).

[0064] Step 103, using a second thread, querying the second material knowledge corresponding to the material information from the second knowledge graph corresponding to the second material field.

[0065] In some embodiments, the execution subject can query the second material knowledge corresponding to the material information from the second knowledge graph corresponding to the material field by using the second thread. The second knowledge graph is a text form knowledge graph. The second material knowledge can be the information content of the text file retrieved from the second knowledge graph. The second thread can be a structured data-based retrieval process. According to the content of the feature tag information in the known material information (for example, "category: copper, impurity content: 5%, main impurities: plastic"). Query the relationship in the text form second knowledge graph (for example, traverse the relationship edge "processing technology, risk, value" according to the "copper" entity node), and obtain the corresponding material knowledge.

[0066] As an example, first, the material information is preprocessed and the key attributes are extracted. And the key attributes are used to construct a query statement. Second, the text content in the knowledge graph is parsed, and a query graph in memory (for example, an RDF graph) is constructed. Finally, the query statement is executed in the query graph to obtain the material knowledge.

[0067] In the process of using the technical solutions to solve the above technical problems, the system often returns material knowledge with fast generation speed, but cannot guarantee the accuracy of the knowledge. At the same time, when a thread has knowledge errors, the other thread's knowledge graph also has the same errors. The generated report information is inaccurate, which makes it impossible to effectively guide the recycling process. To solve these problems, the conventional solution is to check the correctness of the knowledge by manual review, and to check the information source of the knowledge graph. However, the inventors consider adding manual review, which increases the report generation time and makes it impossible to quickly obtain the report to make recycling decisions. We decided to use the following solution:

[0068] Optionally, the execution subject can further perform the following steps:

[0069] First, send the second material knowledge to the intermediate cache area to perform the following third operation steps:

[0070] The first sub-step is to extract the minimum knowledge unit of the second material knowledge in the intermediate cache area to obtain knowledge information as a material knowledge set, wherein the material knowledge in the material knowledge set includes material entities and corresponding material attributes. The minimum knowledge unit can be a basic unit of knowledge in the form of a triple (entity, attribute, value). The knowledge information can be knowledge content in the form of a triple. The material knowledge set can be a set composed of material entities and corresponding material attributes. The intermediate cache area can be a temporary storage area for storing data to be processed. In practice, the intermediate cache area can be a memory cache (for example, Redis) with high throughput.

[0071] As an example, each knowledge entry in the second material knowledge can be parsed to extract entities and attributes to form a minimum knowledge unit. The minimum knowledge units are combined into a material knowledge set.

[0072] The second sub-step is to filter each material knowledge in the material knowledge set according to a preset material knowledge standard library to obtain conflict knowledge as a conflict knowledge set. The preset material knowledge standard library can be a verified material industry standard knowledge base. The conflict knowledge can be a knowledge unit that is inconsistent with the preset material knowledge standard library in the material knowledge set. The conflict knowledge set can be a set composed of conflict knowledge.

[0073] As an example, first, each material knowledge in the material knowledge set is traversed and queried in the preset material knowledge standard library to obtain corresponding standard knowledge. Then, the attributes of the material knowledge and the attributes in the standard library are compared, and if there is inconsistency, the conflict knowledge is marked. All conflict knowledge is collected to form a conflict knowledge set.

[0074] The third sub-step is to perform the following fourth operation step for each conflict knowledge in the conflict knowledge set:

[0075] Sub-step one, according to the material entity of the conflict knowledge, retrieve the corresponding standard knowledge from the preset material knowledge standard library. The standard knowledge can be the correct knowledge content corresponding to the conflict knowledge retrieved from the preset material knowledge standard library.

[0076] Sub-step two, extract the standard attributes in the standard knowledge. The standard attributes can be correct material characteristics (for example, correct material entities and material entity relationships) in the standard knowledge. The correct material characteristics can be material characteristics that have been correctly verified.

[0077] Sub-step three, determine the structural consistency score of the material attributes of the conflict knowledge and the standard attributes. The structural consistency score can be a measure of the structural similarity between the material attributes of the conflict knowledge and the standard attributes.

[0078] As an example, first, using the embedding model, the material attribute and the standard attribute of the conflict knowledge are converted into a feature vector. Finally, the cosine similarity of the two vectors is calculated. The cosine similarity is used as the structural consistency score.

[0079] Sub-step four, determine the confidence score corresponding to the conflict knowledge in the second knowledge graph. The confidence score can be the reliability score of the knowledge source when building the knowledge graph. In practice, the confidence score can be a value between 0 and 100. The higher the value, the higher the reliability of the standard knowledge source.

[0080] As an example, the confidence score of the conflict knowledge can be obtained from the metadata of the second knowledge graph.

[0081] Sub-step five, the structural consistency score and the confidence score are weighted and fused to obtain a comprehensive confidence score. The comprehensive confidence score can be a score that combines the structural consistency score and the confidence score to determine the reliability of the conflict knowledge. In practice, a weighted average method can be used.

[0082] Sub-step six, in response to the comprehensive confidence score being greater than a preset value, the node and relationship corresponding to the conflict knowledge in the second knowledge graph are corrected, and the conflict knowledge in the second material knowledge is corrected. The preset value can be a standard value for judging the error degree of the conflict knowledge. In practice, the standard value can be 0.8. The correction method can be to perform an update operation in the graph database to add the standard knowledge.

[0083] Sub-step seven, in response to the existence of the conflict knowledge in the first knowledge graph, the corresponding node and relationship in the first knowledge graph are updated.

[0084] As an example, the conflict knowledge can be queried in the first knowledge graph. If the conflict knowledge exists, the corresponding part of the first knowledge graph is corrected.

[0085] Sub-step eight, in response to the comprehensive confidence score not being greater than the preset value, the standard knowledge is added to the target material knowledge to generate reference annotations for the recycling report.

[0086] As an example, the standard knowledge is added to the target material knowledge, and is used as a reference annotation when generating a recycling report later.

[0087] The above operation steps, as one of the invention points of the present disclosure, solve the technical problems mentioned in the background art. Due to the dispersion of the knowledge of the recycled material, the relevant knowledge content is scattered in different platforms, documents and industry standards, and there is a lack of unified organization and collection. In practice, the conventional method increases the correctness of the artificial review of the knowledge, and increases the time of the report generation. The present disclosure designs a scheme for conflict detection and correction of the knowledge graph, uses intermediate caching, vector calculation, confidence evaluation and weighted decision mechanism to identify the conflict between the external material knowledge and the standard library. And according to the comprehensive confidence score, the knowledge graph with errors is selected for correction. Or keep the standard knowledge as part of the content of the report. Therefore, by solving the potential conflict in the knowledge acquisition process, the quality of the recycling report can be improved. At the same time, the conflict in the knowledge graph data is corrected, and the quality of the graph itself is improved. The practitioners can follow the recycling report that conforms to the industry standard to carry out the recycling work.

[0088] Step 104, determining the target material knowledge corresponding to the output of the target thread.

[0089] In some embodiments, the above execution subject can determine the target material knowledge corresponding to the output of the target thread. Wherein the target thread is the thread that ends the execution earliest among the first thread and the second thread, and the target material knowledge is the knowledge generated earliest among the first material knowledge and the second material knowledge. Wherein the target material knowledge can be the structured information of the material (for example, the composition, source, form, recycling method and flow scheme of the material).

[0090] In some optional implementations of some embodiments, the execution subject can determine the target material knowledge corresponding to the output of the target thread, which can include the following steps:

[0091] First, determine the path depth of the knowledge vector block in the first knowledge graph to obtain the vector confidence. Wherein the path depth can be the shortest number of hops from the target entity in the knowledge vector block to the root node in the knowledge graph. The vector confidence can be the knowledge reliability score according to the path depth. For example, in the scene of knowledge graph query, the shallower the path depth, the higher the confidence.

[0092] As an example, the breadth-first search can be used to determine the depth of the target entity in the knowledge vector block in the knowledge graph. And determine the confidence of the vector according to the depth size.

[0093] Second, determine the graph node corresponding to the knowledge vector block in the first knowledge graph to obtain the semantic similarity. Wherein the graph node can be the entity corresponding to the material attribute (for example, material type and material feature) in the knowledge graph. The semantic similarity can be the cosine similarity of the embedding vector of the knowledge vector block and the graph node.

[0094] A third step is to determine a graph relationship corresponding to the knowledge vector block in the first knowledge graph, and obtain a relationship similarity. The graph relationship can be a semantic relationship between two entities with types and attributes in the knowledge graph. The relationship similarity can be a matching degree of the relationship embedding in the knowledge vector block and the graph relationship. In practice, the cosine similarity of the relationship in the knowledge vector block and the graph relationship in the knowledge graph can be determined as the relationship similarity.

[0095] A fourth step is to determine a knowledge correctness score of the first material knowledge according to the vector confidence, the semantic similarity and the relationship similarity. The knowledge correctness score can be a weighted evaluation value of the vector confidence, the semantic similarity and the relationship similarity.

[0096] A fifth step is to perform the following first operation step in response to the knowledge correctness score being lower than a preset warning threshold:

[0097] A first sub-step is to send the first material knowledge to an intermediate cache area for knowledge verification. The intermediate cache area can be a queue for temporarily storing data to be verified.

[0098] A second sub-step is to add a to-be-verified mark to the first material knowledge, and take the first material knowledge as a target material knowledge to generate a recovery report. The to-be-verified mark can carry unique identification content (for example, JSON format content).

[0099] In the process of solving the above technical problems by adopting the technical solutions, the following problems often occur: when the secondary material knowledge from different channels is fused, the structures in each knowledge graph are different, and the fusion of the secondary material knowledge can cause inconsistency. At the same time, when different knowledge graphs interpret and infer the material, the obtained knowledge has a certain deviation. The deviated knowledge cannot be recorded. To solve these problems, the conventional solution is to correct by recording logs of the knowledge graph. The conflict data source is labeled for recording. However, the inventors consider that the conflict in the log record lacks structured positioning and cannot distinguish levels (for example, entities, attributes and relationships). The conflict data source is labeled, but cannot be accurately positioned to the entity node in the knowledge graph. We decided to use the following solution:

[0100] Optionally, the execution subject can further perform the following steps:

[0101] A first step is to send the first material knowledge to an intermediate cache area, and perform the following knowledge verification operation steps:

[0102] A first sub-step, obtaining the first material knowledge in the intermediate cache area. As an example, the first material knowledge to be verified can be popped from the intermediate cache area (e.g., Redis). And the first material knowledge is parsed into a structure processable by the program (e.g., a dictionary).

[0103] A second sub-step, feature extraction is performed on the first material knowledge to obtain first material feature information, wherein the first material feature information includes first material entities, first material semantic information, and first material relationships. The feature extraction can extract key components (e.g., entities, semantic information, and relationships) from the material knowledge. The first material feature information is the content information of the key components in the material knowledge. The first material entity can be the main entity object described in the knowledge (e.g., stainless steel 304). The first material semantic information can be the content describing the attribute state of the entity (density 7.93 g / cm 3 ). The first material relationship can be the relationship between the entity and other entities in the material knowledge (e.g., "recyclable" corresponding to "stainless steel ingot").

[0104] As an example, an entity recognition model (e.g., spaCy, a pre-trained NER model) can be used to identify entities in the material knowledge. Regular expressions can be used to extract text information from the material knowledge. A relationship extraction model (e.g., a pre-trained BERT model) can be used to obtain the relationship between entities.

[0105] A third sub-step, feature extraction is performed on the second material knowledge to obtain second material feature information, wherein the second material feature information includes second material entities, second material semantic information, and second material relationships.

[0106] A fourth sub-step, in response to the first material feature information and the second material feature information being inconsistent, the following second operation step is performed:

[0107] Sub-step one, in response to the first material entity and the second material entity being inconsistent, determining the conflict entity in the first material entity and the second material entity. The conflict entity can be the inconsistent part of the first material entity and the second material entity. For example, in the steel recycling scenario, the first material entity can be "stainless steel 304", and the second material entity can be "stainless steel 316". The conflict entity can be "stainless steel 304".

[0108] Sub-step two, in response to the inconsistency between the first material semantic information and the second material semantic information, determine the conflict semantic information in the first material semantic information and the second material semantic information. Wherein, the conflict semantic information can be the inconsistency of semantic description on the same entity or attribute in the material knowledge. For example, in the steel recycling scenario, the first material entity can be "stainless steel 304". The semantic information can be "the density is 7.93 g / cm3". The second material entity can be "stainless steel 304". The semantic information can be "the density is 7.99 g / cm3". The conflict semantic information can be "the density is 7.93 g / cm3".

[0109] Sub-step three, in response to the inconsistency between the first material relationship and the second material relationship, determine the conflict material relationship in the first material relationship and the second material relationship. Wherein, the conflict material relationship can be the different relationship between entities on the same entity in the material knowledge. For example, in the steel recycling scenario, the first material entity can be "stainless steel 304". The material relationship can be "recyclable", "stainless steel ingot". The second material entity can be "stainless steel 304". The material relationship can be "recyclable", "rebar". The conflict material relationship can be "stainless steel ingot" of the first material relationship and "rebar" of the second material relationship.

[0110] Sub-step four, determine the query path of the conflict entity in the first material knowledge as the conflict point. Wherein, the query path can be the query path of the conflict entity in the first material knowledge. The conflict point can be the query path from the root node to the conflict entity.

[0111] As an example, perform path query in the graph database corresponding to the knowledge graph. Determine the query path from the root node to the conflict entity.

[0112] Sub-step five, determine the knowledge source of the conflict semantic information in the first knowledge graph as the conflict information. Wherein, the knowledge source can be the knowledge data used when constructing the knowledge graph (for example, data set, journal and expert report). The conflict information can be the specific content of the knowledge data.

[0113] As an example, the knowledge graph construction record can be searched to query the source of the conflict semantic information (for example, text content, data source, update time and provider) in the first knowledge graph.

[0114] Sub-step six, determine the node relationship of the conflict material relationship in the first knowledge graph and the second knowledge graph as the abnormal relationship. Wherein, the node relationship can be the contact relationship between entities in the knowledge graph. The abnormal relationship can be the specific manifestation of the conflict material relationship in the two knowledge graphs (for example, different relationship types and different target entities).

[0115] As an example, the conflict material relationship can be queried from the graph database corresponding to the first knowledge graph and the second knowledge graph respectively. The entity type and the connected entity corresponding to the conflict material relationship are obtained.

[0116] Step seven, according to the above conflict point, the above conflict information and the above abnormal relationship, a conflict comparison result is constructed. Wherein, the conflict comparison result can be the structured information of the comprehensive conflict point, the conflict information and the abnormal relationship.

[0117] The fifth sub-step, according to the above conflict comparison result, the query chain of the first thread is analyzed, and an abnormal path is obtained. Wherein, the abnormal path can be the step path in the query chain that causes the conflict in the query process.

[0118] As an example, according to the query log, the step of query can be parsed by calling the chain tracking tool. And the conflict point is mapped to the specific step of the query chain to determine the query path.

[0119] The sixth sub-step, according to the above abnormal path, the material knowledge base and the first knowledge graph are optimized. Wherein, the structure optimization can be the reason for the abnormality to modify the knowledge base and the knowledge graph pruning operation. Wherein, the knowledge base modification can be to complete or delete the content in the knowledge base. The knowledge graph pruning operation can be to constrain and adjust the relationship of the corresponding entity in the graph according to the abnormal relationship.

[0120] The above operation steps, as one of the invention points of the present disclosure, solve the technical problems mentioned in the background art: "Due to the dispersion of the material knowledge, the relevant knowledge content is scattered in different platforms, documents and industry standards, and there is a lack of unified organization and collection." and "The conflict in the log record lacks structured positioning, and there is no distinction between levels (for example, entities, attributes and relationships). The conflict data source is labeled, but it cannot be accurately positioned to the entity node in the knowledge graph." In practice, the conventional method regards the data conflict as a part that can be isolated and processed, ignoring the relevance between the conflict data contents (for example, between entities and between entities and nodes). The present disclosure designs a deep knowledge conflict and correction scheme. After detecting the conflicts of the material knowledge from different sources, the conflicts are accurately classified through fine feature extraction (for example, entities, semantics and relationships). The path, source and specific performance of the conflict point in the knowledge graph are traced. According to the conflict comparison result, the query chain links (for example, abnormal path) that cause the conflict are analyzed, and the underlying knowledge base (for example, content addition and deletion) and the upper knowledge graph (for example, structure pruning and constraint) are actively corrected. Therefore, through the clear conflict content and conflict correction method, the maintenance of the knowledge base and the knowledge graph is targeted. It is more efficient to obtain material knowledge, so as to generate a high-quality recycling report to guide recycling. The efficiency of recycling material by practitioners is improved.

[0121] Optionally, the above execution subject can further perform the following steps:

[0122] First, extract the material entity name in the first material knowledge and the second material knowledge. The material entity name can be the unique identifier of the material.

[0123] As an example, the name of the entity can be extracted in the first knowledge graph using a query method. The material entity name in the text can be identified using natural language processing technology.

[0124] Second, according to the preset material entity standard table, the material entity name is converted into a coded standard material entity. The preset material entity standard table can be a correspondence mapping table that stores entity relationships and unique codes. The standard material entity can be a unique identifier after mapping conversion by the mapping table.

[0125] As an example, each material entity name can be traversed to find the corresponding code in the preset material entity standard table to obtain the standard material entity. In practice, first, the preset material entity standard table can be loaded into a dictionary structure. Second, use the material entity name as the key and the code as the value. Key query is performed on each material entity name in the dictionary to obtain the corresponding standard code. Finally, the standard code obtained is used to replace the material entity name to obtain the standard material entity.

[0126] Thirdly, determine the graph structure and the first attribute set of the standard material entity in the first knowledge graph. The graph structure can be the network structure of entity nodes and edge relationships in the knowledge graph. The first attribute set can be the standard entity attribute set extracted from the first knowledge graph (structured graph). For example, the standard material entity can be "motor". The first attribute can be "material: copper coil" and "weight: 5 kg".

[0127] As an example, first, in the knowledge graph, query the graph structure (e.g., nodes, edge relationships) centered on the standard material entity. Then, parse the attributes of the nodes from the graph structure. Then, represent the graph structure as a data structure (e.g., JSON and graph object). The attributes are stored in the form of a dictionary.

[0128] Fourthly, determine the text attribute structure and the second attribute set of the standard material entity in the second knowledge graph. The text attribute structure can be a set of unstructured text descriptions (e.g., JSON key-value pairs). The second attribute set can be the entity attribute set parsed from the second knowledge graph (text construction graph). For example, the standard material entity can be "motor". The second attribute can be "material: copper, parameter: weight equal to 5 kg".

[0129] As an example, first, locate the relevant records according to the standard material entity in the second knowledge graph to obtain the material entity content. Then, extract the entire attribute object in the material entity content. In practice, the JSON key-value tree can be parsed, and only the entire structure is retained. Secondly, the text attribute structure is used as the entity attribute. In practice, the key-value of all leaf nodes in the JSON can be extracted. Finally, the text attribute structure is stored as JSON, and the attribute set is stored in the form of a dictionary.

[0130] Fifthly, convert the text attribute structure into a text construction graph structure. The text construction graph structure can be a pseudo-graph structure converted from the text attribute.

[0131] As an example, the text attribute key can be used as a graph node. The text attribute value can be used as a leaf node. The edge can represent the text affiliation.

[0132] Sixthly, determine the structural similarity between the graph structure and the text construction graph according to the graph edit distance algorithm. The graph edit distance algorithm can be an algorithm for determining the difference between graph structures (e.g., the GED library in Python). In practice, the cost of modifying the content of the graph structure can be determined as the evaluation standard. The structural similarity can be a similarity score normalized by the graph edit distance.

[0133] As an example, first, the operation cost can be determined. For example, the default node / edge insertion / deletion cost is 1, and the replacement cost is 0.5. Second, the graph edit distance of the graph structure and the text construction graph is determined to obtain the total edit operation and the total cost. For example, the graph structure can be ["motor -> (material) -> copper coil", "motor -> (weight) -> 5kg"]. The text construction graph can be ["motor -> (material) -> copper", "motor -> (parameter) -> weight = 5kg"]. The replacement cost of the node "copper" in the text construction graph for "copper coil" can be 0.5. The replacement cost of the edge "material" in the text construction graph for "material" can be 0.5. The insertion cost of the missing edge "weight" in the text construction graph can be 1. The deletion cost of the node "weight = 5kg" in the text construction graph can be 1. The total edit operation can be 0.5 + 0.5 + 1 + 1 = 3.0. The edit operation of the graph structure can be 3 node replacements and 2 edge replacements, and the graph structure cost is 3 + 2 = 5. The edit operation of the text construction graph can be 3 node replacements and 2 edge replacements. The text construction graph cost is 3 + 2 = 5. The total cost is 5 + 5 = 10.0. Finally, the total edit operation is normalized according to the total cost to obtain the structural similarity. For example, the structural similarity can be 1 - (3.0 / 10.0) = 0.7.

[0134] In the seventh step, in response to the structural similarity being greater than the preset threshold, the merging of the first attribute set and the second attribute set is performed to obtain a structural alignment result. The attribute merging can be merging the entity attribute content obtained from the knowledge graph. In practice, the dictionary merging can be used to merge the contents in the first attribute set and the second attribute set. The structural alignment result can be a description of the merged knowledge graph after alignment. In practice, it can be a new knowledge graph or JSON.

[0135] In the eighth step, in response to the structural similarity being not greater than the preset threshold, the following fifth operation step is performed:

[0136] In the first sub-step, the difference structure of the nodes and edges in the graph structure and the text construction graph is identified according to the graph edit distance algorithm to obtain a difference tuple. The difference tuple can be a triple (operation type, element type, element ID) recording the structural difference.

[0137] As an example, first, the output edit operation sequence can be obtained using the graph edit distance algorithm. Finally, each operation sequence is traversed to create a triple.

[0138] In the second sub-step, the corresponding attributes in the first attribute set and the second attribute set are stored in the local memory for knowledge graph structure repair. The local memory can be a local storage medium (for example, a cache or a file in RAM).

[0139] As an example, traverse the difference tuples, extract the corresponding attributes in the first attribute set and the second attribute set for each element ID, and store them in the cache or file in the RAM.

[0140] In step 105, according to the target material knowledge, a material recycling report corresponding to the material information is generated by using a multi-modal large language model, and the recycled material is recycled according to the material recycling report.

[0141] In some embodiments, the above execution subject can generate a material recycling report corresponding to the above material information according to the above target material knowledge by using the above multi-modal large language model, and recycle the recycled material according to the above material recycling report. Wherein, the material recycling report can be a text file detailing the recycling suggestion of the recycled material, including the recycling scheme suggestion, the recycling precautions and the circulation scheme of the material. The recycling of the recycled material can be the process of actual recycling operation of the recycled material by the relevant personnel according to the content of the material recycling report.

[0142] As an example, first, a natural language instruction template can be defined. For example, "You are a material recycling expert, please generate a recycling report based on the following material attributes" + target material knowledge. Secondly, according to the target material knowledge, the natural language instruction template is supplemented and completed, and is used as a report generation prompt word. Finally, the report generation prompt word is input into the multi-modal large language model to generate the material recycling report.

[0143] The above various embodiments of the present disclosure have the following beneficial effects: through the recycled material recycling method based on the multi-modal large language model of some embodiments of the present disclosure, the recycled material image can be automatically identified, the structured and unstructured knowledge graph is combined to realize fast knowledge retrieval, and a recycling report is automatically generated. In the report generation process, the traditional method often needs to manually identify and classify the material first, and then search and process the specification and industry information through multiple platforms, which takes a long time and cannot quickly generate a recycling report with practical guiding significance according to the material image. Based on this, the recycled material recycling method based on the multi-modal large language model of some embodiments of the present disclosure first generates the material information and query statement information corresponding to the target recycled material image by using the pre-trained multi-modal large language model. In this way, automatic perception and semantic understanding of the recycled material image can be realized, reducing the dependence on manual identification and providing a structured input basis for subsequent knowledge retrieval. Secondly, by using the first thread, the following material knowledge acquisition steps are executed: querying the knowledge vector block corresponding to the query statement information from the material knowledge base; querying the first material knowledge corresponding to the knowledge vector block from the first knowledge graph corresponding to the recycled material field, wherein the first knowledge graph is a knowledge graph in the form of node vectors. In this way, professional material knowledge related to the query statement can be efficiently acquired based on the structured query path, and the knowledge corresponding to the recycled material can be quickly obtained. Then, by using the second thread, the second material knowledge corresponding to the material information is queried from the second knowledge graph corresponding to the recycled material field, wherein the second knowledge graph is a knowledge graph in the form of text. Then, the target material knowledge corresponding to the output of the target thread is determined, wherein the target thread is the thread that ends earliest among the first thread and the second thread, and the target material knowledge is the knowledge generated earliest among the first material knowledge and the second material knowledge. In this way, the knowledge corresponding to the recycled material can be quickly acquired on the premise that the recycled material knowledge is effective. Finally, according to the target material knowledge, the material recycling report corresponding to the material information is generated by using the multi-modal large language model, so as to recycle the recycled material according to the material recycling report. In this way, a recycling recommendation report with practical value and professional depth can be automatically generated. In summary, through automatic identification of the recycled material image, parallel retrieval of the knowledge graph, quick acquisition of the knowledge related to the recycled material, generation of the recycling report, and guidance of the subsequent material sorting, processing and recycling of the practitioners.

[0144] Further reference Figure 2 , as an implementation of the method shown in the above figures, the present disclosure provides some embodiments of a recycled material recycling device based on a multi-modal large language model, which correspond to the method embodiments shown in Figure 1 , the recycled material recycling device based on the multi-modal large language model can be applied to various electronic devices.

[0145] As shown in Figure 2 The multi-modal large language model-based recyclable material recycling apparatus 200 includes an acquisition unit 201, a first thread unit 202, a second thread unit 203, a knowledge generation unit 204, and a report generation unit 205. The acquisition unit 201 is configured to generate material information corresponding to a target recyclable material image and query statement information by using a pre-trained multi-modal large language model. The first thread unit 202 is configured to execute the following material knowledge acquisition steps by using a first thread: querying a knowledge vector block corresponding to the query statement information from a material knowledge base; and querying first material knowledge corresponding to the knowledge vector block from a first knowledge graph corresponding to the recyclable material field, wherein the first knowledge graph is a knowledge graph in the form of node vectors. The second thread unit 203 is configured to query second material knowledge corresponding to the material information from a second knowledge graph corresponding to the recyclable material field by using a second thread, wherein the second knowledge graph is a knowledge graph in the form of text. The knowledge generation unit 204 is configured to determine target material knowledge corresponding to the output of a target thread, wherein the target thread is a thread that ends execution earliest among the first thread and the second thread, and the target material knowledge is the knowledge that is generated earliest among the first material knowledge and the second material knowledge. The report generation unit 205 is configured to generate a material recycling report corresponding to the material information by using the multi-modal large language model according to the target material knowledge, so as to recycle the recyclable material according to the material recycling report.

[0146] It can be understood that the units described in the multi-modal large language model-based recyclable material recycling apparatus 200 correspond to the respective steps in the method described with reference to Figure 1 Therefore, the operations, features, and beneficial effects described above with respect to the method also apply to the multi-modal large language model-based recyclable material recycling apparatus 200 and the units included therein, and will not be described here again.

[0147] Reference is made below to Figure 3 which shows a structural schematic diagram of an electronic device (e.g., an electronic device) 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is merely an example and should not impose any limitation on the function and scope of use of the embodiments of the present disclosure.

[0148] As shown in Figure 3As shown, the electronic device 300 can include a processing device (e.g., a central processor, a graphics processor, etc.) 301 that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 302 or loaded into a random access memory (RAM) 303 from a storage device 308. Various programs and data required for the operation of the electronic device 300 are also stored in the RAM 303. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0149] Generally, the following devices can be connected to the I / O interface 305: input devices 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 308 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 309. The communication devices 309 can allow the electronic device 300 to communicate wirelessly or wired with other devices to exchange data. Although Figure 3 The electronic device 300 is shown with various devices, but it should be understood that all of the illustrated devices are not required, and more or fewer devices can alternatively be implemented. Figure 3 Each block shown in the flowcharts can represent a device, or multiple devices, as necessary.

[0150] In particular, processes described above with reference to the flowcharts can be implemented as a computer software program according to some embodiments of the present disclosure. For example, some embodiments of the present disclosure include a computer program product including a computer program carried on a computer readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In some such embodiments, the computer program can be downloaded and installed from a network through the communication devices 309, or installed from the storage devices 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above-described functions defined in the methods of some embodiments of the present disclosure are performed.

[0151] Note that the computer-readable medium or media used to provide the computer program sequence to the computer system can be accompanied by information sufficient to load the program sequence into an internal memory of the computer system. Generally, such information can be stored in one or more computer system memories or data storage systems from which, upon an execution by the computer system of one or more computer-readable programs, a processor of the computer system can retrieve. Based on these inputs, the computer system is programmed to perform particular tasks according to the instructions of the computer program contained in the computer-readable medium(s).

[0152] In some embodiments, the client, server, and / or other components can communicate information via a communication network. The communication network can be any public or private communication network, including the Internet, intranet, extranet, or any other type of network. In some embodiments, the client, server, and / or other components can communicate information via a communication network using any current or future developed network protocol, including HTTP (HyperText Transfer Protocol), SMTP (Simple Mail Transfer Protocol), FTP (File Transfer Protocol), and / or any other protocol.

[0153] The computer readable medium can be included in the electronic device; or can exist independently of the electronic device. The computer readable medium carries one or more programs, when the one or more programs are executed by the electronic device, the electronic device is caused to: generate material information and query statement information corresponding to a target renewable material image by using a pre-trained multi-modal large language model; perform the following material knowledge acquisition steps by using a first thread: query a knowledge vector block corresponding to the query statement information from a material knowledge base; query first material knowledge corresponding to the knowledge vector block from a first knowledge graph corresponding to a renewable material field, wherein the first knowledge graph is a knowledge graph in the form of node vectors; query second material knowledge corresponding to the material information from a second knowledge graph corresponding to the renewable material field by using a second thread, wherein the second knowledge graph is a knowledge graph in the form of text; determine target material knowledge corresponding to output of a target thread, wherein the target thread is a thread that ends execution earliest among the first thread and the second thread, and the target material knowledge is knowledge that is generated earliest among the first material knowledge and the second material knowledge; and generate a material recycling report corresponding to the material information by using the multi-modal large language model according to the target material knowledge, so as to recycle renewable materials according to the material recycling report.

[0154] Computer program code for carrying out operations of some embodiments of the disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0155] The flow and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow and block diagrams can represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may be executed in the reverse order, depending on the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or combinations of hardware and software.

[0156] The units described in some embodiments of the present disclosure can be implemented in the form of software, or can be implemented in the form of hardware. The described units can also be arranged in a processor, for example, it can be described that: a processor includes an acquisition unit, a first thread unit, a second thread unit, a knowledge generation unit and a report generation unit. Among them, the names of these units do not constitute a limitation to the units themselves in some cases, for example, the acquisition unit can also be described as "a unit that generates material information and query statement information corresponding to the target analog material image by using a pre-trained multi-modal large language model".

[0157] The functions described above in the present document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, example types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.

[0158] The above description is merely some of the preferred embodiments of the present disclosure and a description of the principles of the technology used. Those skilled in the art should understand that the scope of the application involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or equivalent features without departing from the above inventive concept. For example, the above features are replaced with the technical features disclosed in the embodiments of the present disclosure (but not limited to) having similar functions to form technical solutions.

Claims

1. A method for recycling recycled materials based on a multimodal large language model, comprising: Using a pre-trained multimodal large language model, material information and query statement information corresponding to the target recycled material image are generated; Using the first thread, perform the following material knowledge acquisition steps: Retrieve the knowledge vector block corresponding to the query statement from the material knowledge base; The first material knowledge corresponding to the knowledge vector block is queried from the first knowledge graph corresponding to the field of recycled materials, wherein the first knowledge graph is a knowledge graph in the form of node vectors; Using a second thread, the second material knowledge corresponding to the material information is queried from the second knowledge graph corresponding to the field of recycled materials, wherein the second knowledge graph is a text-based knowledge graph; Determine the target material knowledge output corresponding to the target thread, wherein the target thread is the thread that finishes execution earliest among the first thread and the second thread, and the target material knowledge is the knowledge generated earliest among the first material knowledge and the second material knowledge; Based on the target material knowledge, the multimodal large language model is used to generate a material recycling report corresponding to the material information, so as to recycle recycled materials according to the material recycling report; The multimodal large language model is a large language model based on a multimodal dataset after supervised fine-tuning and domain alignment; and the multimodal large language model is trained through the following steps: Obtain an open-source multimodal large language model as an initial multimodal large language model; Obtain a pre-constructed sample dataset, wherein the data in the sample dataset is in the form of images and text pairs of recycled materials, for the purpose of the supervised fine-tuning training; Obtain a pre-constructed preference dataset, wherein the preference dataset includes recycled material image data and recycled material image text data, and the data in the preference dataset is used for the domain alignment training; Based on the sample dataset, perform the following supervised fine-tuning training steps: Input at least one sample data from the sample dataset into the initial multimodal large language model to obtain the material information corresponding to at least one sample data; The initial multimodal large language model is determined based on the first objective loss function to determine whether it has achieved the preset optimization objective. In response to the initial multimodal large language model achieving the optimization objective, the initial multimodal large language model is used as the multimodal large language model completed under supervised fine-tuning.

2. The method according to claim 1, wherein, The material knowledge base and the first knowledge graph collaborate through static anchor points and dynamic entity alignment; and the material knowledge base and the first knowledge graph collaborate through the following steps: Extract the standard fields of material knowledge from the material knowledge base; Based on the standard fields, the graph nodes in the first knowledge graph are matched to obtain the mapping relationship; In response to the mapping relationship being non-empty, the material knowledge is bound to the graph node as a static anchor point; In response to the mapping being empty, the following dynamic entity alignment operation steps are performed: Determine the knowledge vector of the material knowledge; In the entity vector space of the first knowledge graph, determine the similar nodes of the knowledge vector; Determine the confidence score of the similar nodes; In response to the confidence score being greater than a preset threshold, the similar nodes and the material knowledge are bound together as static anchor points.

3. The method according to claim 1, wherein, The determination of the target material knowledge corresponding to the output of the target thread includes: Determine the path depth of the knowledge vector block in the first knowledge graph to obtain the vector confidence. Determine the graph node corresponding to the knowledge vector block in the first knowledge graph to obtain semantic similarity; Determine the graph relationship corresponding to the knowledge vector block in the first knowledge graph, and obtain the relationship similarity; The knowledge correctness score of the first material knowledge is determined based on the vector confidence, the semantic similarity, and the relational similarity. In response to the knowledge correctness score being lower than a preset warning threshold, the following first operation step is performed: The first material knowledge is sent to the intermediate cache for knowledge verification. The first material knowledge is marked with a verification tag, and the first material knowledge is used as the target material knowledge to generate a recycling report.

4. The method according to claim 1, wherein, The method further includes: Extract the material entity names from the first material knowledge and the second material knowledge; Based on the preset material entity standard table, the material entity name is converted into a standard material entity in coded form; Determine the graph structure and first attribute set of the standard material entity in the first knowledge graph; Determine the text attribute structure and second attribute set of the standard material entity in the second knowledge graph; Transform the text attribute structure into a text construction graph structure; The structural similarity between the graph structure and the text-constructed graph is determined based on the graph edit distance algorithm. In response to the structural similarity being greater than a preset threshold, the first attribute set and the second attribute set are merged to obtain a structural alignment result; In response to the fact that the structural similarity is not greater than a preset threshold, the following fifth operation step is performed: Based on the graph edit distance algorithm, the difference structure of nodes and edges in the graph structure and the text construction graph is identified to obtain difference tuples; Based on the difference tuples, the corresponding attributes in the first attribute set and the second attribute set are stored in local memory for knowledge graph structure repair.

5. The method according to claim 4, wherein, The method further includes: In response to the initial multimodal large language model failing to achieve the optimization objective, the training parameters of the initial multimodal large language model are adjusted, and the adjusted initial multimodal large language model is used as the initial multimodal large language model to perform the model-supervised fine-tuning training step again. The supervised fine-tuned multimodal large language model is selected as the initial domain-aligned multimodal large language model. Based on the preference dataset, perform the following model domain alignment steps: Input at least one recycled material image data from the preference dataset into the initial domain-aligned multimodal large language model to obtain the material information corresponding to at least one recycled material image data; In response to determining whether the initial domain-aligned multimodal large language model has achieved the preset optimization objective based on the second objective loss function; In response to the initial domain-aligned multimodal large language model achieving the optimization objective, the initial domain-aligned multimodal large language model is taken as the domain-aligned multimodal large language model. In response to the initial domain-aligned multimodal large language model failing to achieve the optimization objective, the training parameters of the initial domain-aligned multimodal large language model are adjusted, and the adjusted initial domain-aligned multimodal large language model is used as the initial domain-aligned multimodal large language model to perform the model domain alignment step again.

6. A recyclable material recovery device based on a multimodal large language model, comprising: The acquisition unit is configured to use a pre-trained multimodal large language model to generate material information and query statement information corresponding to the target recycled material image; The first thread unit is configured to perform the following material knowledge acquisition steps using the first thread: Retrieve the knowledge vector block corresponding to the query statement from the material knowledge base; The first material knowledge corresponding to the knowledge vector block is queried from the first knowledge graph corresponding to the field of recycled materials, wherein the first knowledge graph is a knowledge graph in the form of node vectors; The second thread unit is configured to use the second thread to query the second material knowledge corresponding to the material information from the second knowledge graph corresponding to the recycled material domain, wherein the second knowledge graph is a text-based knowledge graph; The knowledge generation unit is configured to determine the target material knowledge output by the target thread, wherein the target thread is the thread that finishes execution earliest among the first thread and the second thread, and the target material knowledge is the knowledge generated earliest among the first material knowledge and the second material knowledge. The report generation unit is configured to generate a material recycling report corresponding to the material information based on the target material knowledge and using the multimodal large language model, so as to recycle recycled materials based on the material recycling report; The multimodal large language model is a large language model based on a multimodal dataset after supervised fine-tuning and domain alignment; and the multimodal large language model is trained through the following steps: Obtain an open-source multimodal large language model as an initial multimodal large language model; Obtain a pre-constructed sample dataset, wherein the data in the sample dataset is in the form of images and text pairs of recycled materials, for the purpose of the supervised fine-tuning training; Obtain a pre-constructed preference dataset, wherein the preference dataset includes recycled material image data and recycled material image text data, and the data in the preference dataset is used for the domain alignment training; Based on the sample dataset, perform the following supervised fine-tuning training steps: Input at least one sample data from the sample dataset into the initial multimodal large language model to obtain the material information corresponding to at least one sample data; The initial multimodal large language model is determined based on the first objective loss function to determine whether it has achieved the preset optimization objective. In response to the initial multimodal large language model achieving the optimization objective, the initial multimodal large language model is used as the multimodal large language model completed under supervised fine-tuning.

7. An electronic device, comprising: One or more processors; Storage device, on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-5.

8. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Auxiliary retrieval method fusing knowledge graph and large language model

    CN117633252A

  • KR20250064065A