Regenerated material recovery method, device and equipment based on multi-modal large language model
By combining a multimodal large language model with a knowledge graph, we can automatically identify images of recycled materials and generate recycling reports, solving the problem of low report generation efficiency caused by the dispersion of recycled material knowledge and achieving fast and accurate recycling report generation.
Patent Information
- Application Number
- CN202511299644.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-12
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-09-12
AI Technical Summary
In the existing technology, knowledge about recycled materials is scattered across different platforms and documents, lacking a unified organization. This results in low efficiency in generating recycled material recycling reports and an inability to quickly process recycling needs.
A method based on a multimodal large language model is used, combined with knowledge graphs in node vector form and text form, to automatically identify recycled material images and generate material information. Structured and unstructured knowledge is acquired through parallel threads to automatically generate recycling reports.
It realizes automatic recognition of recycled material images and rapid knowledge retrieval, generates recycling reports with practical guidance, reduces reliance on manual recognition and search, and improves report generation efficiency and accuracy.
Smart Images

Figure CN120804344A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure relate to the field of computer technology, and in particular, to a method, device and equipment for recycling of secondary materials based on a multi-modal large language model. BACKGROUND
[0002] At present, with the global promotion of sustainable development and circular economy, secondary materials (including waste metals, plastics, paper, electronic components, construction waste and various types) have become efficient recycling and value maximization resources. To recycle diversified secondary materials, detailed and accurate recycling reports are needed. For the generation of secondary material recycling reports, the commonly used method is to rely on manual review and writing.
[0003] However, when the above method is used to generate the recycling report of secondary materials, the following technical problems often exist: Due to the dispersion of secondary material knowledge, the relevant knowledge content is scattered in different platforms, documents and industry specifications, and lacks unified organization and collection. The staff spends a long time in the search process and cannot quickly obtain the corresponding knowledge of the secondary materials. This leads to low efficiency of report generation and difficulty in timely processing of secondary material recycling needs.
[0004] The above information disclosed in the background section is only intended to enhance the understanding of the background of the present inventive concept, and therefore, it can include information that does not form the prior art known to those of ordinary skill in the art in the country. SUMMARY
[0005] The summary section of the present disclosure is used to introduce the concepts in a brief form, which will be described in detail in the specific embodiments section. The summary section of the present disclosure is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.
[0006] Some embodiments of the present disclosure propose a method, device and equipment for recycling of secondary materials based on a multi-modal large language model to solve one or more of the technical problems mentioned in the background section.
[0007] In a first aspect, some embodiments of the present disclosure provide a method for recycling of recycled materials based on a multi-modal large language model, comprising: generating, by using a pre-trained multi-modal large language model, material information corresponding to a target recycled material image and query sentence information; performing, by using a first thread, the following material knowledge acquisition steps: querying a knowledge vector block corresponding to the query sentence information from a material knowledge base; querying first material knowledge corresponding to the knowledge vector block from a first knowledge graph corresponding to a recycled material field, wherein the first knowledge graph is a knowledge graph in the form of node vectors; querying, by using a second thread, second material knowledge corresponding to the material information from a second knowledge graph corresponding to the recycled material field, wherein the second knowledge graph is a knowledge graph in the form of text; determining target material knowledge corresponding to an output of a target thread, wherein the target thread is a thread that ends execution earliest among the first thread and the second thread, and the target material knowledge is knowledge that is generated earliest among the first material knowledge and the second material knowledge; and generating, by using the multi-modal large language model, a material recycling report corresponding to the material information according to the target material knowledge, so as to recycle the recycled material according to the material recycling report.
[0008] In a second aspect, some embodiments of the present disclosure provide an apparatus for recycling of recycled materials based on a multi-modal large language model, comprising: an acquisition unit configured to generate, by using a pre-trained multi-modal large language model, material information corresponding to a target recycled material image and query sentence information; a first thread unit configured to perform, by using a first thread, the following material knowledge acquisition steps: querying a knowledge vector block corresponding to the query sentence information from a material knowledge base; querying first material knowledge corresponding to the knowledge vector block from a first knowledge graph corresponding to a recycled material field, wherein the first knowledge graph is a knowledge graph in the form of node vectors; a second thread unit configured to query, by using a second thread, second material knowledge corresponding to the material information from a second knowledge graph corresponding to the recycled material field, wherein the second knowledge graph is a knowledge graph in the form of text; a knowledge generation unit configured to determine target material knowledge corresponding to an output of a target thread, wherein the target thread is a thread that ends execution earliest among the first thread and the second thread, and the target material knowledge is knowledge that is generated earliest among the first material knowledge and the second material knowledge; and a report generation unit configured to generate, by using the multi-modal large language model, a material recycling report corresponding to the material information according to the target material knowledge, so as to recycle the recycled material according to the material recycling report.
[0009] In a third aspect, some embodiments of the present disclosure provide an electronic device, comprising: one or more processors; a storage device having one or more programs stored thereon, when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any implementation manner of the first aspect.
[0010] In a fourth aspect, some embodiments of the present disclosure provide a computer readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method as described in any implementation manner of the first aspect.
[0011] The above various embodiments of the present disclosure have the following beneficial effects: through the multi-modal large language model-based recycled material recycling method of some embodiments of the present disclosure, the recycled material image can be automatically identified, the structured and unstructured knowledge graph is combined to realize fast knowledge retrieval, and a recycling report is automatically generated. In the report generation process, the traditional method often needs manual identification and classification of the material first, and then searches for processing specifications and industry information through multiple platforms, which takes a long time and cannot quickly generate a recycling report with practical guiding significance according to the material image.
[0012] Based on this, the multi-modal large language model-based recycled material recycling method of some embodiments of the present disclosure first generates the material information and query statement information corresponding to the target recycled material image by using the pre-trained multi-modal large language model. In this way, automatic perception and semantic understanding of the recycled material image can be achieved, reducing the dependence on manual identification and providing a structured input basis for subsequent knowledge retrieval. Second, using the first thread, the following material knowledge acquisition steps are performed: querying the knowledge vector block corresponding to the query statement information from the material knowledge base; querying the first material knowledge corresponding to the knowledge vector block from the first knowledge graph corresponding to the recycled material field, wherein the first knowledge graph is a knowledge graph in the form of node vectors. In this way, professional material knowledge related to the query statement can be efficiently obtained based on the structured query path, and the knowledge corresponding to the recycled material can be quickly obtained. Then, using the second thread, the second material knowledge corresponding to the material information is queried from the second knowledge graph corresponding to the recycled material field, wherein the second knowledge graph is a knowledge graph in the form of text. Then, the target material knowledge corresponding to the output of the target thread is determined, wherein the target thread is the thread that ends earliest among the first thread and the second thread, and the target material knowledge is the knowledge generated earliest among the first material knowledge and the second material knowledge. In this way, the knowledge corresponding to the recycled material can be quickly obtained under the premise of effective recycled material knowledge. Finally, according to the target material knowledge, the multi-modal large language model is used to generate the material recycling report corresponding to the material information, so as to recycle the recycled material according to the material recycling report. In this way, a recycling recommendation report with practical value and professional depth can be automatically generated. In summary, through automatic identification of the recycled material image, combined with knowledge graph parallel retrieval, the knowledge related to the recycled material is quickly obtained, the recycling report is generated, and the subsequent material sorting, processing and recycling of practitioners are guided. BRIEF DESCRIPTION OF DRAWINGS
[0013] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by referring to the following detailed description in conjunction with the accompanying drawings. In the drawings, like reference numerals refer to like elements or features throughout. It should be understood that the drawings are schematic and elements and features are not necessarily to scale.
[0014] Figure 1 is a flowchart of some embodiments of the multi-modal large language model-based recycled material recycling method according to the present disclosure; Figure 2 is a structural schematic diagram of some embodiments of the multi-modal large language model-based recycled material recycling device according to the present disclosure; Figure 3 is a structural schematic diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0015] Embodiments of the present disclosure will be described below in greater detail with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure will be more thoroughly and completely understood. It should be understood that the drawings of the present disclosure are only for illustrative purposes and are not intended to limit the scope of protection of the present disclosure.
[0016] In addition, it should be further noted that only parts related to the present application are shown in the drawings for ease of description. The embodiments in the present disclosure and the features in the embodiments can be combined with each other without conflict.
[0017] It should be noted that the terms "first", "second", and the like mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.
[0018] It should be noted that the adjectives "one", "multiple" mentioned in the present disclosure are illustrative and not limiting, and those skilled in the art should understand that unless otherwise explicitly stated in the context, it should be understood as "one or more".
[0019] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0020] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.
[0021] Reference Figure 1 , shows the flow 100 of some embodiments of the multi-modal large language model-based recycled material recycling method according to the present disclosure. The multi-modal large language model-based recycled material recycling method includes the following steps: Step 101, using a pre-trained multi-modal large language model, generating material information and query statement information corresponding to a target recycled material image.
[0022] In some embodiments, the execution subject (e.g., an electronic device) of the above-mentioned multi-modal large language model-based secondary material recycling method can be a server cluster deployed locally or in the cloud. Among them, the above-mentioned pre-trained multi-modal large language model can support a large language model that can process multi-modal data. In practice, the input of the multi-modal large language model can be multi-modal data, and the corresponding output can also be multi-modal data. Multi-modal data can include, but is not limited to, at least one of the following: video modal corresponding modal data, audio modal corresponding modal data, text modal corresponding modal data, and image modal corresponding modal data. The multi-modal large language model can directly process images and generate text corresponding large language models.
[0023] The above-mentioned large language model can be a model trained based on massive image texts (e.g., GPT-4V and Gemini). The image-text pair is a binary tuple data composed of a secondary material image and corresponding structured annotation information. For example, the image-text pair can be a binary tuple data composed of a brass scrap image and the text "brass, water yield rate about 95%, copper content about 65%, containing a small amount of plastic impurities".
[0024] The above-mentioned target secondary material image can be a real photo of the material to be recycled, which has identifiable visual feature information (e.g., material, morphology, and type). The above-mentioned material information can be structured description information (e.g., material type, physical state, component identification, and pollution degree) extracted from the image. The above-mentioned query statement information can be a query instruction automatically generated by the multi-modal large language model for accurate matching of knowledge base retrieval.
[0025] As an example, first, input the secondary material image into the multi-modal large language model, use the multi-modal joint inference of the large model to extract and infer the features contained in the image, and obtain the feature vector of the material information. Finally, use the decoding part of the multi-modal large language model to decode the feature vector of the material information, and construct the query statement information to obtain the structured description of the material information and the query statement.
[0026] In some optional implementations of some embodiments, the multi-modal large language model is a large language model that is supervised fine-tuned and domain-aligned based on a multi-modal data set; and the above-mentioned multi-modal large language model can be trained by the following steps: First, obtain an open source multi-modal large language model as an initial multi-modal large language model. Among them, the above-mentioned open source multi-modal large language model can be a multi-modal large model (e.g., LLaVA and MiniGPT-4) with public source code that can process image and text inputs simultaneously. The above-mentioned initial multi-modal large language model can be a multi-modal large model that has not been trained and fine-tuned. In practice, the multi-modal large model can be obtained from an open source code repository (e.g., GitHub).
[0027] In the second step, a pre-constructed sample dataset is obtained, wherein the data in the sample dataset is in the form of a material image-text pair, and is used for the above-mentioned supervised fine-tuning training. The sample dataset can be a training dataset for supervised fine-tuning. The material image-text pair can be in the form of (image, text), for example, {“image”: “a photo of a scrap metal part”, “text”: [“material”: “steel”, “purity”: “95%”, “process”: “electrolysis”]}. The supervised fine-tuning can refer to a process of further optimizing the performance of a pre-trained model by using labeled data.
[0028] In the third step, a pre-constructed preference dataset is obtained, wherein the preference dataset includes material image data and material image text data, and the data in the preference dataset is used for the above-mentioned domain alignment training. The preference dataset can be a dataset for alignment training. The material image data can be the image in the sample dataset. The material image text data can be expert-preferred high-quality answers (positive examples) and ordinary answers (negative examples). In practice, domain experts (e.g., experienced recycling practitioners and quality inspectors) can be used to judge the recognition results corresponding to the images to obtain the material image text data. As an example, the recognition results can be obtained using the PPO (Proximal Policy Optimization) algorithm.
[0029] In the fourth step, the following model supervised fine-tuning training steps are performed based on the sample dataset: In the first sub-step, at least one sample data in the sample dataset is input into the initial multi-modal large language model to obtain material information corresponding to the at least one sample data. The material information can be structured text content generated by the model corresponding to the sample data.
[0030] In the second sub-step, it is determined whether the initial multi-modal large language model reaches a preset optimization target according to a first target loss function. The optimization target can be a condition for stopping the training of the preset model. For example, the optimization target can be that the cross-entropy loss function value is reduced to below a preset value (e.g., 0.1). The first target loss function can be a cross-entropy loss function for measuring the difference between the text content output by the model and the standard answer text.
[0031] In the third sub-step, in response to the initial multi-modal large language model reaching the optimization target, the initial multi-modal large language model is used as a multi-modal large language model after supervised fine-tuning.
[0032] Optionally, the above execution subject can further perform the following steps: In a first step, in response to the initial multi-modal large language model not meeting the optimization target, the training parameters of the initial multi-modal large language model are adjusted, and the adjusted initial multi-modal large language model is used as the initial multi-modal large language model to perform the model supervised fine-tuning training step again. The training parameters can include learning rate, training batch size, and maximum training rounds. In practice, the parameters can be modified according to the state of model training. For example, if the model training appears gradient shock, the learning rate is reduced.
[0033] In a second step, the supervised fine-tuned multi-modal large language model is determined as an initial domain-aligned multi-modal large language model. The initial domain-aligned multi-modal large language model can be an initial model that is subjected to domain alignment training after supervised fine-tuning.
[0034] In a third step, according to the preference data set, the following model domain alignment step is performed: In a first sub-step, at least one regenerated material image data in the preference data set is input into the initial domain-aligned multi-modal large language model to obtain corresponding material information of the at least one regenerated material image data.
[0035] In a second sub-step, the corresponding material information of the at least one regenerated material image data is compared with the corresponding regenerated material image text data to obtain a second target loss function. The second target loss function can be a comparison loss function for the preference data. The comparison loss function can be a function used to measure the matching degree of the model generated result and the domain expert preference. For example, DPO (Direct Preference Optimization).
[0036] In a third sub-step, in response to the initial domain-aligned multi-modal large language model not meeting the optimization target, the training parameters of the initial domain-aligned multi-modal large language model are adjusted, and the adjusted initial domain-aligned multi-modal large language model is used as the initial domain-aligned multi-modal large language model to perform the model domain alignment step again.
[0037] In a fourth sub-step, in response to the initial domain-aligned multi-modal large language model meeting the optimization target, the initial domain-aligned multi-modal large language model is used as a domain-aligned multi-modal large language model.
[0038] In a fourth step, in response to the initial domain-aligned multi-modal large language model not meeting the optimization target, the training parameters of the initial domain-aligned multi-modal large language model are adjusted, and the adjusted initial domain-aligned multi-modal large language model is used as the initial domain-aligned multi-modal large language model to perform the model domain alignment step again.
[0039] At step 102, the following material knowledge acquisition steps are performed using a first thread: At step 1021, a knowledge vector block corresponding to the query statement information is queried from the material knowledge base.
[0040] In some embodiments, the execution subject can query the knowledge vector block corresponding to the query statement information from the material knowledge base. The material knowledge base can be a database that stores structured or semi-structured data related to materials (e.g., raw materials, parts). The database can be related information of the material, including attributes of the material (such as material, specifications, supplier information), use scenarios, processing technology, industry standards and specifications, and professional literature and publications. The knowledge vector block can be a vector representation of a knowledge fragment (e.g., a sentence, a paragraph, or a set of attributes) stored in the material knowledge base. For example, the query statement information can be a query vector converted by the same embedding model as the knowledge vector block. By determining the similarity between the query vector and the knowledge vector block in the knowledge base, the knowledge vector block corresponding to the query statement information is determined. The first thread can be a retrieval process based on semantic understanding. The natural language information is converted into a vector form, and a semantic match is performed in the first knowledge graph in the form of node vectors to obtain the corresponding material knowledge.
[0041] As an example, first, the query statement information is preprocessed and converted into a query vector. Second, the query vector is used to perform similarity retrieval in the vector database to find the most similar knowledge vector block to the query vector. Finally, the text information of the original knowledge corresponding to the knowledge vector block is returned.
[0042] At step 1022, the first material knowledge corresponding to the knowledge vector block is queried from the first knowledge graph corresponding to the regenerated material field.
[0043] In some embodiments, the execution subject can query the knowledge vector block corresponding to the query statement information from the material knowledge base. The material knowledge base can be a database that stores structured or semi-structured data related to materials (e.g., raw materials, parts). The database can be related information of the material, including attributes of the material (such as material, specifications, supplier information), use scenarios, processing technology, industry standards and specifications, and professional literature and publications. The knowledge vector block can be a vector representation of a knowledge fragment (e.g., a sentence, a paragraph, or a set of attributes) stored in the material knowledge base. For example, the query statement information can be a query vector converted by the same embedding model as the knowledge vector block. By determining the similarity between the query vector and the knowledge vector block in the knowledge base, the knowledge vector block corresponding to the query statement information is determined. The first thread can be a retrieval process based on semantic understanding. The natural language information is converted into a vector form, and a semantic match is performed in the first knowledge graph in the form of node vectors to obtain the corresponding material knowledge.
[0044] As an example, first, a nearest neighbor search algorithm (e.g., HNSW, Hierarchical Navigable Small World) can be performed to determine semantically similar entities of the knowledge vector block in the knowledge graph as a candidate entity set. Finally, the complete data record information of the candidate entity nodes and their associated edges in the candidate entity set is called through the graph database query interface as the output result of the target material knowledge.
[0045] Optionally, the above-mentioned material knowledge base and the above-mentioned first knowledge graph are cooperated through static anchor points and dynamic entity alignment; and the above-mentioned material knowledge base and the above-mentioned first knowledge graph are cooperated through the following steps: First, the standard fields of the material knowledge in the above-mentioned material knowledge base are extracted. Among them, the above-mentioned standard fields can be pre-defined key attribute fields (e.g., material name, material number and international standard number) for cross-knowledge base alignment. In practice, the pre-defined field extraction can be performed on each material knowledge in the material database, and stored as a set or list.
[0046] Second, according to the above-mentioned standard fields, the graph nodes in the above-mentioned first knowledge graph are matched to obtain a mapping relationship. Among them, the above-mentioned graph nodes can be entity nodes in the knowledge graph with unique identifiers and attributes (e.g., material name and material number). The above-mentioned mapping relationship can be the correspondence between the record in the material knowledge base and the knowledge graph node. In practice, the standard field can be used as a query condition to match the node in the knowledge graph through query. And the queried node is associated with the material knowledge record to form a mapping dictionary.
[0047] Third, in response to the above-mentioned mapping relationship being not empty, the above-mentioned material knowledge is bound to the above-mentioned graph node as a static anchor point. Among them, the above-mentioned static anchor point can be a binding relationship used to avoid repeated matching.
[0048] As an example, first, a new relationship edge is created between the material knowledge entity node and the corresponding graph node in the knowledge graph. Second, the material knowledge record is located in the material knowledge base, and a unique identifier pointing to the above-mentioned graph node is added in the metadata. Finally, the identification information of the record in the knowledge base is filled into the newly created relationship edge to obtain a record bidirectional relationship static anchor point.
[0049] Fourth, in response to the above-mentioned mapping relationship being empty, the following dynamic entity alignment operation steps are performed: First sub-step, determine the knowledge vector of the above-mentioned material knowledge. Among them, the above-mentioned knowledge vector can be a kind of representation form of the material knowledge in the vector space. In practice, the material knowledge can be input into an embedding model (e.g., a pre-trained BERT model) to generate a knowledge vector.
[0050] A second sub-step, determining similar nodes of the knowledge vector in the entity vector space of the first knowledge graph. The entity vector space can be a set of vector representations of all graph nodes in the knowledge graph. The similar nodes can be graph nodes that are close to the knowledge vector in terms of mathematical distance (e.g., Euclidean distance) or angle (e.g., cosine similarity). In practice, the similar nodes can be determined based on the cosine similarity between the knowledge vector and the entity vectors of the graph nodes.
[0051] A third sub-step, determining a confidence score of the similar nodes. The confidence score can be a reliability score (e.g., cosine similarity) of the entity alignment result. In practice, the confidence score can be determined by normalizing the cosine similarity of each similar node.
[0052] A fourth sub-step, binding the similar nodes and the material knowledge as static anchors in response to the confidence score being greater than a preset threshold. The preset threshold can be a pre-set minimum value of the entity alignment confidence score (e.g., the preset threshold can be 0.8).
[0053] Step 103, querying the second material knowledge corresponding to the material information from the second knowledge graph corresponding to the recycled material field using the second thread.
[0054] In some embodiments, the execution subject can query the second material knowledge corresponding to the material information from the second knowledge graph corresponding to the recycled material field using the second thread. The second knowledge graph can be a text-based knowledge graph. The second material knowledge can be the information content of the text file retrieved from the second knowledge graph. The second thread can be a structured data-based retrieval process. Based on the content of the feature tag information in the known material information (e.g., "category: copper, impurity content: 5%, main impurities: plastic"), the relationship in the text-based second knowledge graph is queried (e.g., the relationship edges "processing technology, risk, value" are traversed based on the "copper" entity node), and the corresponding material knowledge is obtained.
[0055] As an example, first, the material information is pre-processed and the key attributes are extracted. Then, the query statement is constructed using the key attributes. Second, the text content in the knowledge graph is parsed, and the in-memory query graph (e.g., RDF graph) is constructed. Finally, the query statement is executed in the query graph, and the material knowledge is obtained.
[0056] In the process of adopting technical solutions to solve the above technical problems, the system often returns material knowledge with fast generation speed, but cannot guarantee the accuracy of the knowledge. At the same time, when a thread has knowledge errors, the same errors exist in the knowledge graph of another thread. The generated report information is inaccurate, which cannot effectively guide the recycling process. To solve these problems, the conventional solution is generally: check the correctness of the knowledge through manual audit. And check the information source of the knowledge graph. The inventors consider adding manual audit, which increases the report generation time and cannot quickly obtain the report to make recycling decisions. We decided to use the following solutions: Optionally, the above execution subject can also perform the following steps: First, send the above second material knowledge to the intermediate cache area to perform the following third operation steps: First substep, extract the minimum knowledge unit of the second material knowledge in the above intermediate cache area to obtain knowledge information as a material knowledge set, wherein the material knowledge in the material knowledge set includes material entities and corresponding material attributes. Wherein the minimum knowledge unit can be a basic unit of knowledge in the form of a triple (entity, attribute, value). The above knowledge information can be the knowledge content in the form of a triple. The above material knowledge set can be a set composed of material entities and corresponding material attributes. Wherein the intermediate cache area can be a temporary storage area for storing data to be processed. In practice, the intermediate cache area can be a high-throughput memory cache (for example, Redis).
[0057] As an example, each knowledge entry in the second material knowledge can be parsed to extract entities and attributes to form minimum knowledge units. And combine into a material knowledge set.
[0058] Second substep, according to the preset material knowledge standard library, each material knowledge in the above material knowledge set is screened to obtain conflict knowledge as a conflict knowledge set. Wherein the above preset material knowledge standard library can be a verified secondary material industry standard knowledge base. The above conflict knowledge can be a knowledge unit that is inconsistent with the preset material knowledge standard library in the material knowledge set. The above conflict knowledge set can be a set composed of conflict knowledge.
[0059] As an example, first, traverse each material knowledge in the material knowledge set and query in the preset material knowledge standard library to obtain the corresponding standard knowledge. Then, compare the attributes of the material knowledge with the attributes in the standard library, and mark as conflict knowledge if there is inconsistency. Collect all conflict knowledge to form a conflict knowledge set.
[0060] Third substep, for each conflict knowledge in the above conflict knowledge set, perform the following fourth operation steps: Sub-step one, retrieving corresponding standard knowledge from the preset material knowledge standard library according to the material entity of the conflict knowledge. The standard knowledge can be the content of the correct knowledge corresponding to the conflict knowledge retrieved from the preset material knowledge standard library.
[0061] Sub-step two, extracting standard attributes in the standard knowledge. The standard attributes can be correct material characteristics (e.g., correct material entities and material entity relationships) in the standard knowledge. The correct material characteristics can be material characteristics that have been correctly verified.
[0062] Sub-step three, determining the structural consistency score of the material attributes of the conflict knowledge and the standard attributes. The structural consistency score can be a measure of the structural similarity between the material attributes of the conflict knowledge and the standard attributes.
[0063] As an example, first, using an embedding model, the material attributes of the conflict knowledge and the standard attributes are converted into feature vectors. Finally, the cosine similarity of the two vectors is calculated. The cosine similarity is used as the structural consistency score.
[0064] Sub-step four, determining the confidence score corresponding to the conflict knowledge in the second knowledge graph. The confidence score can be a measure of the reliability of the knowledge source when constructing the knowledge graph. In practice, the confidence score can be a value between 0 and 100. The higher the value, the higher the reliability of the standard knowledge source.
[0065] As an example, the confidence score of the conflict knowledge can be obtained from the metadata of the second knowledge graph.
[0066] Sub-step five, weighting and fusing the structural consistency score and the confidence score to obtain a comprehensive confidence score. The comprehensive confidence score can be a score that combines the structural consistency score and the confidence score to determine the reliability of the conflict knowledge. In practice, a weighted average method can be used.
[0067] Sub-step six, in response to the comprehensive confidence score being greater than a preset value, correcting the nodes and relationships corresponding to the conflict knowledge in the second knowledge graph and correcting the conflict knowledge in the second material knowledge. The preset value can be a standard value for judging the degree of error of the conflict knowledge. In practice, the standard value can be 0.8. The correction method can be an update operation in a graph database to add standard knowledge.
[0068] Sub-step seven, in response to the existence of the conflict knowledge in the first knowledge graph, updating the corresponding nodes and relationships in the first knowledge graph.
[0069] As an example, the conflicting knowledge can be queried in the first knowledge graph. The corresponding part in the first knowledge graph is modified if there is conflicting knowledge.
[0070] Sub-step eight, in response to the above-mentioned comprehensive confidence score being not greater than a preset value, the above-mentioned standard knowledge is added to the above-mentioned target material knowledge to generate reference annotations in the recycling report.
[0071] As an example, the standard knowledge is added to the target material knowledge, and is used as a reference annotation when generating a recycling report subsequently.
[0072] The above-mentioned operation steps, as one of the invention points of the present disclosure, solve the technical problems mentioned in the background art, that is, “due to the dispersion of recycled material knowledge, the related knowledge content is scattered in different platforms, documents and industry standards, and lacks unified organization and collection.” In practice, the conventional method increases the correctness of the knowledge through manual auditing, which increases the time of report generation. The present disclosure designs a knowledge graph conflict detection and correction scheme, which uses intermediate caching, vector calculation, confidence evaluation and weighted decision mechanism to identify the conflict between external material knowledge and standard library. And according to the comprehensive confidence score, the knowledge graph with errors is selected to be corrected. Or keep the standard knowledge as part of the report reference content. Therefore, by solving the potential conflict in the knowledge acquisition process, the quality of the recycling report can be improved. At the same time, the conflict in the knowledge graph data is corrected, and the quality of the graph itself is improved. The practitioners can follow the recycling report that conforms to the industry standard to carry out recycling work.
[0073] Step 104, determining the target material knowledge corresponding to the output of the target thread.
[0074] In some embodiments, the execution subject can determine the target material knowledge corresponding to the output of the target thread. Wherein the target thread is the thread that ends execution earliest among the first thread and the second thread, and the target material knowledge is the knowledge generated earliest among the first material knowledge and the second material knowledge. Wherein the target material knowledge can be the structured information of the material (for example, the composition, source, form, recycling method and flow scheme of the material).
[0075] In some optional implementations of some embodiments, the execution subject can determine the target material knowledge corresponding to the output of the target thread, which can include the following steps: First, determine the path depth of the knowledge vector block in the first knowledge graph to obtain the vector confidence. Wherein the path depth can be the shortest number of hops from the target entity in the knowledge vector block to the root node in the knowledge graph. The vector confidence can be a knowledge reliability score according to the path depth. For example, in the scene of knowledge graph query, the shallower the path depth, the higher the confidence.
[0076] As an example, a breadth-first search can be used to determine the depth of a target entity in the knowledge graph from a knowledge vector block. The confidence of the vector is determined based on the depth.
[0077] In a second step, a corresponding graph node of the knowledge vector block in the first knowledge graph is determined to obtain a semantic similarity. The graph node can be an entity corresponding to a material attribute (e.g., material type and material feature) in the knowledge graph. The semantic similarity can be a cosine similarity between the knowledge vector block and the embedding vector of the graph node.
[0078] In a third step, a corresponding graph relationship of the knowledge vector block in the first knowledge graph is determined to obtain a relationship similarity. The graph relationship can be a semantic relationship between two entities with types and attributes in the knowledge graph. The relationship similarity can be a matching degree between the relationship embedding in the knowledge vector block and the graph relationship. In practice, the cosine similarity between the relationship in the knowledge vector block and the graph relationship in the knowledge graph can be determined as the relationship similarity.
[0079] In a fourth step, a knowledge correctness score of the first material knowledge is determined based on the vector confidence, the semantic similarity, and the relationship similarity. The knowledge correctness score can be a weighted evaluation value of the vector confidence, the semantic similarity, and the relationship similarity.
[0080] In a fifth step, in response to the knowledge correctness score being lower than a preset warning threshold, the following first operation steps are performed: In a first sub-step, the first material knowledge is sent to an intermediate cache area for knowledge verification operations. The intermediate cache area can be a queue that supports high-concurrency writing and transaction processing for temporarily storing data to be verified.
[0081] In a second sub-step, the first material knowledge is added with a to-be-verified mark, and the first material knowledge is taken as a target material knowledge to generate a recovery report. The to-be-verified mark can carry unique identification content (e.g., JSON format content).
[0082] In the process of adopting technical solutions to solve the above technical problems, the following problems often occur: when fusing knowledge of recycled materials from different channels, the structure in each knowledge graph is different, and fusing knowledge of recycled materials may have inconsistency problems. At the same time, when different knowledge graphs interpret and reason about materials, there is a certain deviation in the knowledge obtained. It is not possible to record the knowledge that deviates. In view of these problems, the conventional solution is generally: modify through the record log of the knowledge graph. Label at the conflict data source to record. The inventors consider that the conflict in the record log lacks structured positioning and does not distinguish levels (for example, entities, attributes, and relationships). Label at the conflict data source, but cannot accurately locate the entity node in the knowledge graph. We decided to use the following solution: Optionally, the above execution subject can also perform the following steps: First, send the first material knowledge to the intermediate cache area, and perform the following knowledge verification operation steps: First substep, get the first material knowledge in the above intermediate cache area. As an example, the first material knowledge to be verified can be popped out from the intermediate cache area (for example, Redis). And parse the first material knowledge into a structure that the program can process (for example, a dictionary).
[0083] Second substep, feature extraction is performed on the above first material knowledge to obtain first material feature information, wherein the above first material feature information includes first material entity, first material semantic information and first material relationship. Wherein the feature extraction can be to extract the key components (such as entities, semantic information and relationships) in the material knowledge. The above first material feature information is the content information of the key components in the material knowledge. The above first material entity can be the entity object mainly described in the knowledge (for example, stainless steel 304). The above first material semantic information can be the content describing the attribute state of the entity (the density is 7.93 g / cm 3 ). The above first material relationship can be the relationship between the entity and other entities in the material knowledge (for example, "recyclable" corresponds to "stainless steel ingot").
[0084] As an example, a body recognition model (for example, spaCy, a pre-trained NER model) can be used to recognize entities in the material knowledge. Regular expressions can be used to extract text information in the material knowledge. A relationship extraction model (for example, a pre-trained BERT model) can be used to obtain the relationship between entities.
[0085] Third substep, feature extraction is performed on the above second material knowledge to obtain second material feature information, wherein the above second material feature information includes second material entity, second material semantic information and second material relationship.
[0086] A fourth sub-step, in response to the first material characteristic information and the second material characteristic information being inconsistent, performing the following second operation step: A first sub-step, in response to the first material entity and the second material entity being inconsistent, determining a conflict entity in the first material entity and the second material entity. The conflict entity can be a part of the first material entity and the second material entity that is inconsistent. For example, in a steel recycling scenario, the first material entity can be "stainless steel 304", and the second material entity can be "stainless steel 316". The conflict entity can be "stainless steel 304".
[0087] A second sub-step, in response to the first material semantic information and the second material semantic information being inconsistent, determining a conflict semantic information in the first material semantic information and the second material semantic information. The conflict semantic information can be a semantic description that is inconsistent on the same entity or attribute in the material knowledge. For example, in a steel recycling scenario, the first material entity can be "stainless steel 304". The semantic information can be "density 7.93 g / cm3". The second material entity can be "stainless steel 304". The semantic information can be "density 7.99 g / cm3". The conflict semantic information can be "density 7.93 g / cm3".
[0088] A third sub-step, in response to the first material relationship and the second material relationship being inconsistent, determining a conflict material relationship in the first material relationship and the second material relationship. The conflict material relationship can be a relationship between entities that points to different entities on the same entity in the material knowledge. For example, in a steel recycling scenario, the first material entity can be "stainless steel 304". The material relationship can be "recyclable", "stainless steel ingot". The second material entity can be "stainless steel 304". The material relationship can be "recyclable", "rebar". The conflict material relationship can be "stainless steel ingot" of the first material relationship and "rebar" of the second material relationship.
[0089] A fourth sub-step, determining a query path of the conflict entity in the first material knowledge as a conflict point. The query path can be a query path of the conflict entity in the first material knowledge. The conflict point can be a query path from the root node to the conflict entity.
[0090] As an example, a path query is performed in a graph database corresponding to the knowledge graph. A query path from the root node to the conflict entity is determined.
[0091] Sub-step five, determine the knowledge source of the conflict semantic information in the first knowledge graph as conflict information. The knowledge source can be the knowledge data used when constructing the knowledge graph (e.g., data sets, journals, and expert reports). The conflict information can be the specific content of the knowledge data.
[0092] As an example, the knowledge graph construction record can be searched to query the source of the conflict semantic information (e.g., text content, data source, update time, and provider) from the first knowledge graph.
[0093] Sub-step six, determine the node relationship of the conflict material relationship in the first knowledge graph and the second knowledge graph as an abnormal relationship. The node relationship can be the connection relationship between entities in the knowledge graph. The abnormal relationship can be the specific manifestation of the conflict material relationship in the two knowledge graphs (e.g., different relationship types and different target entities).
[0094] As an example, the conflict material relationship can be queried from the respective graph databases of the first knowledge graph and the second knowledge graph. The entity types and connected entities corresponding to the conflict material relationship are obtained.
[0095] Sub-step seven, construct a conflict comparison result according to the conflict point, the conflict information, and the abnormal relationship. The conflict comparison result can be structured information that integrates the conflict point, the conflict information, and the abnormal relationship.
[0096] Fifth sub-step, analyze the query chain of the first thread according to the conflict comparison result to obtain an abnormal path. The abnormal path can be the step path in the query chain that causes the conflict in the query process.
[0097] As an example, the step of the query can be parsed by calling a chain tracking tool according to the query log. The conflict point is mapped to a specific step in the query chain to determine the query path.
[0098] Sixth sub-step, structure optimization is performed on the material knowledge base and the first knowledge graph according to the abnormal path. The structure optimization can be the cause of the abnormality to modify the knowledge base and the pruning operation of the knowledge graph. The knowledge base modification can be the completion or deletion of the content in the knowledge base. The knowledge graph pruning operation can be the constraint and adjustment of the relationship of the corresponding entity in the graph according to the abnormal relationship.
[0099] The above operation steps, as one of the invention points of the present disclosure, solve the technical problems mentioned in the background art: "due to the dispersion of the material knowledge, the relevant knowledge content is scattered in different platforms, documents and industry standards, and there is a lack of unified organization and collection"; and "the conflict in the log record lacks structured positioning, and there is no distinction between levels (for example, entities, attributes and relationships). The conflict data source is labeled, but it cannot be accurately positioned to the entity node in the knowledge graph." In practice, the conventional method regards the data conflict as a part that can be isolated and processed, ignoring the relevance between the conflict data contents (for example, between entities and between entities and nodes). The present disclosure designs a deep knowledge conflict and correction scheme. After detecting the conflicts of the material knowledge from different sources, the conflicts are accurately classified through fine feature extraction (for example, entities, semantics and relationships). The path, source and specific performance of the conflict point in the knowledge graph are traced. According to the conflict comparison result, the query chain links (for example, abnormal path) that cause the conflict are analyzed, and the underlying knowledge base (for example, content addition and deletion) and the upper knowledge graph (for example, structure pruning and constraint) are actively corrected. Therefore, through the clear conflict content and conflict correction method, the maintenance of the knowledge base and the knowledge graph is targeted. It is more efficient to obtain material knowledge, so as to generate a high-quality recycling report to guide recycling. The efficiency of recycling material by practitioners is improved.
[0100] Optionally, the above execution subject can further perform the following steps: First, extract the material entity name in the first material knowledge and the second material knowledge. The material entity name can be the unique identifier of the material.
[0101] As an example, the name of the entity can be extracted in the first knowledge graph using a query method. The material entity name in the text can be identified using natural language processing technology.
[0102] Second, according to the preset material entity standard table, the material entity name is converted into a coded standard material entity. The preset material entity standard table can be a correspondence mapping table that stores entity relationships and unique codes. The standard material entity can be a unique identifier after mapping conversion by the mapping table.
[0103] As an example, each material entity name can be traversed to find the corresponding code in the preset material entity standard table to obtain the standard material entity. In practice, first, the preset material entity standard table can be loaded into a dictionary structure. Second, the material entity name is used as the key and the code is used as the value. Each material entity name is queried in the dictionary to obtain the corresponding standard code. Finally, the standard code obtained is used to replace the material entity name to obtain the standard material entity.
[0104] Thirdly, determine the graph structure and the first attribute set of the standard material entity in the first knowledge graph. The graph structure can be the network structure of entity nodes and edge relationships in the knowledge graph. The first attribute set can be the standard entity attribute set extracted from the first knowledge graph (structured graph). For example, the standard material entity can be "motor". The first attribute can be "material: copper coil" and "weight: 5 kg".
[0105] As an example, first, in the knowledge graph, query the graph structure (e.g., nodes, edge relationships) centered on the standard material entity. Then, parse the attributes of the nodes from the graph structure. Then, represent the graph structure as a data structure (e.g., JSON and graph object). The attributes are stored in the form of a dictionary.
[0106] Fourthly, determine the text attribute structure and the second attribute set of the standard material entity in the second knowledge graph. The text attribute structure can be a set of unstructured text descriptions (e.g., JSON key-value pairs). The second attribute set can be the entity attribute set parsed from the second knowledge graph (text construction graph). For example, the standard material entity can be "motor". The second attribute can be "material: copper, parameter: weight equal to 5 kg".
[0107] As an example, first, locate the relevant records according to the standard material entity in the second knowledge graph to obtain the material entity content. Then, extract the entire attribute object in the material entity content. In practice, the JSON key-value tree can be parsed, and only the entire structure is retained. Secondly, the text attribute structure is used as the entity attribute. In practice, the key-value of all leaf nodes in the JSON can be extracted. Finally, the text attribute structure is stored as JSON, and the attribute set is stored in the form of a dictionary.
[0108] Fifthly, convert the text attribute structure into a text construction graph structure. The text construction graph structure can be a pseudo-graph structure converted from the text attribute.
[0109] As an example, the text attribute key can be used as a graph node. The text attribute value can be used as a leaf node. The edge can represent the text affiliation.
[0110] Sixthly, determine the structural similarity between the graph structure and the text construction graph according to the graph edit distance algorithm. The graph edit distance algorithm can be an algorithm for determining the difference between graph structures (e.g., the GED library in Python). In practice, the cost of modifying the content of the graph structure can be determined as the evaluation standard. The structural similarity can be a similarity score normalized by the graph edit distance.
[0111] As an example, first, the operation cost can be determined. For example, the default node / edge insertion / deletion cost is 1, and the replacement cost is 0.5. Second, the graph edit distance of the graph structure and the text construction graph is determined to obtain the total edit operation and the total cost. For example, the graph structure can be ["motor -> (material) -> copper coil", "motor -> (weight) -> 5kg"]. The text construction graph can be ["motor -> (material) -> copper", "motor -> (parameter) -> weight = 5kg"]. The replacement cost of the node "copper" in the text construction graph for "copper coil" can be 0.5. The replacement cost of the edge "material" in the text construction graph for "material" can be 0.5. The insertion cost of the missing edge "weight" in the text construction graph can be 1. The deletion cost of the node "weight = 5kg" in the text construction graph can be 1. The total edit operation can be 0.5+0.5+1+1=3.0. The edit operation of the graph structure can be 3 node replacements and 2 edge replacements, and the graph structure cost is 3+2=5. The edit operation of the text construction graph can be 3 node replacements and 2 edge replacements. The text construction graph cost is 3+2=5. The total cost is 5+5=10.0. Finally, the total edit operation is normalized according to the total cost to obtain the structural similarity. For example, the structural similarity can be 1-(3.0 / 10.0)=0.7.
[0112] In the seventh step, in response to the structural similarity being greater than the preset threshold, the merging of the first attribute set and the second attribute set is performed to obtain a structure alignment result. The attribute merging can be merging the entity attribute content obtained from the knowledge graph. In practice, the dictionary merging can be used to merge the contents in the first attribute set and the second attribute set. The structure alignment result can be a description of the merged knowledge graph after alignment. In practice, it can be a new knowledge graph or JSON.
[0113] In the eighth step, in response to the structural similarity being not greater than the preset threshold, the following fifth operation step is performed: In the first sub-step, the difference structure of the nodes and edges in the graph structure and the text construction graph is identified according to the graph edit distance algorithm to obtain a difference tuple. The difference tuple can be a triple (operation type, element type, element ID) recording the structural difference.
[0114] As an example, first, the operation sequence output by the graph edit distance algorithm can be used. Finally, each operation sequence is traversed to create a triple.
[0115] In the second sub-step, the corresponding attributes in the first attribute set and the second attribute set are stored in the local memory for knowledge graph structure repair. The local memory can be a local storage medium (for example, a cache or a file in RAM).
[0116] As an example, the difference tuples are traversed, and for each element ID, corresponding attributes in the first attribute set and the second attribute set are extracted and stored in a cache or file in the RAM.
[0117] Step 105 : Based on the target material knowledge, a multimodal large language model is used to generate a material recycling report corresponding to the material information, so as to recycle the recycled materials according to the material recycling report.
[0118] In some embodiments, the execution entity may generate a material recycling report corresponding to the material information based on the target material knowledge and the multimodal large language model, and perform recycled material recycling based on the material recycling report. The material recycling report may be a text file that details recommendations for recycling renewable materials, including recycling plan recommendations, recycling precautions, and a circulation plan. The recycled material recycling process may involve relevant personnel performing actual recycling operations on renewable materials based on the contents of the material recycling report.
[0119] As an example, first, a natural language instruction template can be defined. For example, "You are a material recycling expert. Please generate a recycling report based on the following material attributes" plus target material knowledge. Next, the natural language instruction template is supplemented with target material knowledge and used as a report generation prompt. Finally, the report generation prompt is input into a multimodal large language model to generate a material recycling report.
[0120] The above various embodiments of the present disclosure have the following beneficial effects: through the recycled material recycling method based on the multi-modal large language model of some embodiments of the present disclosure, the recycled material image can be automatically identified, the structured and unstructured knowledge graph is combined to realize fast knowledge retrieval, and a recycling report is automatically generated. In the report generation process, the traditional method often needs to manually identify and classify the material first, and then search and process the specification and industry information through multiple platforms, which takes a long time and cannot quickly generate a recycling report with practical guiding significance according to the material image. Based on this, the recycled material recycling method based on the multi-modal large language model of some embodiments of the present disclosure first generates the material information and query statement information corresponding to the target recycled material image by using the pre-trained multi-modal large language model. In this way, automatic perception and semantic understanding of the recycled material image can be realized, reducing the dependence on manual identification and providing a structured input basis for subsequent knowledge retrieval. Secondly, by using the first thread, the following material knowledge acquisition steps are executed: querying the knowledge vector block corresponding to the query statement information from the material knowledge base; querying the first material knowledge corresponding to the knowledge vector block from the first knowledge graph corresponding to the recycled material field, wherein the first knowledge graph is a knowledge graph in the form of node vectors. In this way, professional material knowledge related to the query statement can be efficiently acquired based on the structured query path, and the knowledge corresponding to the recycled material can be quickly obtained. Then, by using the second thread, the second material knowledge corresponding to the material information is queried from the second knowledge graph corresponding to the recycled material field, wherein the second knowledge graph is a knowledge graph in the form of text. Then, the target material knowledge corresponding to the output of the target thread is determined, wherein the target thread is the thread that ends earliest among the first thread and the second thread, and the target material knowledge is the knowledge generated earliest among the first material knowledge and the second material knowledge. In this way, the knowledge corresponding to the recycled material can be quickly acquired on the premise that the recycled material knowledge is effective. Finally, according to the target material knowledge, the material recycling report corresponding to the material information is generated by using the multi-modal large language model, so as to recycle the recycled material according to the material recycling report. In this way, a recycling recommendation report with practical value and professional depth can be automatically generated. In summary, through automatic identification of the recycled material image, parallel retrieval of the knowledge graph, quick acquisition of the knowledge related to the recycled material, generation of the recycling report, and guidance of the subsequent material sorting, processing and recycling of the practitioners.
[0121] Further reference Figure 2 , as an implementation of the method shown in the above figures, the present disclosure provides some embodiments of a recycled material recycling device based on a multi-modal large language model, which correspond to the method embodiments shown in Figure 1 The recycled material recycling device based on the multi-modal large language model can be applied in various electronic devices.
[0122] As shown in Figure 2 The multi-modal large language model-based recyclable material recycling apparatus 200 includes an acquisition unit 201, a first thread unit 202, a second thread unit 203, a knowledge generation unit 204, and a report generation unit 205. The acquisition unit 201 is configured to generate material information corresponding to a target recyclable material image and query statement information by using a pre-trained multi-modal large language model. The first thread unit 202 is configured to execute the following material knowledge acquisition steps by using a first thread: querying a knowledge vector block corresponding to the query statement information from a material knowledge base; and querying first material knowledge corresponding to the knowledge vector block from a first knowledge graph corresponding to the recyclable material field, wherein the first knowledge graph is a knowledge graph in the form of node vectors. The second thread unit 203 is configured to query second material knowledge corresponding to the material information from a second knowledge graph corresponding to the recyclable material field by using a second thread, wherein the second knowledge graph is a knowledge graph in the form of text. The knowledge generation unit 204 is configured to determine target material knowledge corresponding to the output of a target thread, wherein the target thread is a thread that ends execution earliest among the first thread and the second thread, and the target material knowledge is the knowledge that is generated earliest among the first material knowledge and the second material knowledge. The report generation unit 205 is configured to generate a material recycling report corresponding to the material information by using the multi-modal large language model according to the target material knowledge, so as to recycle the recyclable material according to the material recycling report.
[0123] It can be understood that the units described in the multi-modal large language model-based recyclable material recycling apparatus 200 correspond to the respective steps in the method described with reference to Figure 1 Therefore, the operations, features, and beneficial effects described above with respect to the method also apply to the multi-modal large language model-based recyclable material recycling apparatus 200 and the units included therein, and will not be described here again.
[0124] Reference is made below to Figure 3 which shows a structural schematic diagram of an electronic device (e.g., an electronic device) 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is merely an example and should not impose any limitation on the functions and use range of the embodiments of the present disclosure.
[0125] As shown in Figure 3As shown, the electronic device 300 can include a processing device (e.g., a central processor, a graphics processor, etc.) 301 that can perform various appropriate actions and processes according to programs stored in a read only memory (ROM) 302 or loaded into a random access memory (RAM) 303 from a storage device 308. Various programs and data required for the operation of the electronic device 300 are also stored in the RAM 303. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0126] Generally, the following devices can be connected to the I / O interface 305: input devices 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 308 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 309. The communication devices 309 can allow the electronic device 300 to communicate wirelessly or wired with other devices to exchange data. Although Figure 3 The electronic device 300 is shown with various devices, but it should be understood that all of the illustrated devices are not required, and more or fewer devices can alternatively be implemented. Figure 3 Each block shown in the flowcharts can represent a device, or multiple devices, as necessary.
[0127] In particular, processes described above with reference to the flowcharts can be implemented as a computer software program according to some embodiments of the present disclosure. For example, some embodiments of the present disclosure include a computer program product including a computer program carried on a computer readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In some such embodiments, the computer program can be downloaded and installed from a network through the communication devices 309, or installed from the storage devices 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above-described functions defined in the methods of some embodiments of the present disclosure are performed.
[0128] Note that the computer-readable medium or media used to provide the computer program sequence to the computer system can be accompanied by information sufficient to load the program sequence into an internal memory of the computer system. Generally, such information can be stored in one or more computer system memories or data storage systems from which, upon an execution by the computer system of one or more computer-readable programs, a processor of the computer system can retrieve. Based on these inputs, the computer system is programmed to perform particular tasks according to the instructions of the computer program contained in the computer-readable medium(s).
[0129] In some embodiments, the client, server, or both can communicate using any known or future developed network protocols, such as the HyperText Transfer Protocol (HTTP), and can be interconnected with any form or medium of digital data communication (for example, a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), the Internet, and peer-to-peer networks (for example, ad hoc peer-to-peer networks), as well as any current or future developed network.
[0130] The computer readable medium can be included in the electronic device; or can exist independently of the electronic device. The computer readable medium carries one or more programs, when the one or more programs are executed by the electronic device, the electronic device is caused to: generate material information and query statement information corresponding to a target renewable material image by using a pre-trained multi-modal large language model; perform the following material knowledge acquisition steps by using a first thread: query a knowledge vector block corresponding to the query statement information from a material knowledge base; query first material knowledge corresponding to the knowledge vector block from a first knowledge graph corresponding to a renewable material field, wherein the first knowledge graph is a knowledge graph in the form of node vectors; query second material knowledge corresponding to the material information from a second knowledge graph corresponding to the renewable material field by using a second thread, wherein the second knowledge graph is a knowledge graph in the form of text; determine target material knowledge corresponding to output of a target thread, wherein the target thread is a thread that ends execution earliest among the first thread and the second thread, and the target material knowledge is knowledge that is generated earliest among the first material knowledge and the second material knowledge; and generate a material recycling report corresponding to the material information by using the multi-modal large language model according to the target material knowledge, so as to recycle renewable materials according to the material recycling report.
[0131] Computer program code for carrying out operations of some embodiments of the disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0132] The flow and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow and block diagrams can represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may be executed in the reverse order, depending on the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or combinations of hardware and software.
[0133] The units described in some embodiments of the present disclosure can be implemented in the form of software, or can be implemented in the form of hardware. The described units can also be arranged in a processor, for example, it can be described that: a processor includes an acquisition unit, a first thread unit, a second thread unit, a knowledge generation unit and a report generation unit. Among them, the names of these units do not constitute a limitation to the units themselves in some cases, for example, the acquisition unit can also be described as "a unit that generates material information and query statement information corresponding to the target analog material image by using a pre-trained multi-modal large language model".
[0134] The functions described above in the specification can be implemented, at least in part, by one or more hardware logic components. For example, and without limitation, non-limiting examples of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on a chip (SOCs), complex programmable logic devices (CPLDs), etc.
[0135] The above description is merely some of the preferred embodiments of the present disclosure and a description of the principles of the technology used. Those skilled in the art should understand that the scope of the application involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or equivalent features without departing from the above inventive concept. For example, the above features are replaced with the technical features disclosed in the embodiments of the present disclosure (but not limited to) having similar functions to form technical solutions.
Claims
1. A recycled material recycling method based on a multimodal large language model, comprising: Utilize the pre-trained multimodal large language model to generate material information and query statement information corresponding to the target recycled material image; Using the first thread, perform the following material knowledge acquisition steps: Querying the knowledge vector block corresponding to the query statement information from the material knowledge base; Querying first material knowledge corresponding to the knowledge vector block from a first knowledge graph corresponding to the field of recycled materials, wherein the first knowledge graph is a knowledge graph in the form of node vectors; Using the second thread, querying the second material knowledge corresponding to the material information from the second knowledge graph corresponding to the recycled materials field, wherein the second knowledge graph is a knowledge graph in text form; Determine target material knowledge outputted by a target thread, wherein the target thread is the thread that finishes execution earliest between the first thread and the second thread, and the target material knowledge is the knowledge generated earliest between the first material knowledge and the second material knowledge; According to the target material knowledge, the multimodal large language model is used to generate a material recycling report corresponding to the material information, so as to recycle the recycled materials according to the material recycling report.
2. The method according to claim 1, wherein The material knowledge base and the first knowledge graph operate in collaboration through static anchor points and dynamic entity alignment; and the material knowledge base and the first knowledge graph operate in collaboration through the following steps: Extracting standard fields of material knowledge in the material knowledge base; According to the standard field, the graph nodes in the first knowledge graph are matched to obtain a mapping relationship; In response to the mapping relationship being not empty, binding the material knowledge to the graph node as a static anchor point; In response to the mapping relationship being empty, the following dynamic entity alignment operation steps are performed: determining a knowledge vector of the material knowledge; Determining similar nodes of the knowledge vector in the entity vector space of the first knowledge graph; Determining confidence scores of the similar nodes; In response to the confidence score being greater than a preset threshold, the similar node and the material knowledge are bound as static anchor points.
3. The method according to claim 1, wherein The step of determining target material knowledge corresponding to the output of the target thread includes: Determine the path depth of the knowledge vector block in the first knowledge graph to obtain vector confidence; Determine the graph node corresponding to the knowledge vector block in the first knowledge graph to obtain semantic similarity; Determine the graph relationship corresponding to the knowledge vector block in the first knowledge graph to obtain the relationship similarity; determining a knowledge correctness score of the first material knowledge according to the vector confidence, the semantic similarity, and the relationship similarity; In response to the knowledge accuracy score being lower than a preset warning threshold, the following first operation step is performed: Sending the first material knowledge to an intermediate buffer area for knowledge verification operation; A to-be-verified mark is added to the first material knowledge, and the first material knowledge is used as target material knowledge to generate a recycling report.
4. The method according to claim 1, wherein The method further comprises: Extracting material entity names from the first material knowledge and the second material knowledge; According to the preset material entity standard table, the material entity name is converted into a standard material entity in coded form; Determining a graph structure and a first attribute set of the standard material entity in the first knowledge graph; Determining a text attribute structure and a second attribute set of the standard material entity in the second knowledge graph; Converting the text attribute structure into a text construction graph structure; Determining the structural similarity between the graph structure and the text structure graph according to a graph edit distance algorithm; In response to the structural similarity being greater than a preset threshold, merging the first attribute set and the second attribute set to obtain a structural alignment result; In response to the structural similarity being no greater than a preset threshold, performing the following fifth operation step: Identifying the difference in the nodes and edges between the graph structure and the text structure graph according to the graph edit distance algorithm to obtain difference tuples; According to the difference tuple, corresponding attributes in the first attribute set and the second attribute set are stored in a local memory to repair the knowledge graph structure.
5. The method according to claim 1, wherein The multimodal large language model is a large language model that is fine-tuned and domain-aligned based on a multimodal dataset; and the multimodal large language model is trained by the following steps: Obtain an open-source multimodal large language model as the initial multimodal large language model; Obtaining a pre-constructed sample dataset, wherein the data in the sample dataset is in the form of recycled material image-text pairs, for performing the supervised fine-tuning training; Obtaining a pre-built preference dataset, wherein the preference dataset includes recycled material image data and recycled material image text data, and the data in the preference dataset is for the domain alignment training; Based on the sample dataset, perform the following model supervised fine-tuning training steps: Inputting at least one sample data in the sample data set into the initial multimodal large language model to obtain material information corresponding to the at least one sample data; Determining whether the initial multimodal large language model achieves a preset optimization goal according to a first objective loss function; In response to the initial multimodal large language model reaching the optimization target, the initial multimodal large language model is used as the multimodal large language model completed by supervised fine-tuning.
6. The method according to claim 5, wherein: The method further comprises: In response to the initial multimodal large language model failing to achieve the optimization goal, adjusting training parameters of the initial multimodal large language model, and using the adjusted initial multimodal large language model as the initial multimodal large language model to perform the model supervised fine-tuning training step again; Determining the multimodal large language model completed by the supervised fine-tuning as the initial domain-aligned multimodal large language model; Based on the preferred dataset, the following model domain alignment steps are performed: Inputting at least one recycled material image data in the preference data set into the initial domain aligned multimodal large language model to obtain material information corresponding to the at least one recycled material image data; In response to determining, according to the second objective loss function, whether the initial domain-aligned multimodal large language model achieves a preset optimization goal; In response to the initial domain-aligned multimodal large language model achieving the optimization goal, using the initial domain-aligned multimodal large language model as the domain-aligned multimodal large language model; In response to the initial domain-aligned multimodal large language model failing to achieve the optimization goal, adjusting the training parameters of the initial domain-aligned multimodal large language model, and using the adjusted initial domain-aligned multimodal large language model as the initial domain-aligned multimodal large language model, and performing the model domain alignment step again.
7. A recycled material recovery device based on a multimodal large language model, comprising: an acquisition unit configured to generate material information and query statement information corresponding to the target recycled material image using a pre-trained multimodal large language model; The first thread unit is configured to use the first thread to perform the following material knowledge acquisition steps: Querying the knowledge vector block corresponding to the query statement information from the material knowledge base; Querying first material knowledge corresponding to the knowledge vector block from a first knowledge graph corresponding to the field of recycled materials, wherein the first knowledge graph is a knowledge graph in the form of node vectors; The second thread unit is configured to use the second thread to query second material knowledge corresponding to the material information from a second knowledge graph corresponding to the field of recycled materials, wherein the second knowledge graph is a knowledge graph in text form; a knowledge generation unit configured to determine target material knowledge outputted by a target thread, wherein the target thread is the thread that finishes execution earliest between the first thread and the second thread, and the target material knowledge is the knowledge generated earliest between the first material knowledge and the second material knowledge; The report generating unit is configured to generate a material recycling report corresponding to the material information based on the target material knowledge and using the multimodal large language model, so as to recycle the recycled materials according to the material recycling report.
8. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
9. A computer-readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Image recognition method and device, readable storage medium and electronic equipment
CN112214626A
Dynamic updating method and device for power grid dispatching knowledge graph
CN112905804A
Target object identification method, and training method and device of multi-modal identification model
CN116226785A
Auxiliary retrieval method fusing knowledge graph and large language model
CN117633252A
Knowledge graph construction method and system for chemical waste treatment technology
CN118734958A