Method and device for constructing a specialized text embedding model for solid waste management

By constructing a text embedding model specifically for solid waste management and utilizing a professional corpus and fine-tuning techniques for large language models, the problems of inaccurate semantic representation and insufficient robustness of general models in the field of solid waste were solved. This resulted in efficient semantic understanding and anti-interference capabilities, improving the model's application performance in solid waste management.

CN122491267APending Publication Date: 2026-07-31TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TSINGHUA UNIVERSITY
Filing Date
2026-05-22
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing general natural language processing models suffer from inaccurate semantic representation and insufficient robustness when dealing with solid waste management. They struggle to effectively identify low-frequency technical terms and complex policy logic, resulting in low recognition accuracy and susceptibility to interference from semantically similar texts. Furthermore, they lack systematic training datasets and evaluation benchmarks.

Method used

We construct a dedicated text embedding model for solid waste management. By building a professional corpus with eight core dimensions, and combining self-supervised learning with fine-tuning techniques driven by a large language model, we generate a triplet dataset and a benchmark dataset. We then use a contrastive learning loss function and a cosine annealing scheduling strategy for full-parameter supervised fine-tuning to enhance the model's semantic understanding and robustness.

Benefits of technology

It significantly improves the retrieval accuracy and classification precision of the model in the field of solid waste management, enhances the robustness of semantic representation, provides underlying support for a high-performance intelligent retrieval enhancement generation system, and solves the problem of insufficient adaptability and robustness of general models in the field of solid waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122491267A_ABST
    Figure CN122491267A_ABST
Patent Text Reader

Abstract

This invention relates to the fields of artificial intelligence and natural language processing (NLP), and particularly to a method and apparatus for constructing a dedicated text embedding model for solid waste management. The method includes: converting multi-format documents from a pre-built underlying corpus for solid waste management into structured text; performing semantically sensitive segmentation on the structured text; performing rule-based filtering and prompting engineering dual verification on the segmented structured text to obtain a triplet dataset and a benchmark dataset; and using the triplet dataset to perform fully parameter-supervised fine-tuning of a pre-built initial base model to obtain a dedicated text embedding model for solid waste. This solves the problems of weak semantic capture ability and insufficient robustness of general NLP models in the solid waste vertical domain, significantly improving the accuracy of key tasks such as retrieval and classification, and providing high-performance underlying semantic representation support for the intelligent transformation of solid waste management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and natural language processing, and in particular to a method and apparatus for constructing a dedicated text embedding model for solid waste management. Background Technology

[0002] With the acceleration of global urbanization and the booming development of industrial production, the generation of solid waste is surging at an unprecedented rate. Statistics show that the world generates over 2 billion tons of municipal solid waste (MSW) annually, and this figure is projected to reach 3.4 billion tons by 2050. Traditional management methods are inefficient, costly, and unable to cope with increasingly complex waste flows, posing a serious threat to the environment and health. Therefore, utilizing artificial intelligence (AI) technology to intelligently upgrade solid waste management systems has become an inevitable choice for achieving sustainable development and the goals of a circular economy.

[0003] In recent years, large language models based on the Transformer architecture have demonstrated significant "emergent capabilities" through a leap in parameter count driven by scaling laws, and are becoming a cornerstone for the advancement of general artificial intelligence. This technology achieves a paradigm shift from discriminative to generative models, and from a "model-centric" to a "data center" approach, possessing powerful multimodal processing capabilities, logical deduction, and complex contextual understanding. In multiple vertical scientific fields such as natural language processing, healthcare, finance, and autonomous driving, general-purpose LLMs have shown remarkable adaptability and highly scalable optimization potential.

[0004] In the field of solid waste management, LLM is leading a paradigm shift from single-task to general intelligence. Unlike traditional CNN models, large-scale models, with their multimodal understanding and powerful logical reasoning capabilities, can simultaneously integrate visual recognition, sensor signals, and policy texts, improving the generalization of solid waste classification in complex scenarios. Through prompt word engineering and fine-tuning techniques, LLM can serve as an intelligent assistant to optimize waste collection scheduling, identify illegal dumping, and generate compliance reports, providing new optimization space for building a full-chain, highly scalable intelligent solid waste governance system.

[0005] Currently, the application of large language models in the field of solid waste management (SWM) is still in its early stages, with significant issues such as a lack of targeted training and an imperfect evaluation system. Most existing applications rely on general-purpose large language models, lacking large-scale, systematic corpora specific to the solid waste domain for deep training. This results in limited model understanding and processing capabilities for high-order tasks in solid waste management, such as complex reasoning, policy logic analysis, and technology roadmap planning. Furthermore, the models are ineffective in handling vertical scenarios such as environmental engineering experimental design, cross-language solid waste policy alignment, and intelligent decision support, thus limiting the reliability and scalability of intelligent agent systems in practical applications.

[0006] At the text vectorization level, general embedding models show significant inadequacy when processing texts in the solid waste field. Solid waste texts are highly specialized, containing numerous low-frequency technical terms and complex policy clauses. General models struggle to accurately capture the deep semantic relationships between these terms, resulting in low-quality word vector representations. Furthermore, existing models lack robustness in feature extraction from scarce domain data, making it difficult to achieve fine-grained semantic differentiation in massive amounts of text and highly susceptible to interference from semantically similar but logically inconsistent texts. This insufficient representation capability directly restricts the efficient construction of solid waste knowledge bases and the response accuracy of retrieval-enhanced generation (RAG) systems.

[0007] In summary, the application of large language models in solid waste management still faces the challenge of inaccurate representation of specialized semantics. General embedding models, lacking deep learning capabilities for low-frequency specialized terminology, complex policy logic, and technological connections within the solid waste domain, suffer from low accuracy in processing specialized texts and are susceptible to interference from semantically similar texts. Furthermore, the industry lacks systematic training datasets and refined evaluation benchmarks for text embedding models in the solid waste field, making it difficult to support the high-performance requirements of intelligent management systems in complex retrieval and classification tasks. Summary of the Invention

[0008] This invention provides a method and apparatus for constructing a dedicated text embedding model for solid waste management, in order to solve the problems of insufficient accuracy, limited semantic understanding depth, and insufficient robustness of existing general natural language processing models when facing highly specialized terminology recognition, policy semantic deep analysis, and complex technical document classification in the field of solid waste.

[0009] A first aspect of this invention provides a method for constructing a dedicated text embedding model for solid waste management, comprising the following steps: Convert multi-format documents from a pre-built corpus for solid waste management into structured text; The structured text is subjected to semantically sensitive segmentation to obtain the segmented structured text; The segmented structured text is subjected to rule-based filtering and prompting engineering for dual verification to obtain a triplet dataset and a benchmark dataset. The pre-built initial pedestal model was fine-tuned with full parameter supervision using the triple dataset to obtain a text embedding model specifically for solid waste.

[0010] Optionally, the underlying corpus for solid waste management is constructed based on a classification system of eight dimensions: basic concepts of solid waste, laws and policies, management system, transfer and collection, disposal technology, risk assessment, zero-waste cities and circular economy, and artificial intelligence applications.

[0011] Optionally, the step of performing rule-based filtering and cue engineering dual validation on the segmented structured text to obtain a triplet dataset and a benchmark dataset includes: The filtered structured text is obtained by removing hyperlinks, non-text symbols, bulk references, and residual information with formatting errors from the segmented structured text. The filtered structured text is subjected to solid waste relevance determination by the prompting engineering to obtain the triple dataset and the benchmark dataset, wherein both the triple dataset and the benchmark dataset include multiple query-positive-negative triples.

[0012] Optionally, the step of performing fully parameter-supervised fine-tuning of the pre-built initial pedestal model using the triplet dataset and the benchmark dataset to obtain a text embedding model specifically for solid waste includes: A preset Chinese retrieval dataset is mixed into the triplet dataset to construct a training sample set; Based on the contrastive learning loss function and the cosine annealing scheduling strategy, the pre-constructed initial base model is fine-tuned with full parameter supervision using the training sample set to obtain the solid waste-specific text embedding model.

[0013] Optionally, it also includes: A pre-set Chinese retrieval dataset is mixed into the benchmark dataset to construct a test training sample set; The test training sample set is input into the solid waste-specific text embedding model to obtain the actual embedding results; The test training sample set is input into a pre-built large language model to construct a multi-dimensional stress test benchmark set; The effectiveness is evaluated by comparing the actual embedding results with the multi-dimensional stress test benchmark set.

[0014] A second aspect of the present invention provides an apparatus for constructing a dedicated text embedding model for solid waste management, comprising: The document conversion module is used to convert multi-format documents in a pre-built underlying corpus for solid waste management into structured text; The semantic segmentation module is used to perform semantically sensitive segmentation on the structured text to obtain the segmented structured text; The dual verification module is used to perform rule-based filtering and prompting engineering dual verification on the segmented structured text to obtain the triple dataset and the evaluation benchmark dataset. The supervised fine-tuning module is used to perform full-parameter supervised fine-tuning of the pre-built initial pedestal model using the triplet dataset and the benchmark dataset to obtain a text embedding model specifically for solid waste.

[0015] Optionally, the underlying corpus for solid waste management is constructed based on a classification system of eight dimensions: basic concepts of solid waste, laws and policies, management system, transfer and collection, disposal technology, risk assessment, zero-waste cities and circular economy, and artificial intelligence applications.

[0016] Optionally, the dual verification module includes: The removal unit is used to remove residual information such as hyperlinks, non-text symbols, bulk references, and format errors from the segmented structured text using a filter, so as to obtain the filtered structured text. The determination unit is used to determine the solid waste relevance of the filtered structured text through the prompting process to obtain the triplet dataset and the benchmark dataset, wherein both the triplet dataset and the benchmark dataset include multiple query-positive-negative triplets.

[0017] Optionally, the supervised fine-tuning module includes: The training sample set unit is used to mix a preset Chinese retrieval dataset into the triplet dataset to construct a training sample set; The supervised fine-tuning unit is used to perform full-parameter supervised fine-tuning of the pre-constructed initial pedestal model using the training sample set based on the contrastive learning loss function and the cosine annealing scheduling strategy, so as to obtain the solid waste-specific text embedding model.

[0018] Optionally, it also includes: The test sample set construction module is used to mix a preset Chinese retrieval dataset into the benchmark dataset to construct a test training sample set; The actual embedding module is used to input the test training sample set into the solid waste-specific text embedding model to obtain the actual embedding result; The test embedding module is used to input the test training sample set into a pre-built large language model to construct a multi-dimensional stress test benchmark set. The evaluation module is used to compare and evaluate the effects based on the actual embedding results and the multi-dimensional stress test benchmark set.

[0019] This invention presents a method and apparatus for constructing a dedicated text embedding model for solid waste management. By building a supervised fine-tuning framework based on a professional corpus of eight core dimensions of solid waste management and a Large Language Model (LLM)-driven technique for generating difficult-to-bear samples, it effectively solves the problems of inaccurate representation and weak anti-interference ability of general embedding models when dealing with low-frequency terms and complex policy logic in the solid waste vertical domain. It significantly improves the discriminativeness of the semantic space in the vertical domain. The developed solid waste-specific embedding model WuYu-E, while maintaining the ability to understand general language, greatly enhances the model's retrieval accuracy, classification accuracy, and semantic representation robustness in a real-world environment through the application of full-parameter supervised fine-tuning, adaptive sampling, and specialized loss functions. It provides quantifiable underlying support for the intelligent transformation of solid waste management combined with large language description and high-performance retrieval enhancement generation systems.

[0020] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0021] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 A flowchart illustrating a method for constructing a dedicated text embedding model for solid waste management according to an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating the specific execution of a method for constructing a dedicated text embedding model for solid waste management according to an embodiment of the present invention; Figure 3 This is a block diagram of an apparatus for constructing a dedicated text embedding model for solid waste management according to an embodiment of the present invention.

[0022] Explanation of reference numerals in the attached figures: 30 - Construction device for a dedicated text embedding model for solid waste management; 301 - Document conversion module; 302 - Semantic segmentation module; 303 - Dual verification module; 304 - Supervised fine-tuning module. Detailed Implementation

[0023] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0024] The following description, with reference to the accompanying drawings, illustrates a method and apparatus for constructing a dedicated text embedding model for solid waste management according to embodiments of the present invention. Addressing the key technical bottlenecks mentioned in the background section, such as insufficient accuracy and limited semantic understanding depth in existing general-purpose natural language processing models when facing highly specialized terminology recognition, deep policy semantic analysis, and complex technical document classification in the solid waste field, this invention provides a method for constructing a dedicated text embedding model for solid waste management. This method constructs a domain-specific corpus covering eight core knowledge topics in solid waste management and develops a fine-tuning data synthesis technique driven by self-supervised learning and a large language model (LLM). This achieves automated transformation from original professional documents to high-quality "query-positive example-negative example" triple datasets, thereby guiding the model to systematically improve its capabilities in intelligent retrieval, accurate matching, and knowledge integration of professional texts in the solid waste management field.

[0025] Specifically, Figure 1 This is a flowchart illustrating a method for constructing a dedicated text embedding model for solid waste management, as provided in an embodiment of the present invention.

[0026] like Figure 1 As shown, the method for constructing this specialized text embedding model for solid waste management includes the following steps: In step S101, the multi-format documents in the pre-built underlying corpus for solid waste management are converted into structured text.

[0027] In step S102, semantically sensitive segmentation is performed on the structured text to obtain the segmented structured text.

[0028] In some embodiments, the underlying corpus for solid waste management is constructed based on a classification system encompassing eight dimensions: basic concepts of solid waste, laws and policies, management systems, transfer and collection, disposal technologies, risk assessment, zero-waste cities and circular economy, and artificial intelligence applications.

[0029] In practical implementation, considering the highly fragmented nature and wide professional scope of knowledge in the field of solid waste management, this invention constructs a classification system covering eight core dimensions: basic concepts of solid waste, laws and policies, management systems, transfer and collection, disposal technologies, risk assessment, waste-free cities and circular economy, and artificial intelligence applications. Its core logic lies in breaking down semantic confusion in general models when processing solid waste data through refined topic segmentation, ensuring that the model can deeply cover the entire industrial chain knowledge network from macro-policy guidance to micro-process flow. During construction, the system utilizes a large language model as an intelligent classification engine. By analyzing the deep semantic features of text blocks and comparing them with a pre-set professional knowledge graph, it accurately directs massive amounts of unstructured corpus to their respective topics. This classification mapping mechanism not only effectively solves the long-tail effect of uneven data distribution in vertical fields but also significantly improves the discriminative power and retrieval accuracy of the vector space in different professional contexts through topic-enhanced training.

[0030] Specifically, such as Figure 2 As shown, in the data source screening stage, this embodiment of the invention follows the principle of giving equal importance to authority and diversity. Based on a classification system of eight dimensions of solid waste, including basic concepts, laws and policies, management systems, transfer and collection, disposal technologies, risk assessment, zero-waste cities and circular economy, and artificial intelligence applications, it systematically integrates policy and regulatory documents from the official websites of governments and environmental protection departments at all levels, open-source professional dissertations and journal articles from academic platforms such as CNKI and Web of Science, "zero-waste city" planning reports from government gazettes, and structured basic theoretical books from public e-book libraries. Through these bilingual (Chinese and English) data covering policy, scientific research, and practical dimensions, a low-level corpus for solid waste management with an international perspective and accurate cross-language semantic alignment is constructed.

[0031] Furthermore, the advanced MinerU tool can be used to convert the massive amounts of PDF, DOCX, and other format documents in the underlying corpus into clean structured text. Then, the SentenceSplitter within the LlamaIndex framework is introduced to perform semantically sensitive segmentation of the structured text. Unlike traditional fixed-length chunking, this embodiment of the invention employs a hybrid chunking strategy during semantically sensitive segmentation. This strategy automatically adjusts the segmentation position based on the paragraph logic and syntactic structure of the text, ensuring that the resulting segmented structured text retains semantic integrity to the maximum extent while maintaining chunk lengths that meet the input constraints of the Transformer model.

[0032] In step S103, the segmented structured text is subjected to rule-based filtering and prompting engineering for dual verification to obtain a triplet dataset and an evaluation benchmark dataset.

[0033] In some embodiments, the segmented structured text undergoes dual validation using rule-based filtering and cue engineering to obtain a triplet dataset and a benchmark dataset, including: The filter removes residual information such as hyperlinks, non-text symbols, bulk references, and formatting errors from the segmented structured text to obtain the filtered structured text. By proposing a method to determine the solid waste relevance of the filtered structured text, a triplet dataset and a benchmark dataset are obtained. Both the triplet dataset and the benchmark dataset include multiple query-positive-negative triplets.

[0034] In actual implementation, such as Figure 2 As shown, to ensure that the data entering the training phase has extremely high quality, this embodiment of the invention uses a dual verification based on rule-based filtering and prompt engineering. The filter is responsible for removing residual information such as hyperlinks, non-text symbols, batch citations, and format errors from the text, while prompt engineering drives a large language model to determine the relevance of the text content. Through targeted classification prompts, each piece of text is accurately assigned to eight predefined core research topics to obtain a triplet dataset composed of multiple query-positive-negative triplets and an evaluation benchmark dataset.

[0035] In step S104, the pre-built initial pedestal model is fine-tuned with full parameter supervision using the triple dataset to obtain a text embedding model specifically for solid waste.

[0036] In some embodiments, a pre-built initial pedestal model is fine-tuned with full parameter supervision using a triplet dataset and a benchmark dataset to obtain a text embedding model specifically for solid waste, including: A pre-defined Chinese retrieval dataset is mixed into the triplet dataset to construct a training sample set; Based on a contrastive learning loss function and a cosine annealing scheduling strategy, a pre-built initial pedestal model is fine-tuned with full parameter supervision using a training sample set to obtain a text embedding model specifically for solid waste.

[0037] In practical implementation, the core of this invention lies in establishing a complete solid waste domain embedding representation learning system. This system can support intelligent agent systems in real-time response, long-term memory construction, and knowledge base invocation under privacy protection. The technical solution proposed in this invention adopts a supervised fine-tuning contrastive learning paradigm. The entire framework starts from the underlying data preprocessing, intervening in refined semantic unit segmentation and standardized governance of multi-source heterogeneous data, ensuring that every knowledge fragment extracted from bilingual (Chinese and English) materials such as professional books, research reports, and policy documents maintains extremely high semantic integrity. Addressing the low-frequency professional terminology, conditional regulatory clauses, and complex technical logic structures unique to the solid waste domain, this invention designs a specialized loss function combination and adaptive sampling strategy. By bringing semantically similar samples closer together and pushing away interfering negative samples in the vector space, contrastive learning guides the model to capture deep-level vertical domain semantic relationships, thereby providing solid infrastructure support for the intelligent upgrading of all aspects of solid waste management.

[0038] This invention also proposes an LLM-driven difficult negative sample generation technique. This technique no longer relies on simple random sampling or lexical overlap filtering. Instead, it utilizes a large language model with optimized prompting engineering to simulate the role of a "domain expert." Through semantic generalization, logical reasoning, and role-playing, it generates logically rigorous question pairs (Query) with different perspectives for each positive example paragraph. Furthermore, this invention extracts Query keywords and combines them with expert knowledge graph review to uncover "difficult negative examples" that are highly related in topic or keyword but semantically mismatched. Simultaneously, it uses generative AI to construct "adversarial negative examples" that are highly similar in style and core logic but semantically inconsistent. This multi-layered, multi-dimensional negative sample construction strategy greatly enriches the discriminative depth of the training corpus, enabling the model to accurately identify subtle differences in the meaning of domain-specific texts, thereby fundamentally improving the model's robustness and generalization boundary.

[0039] Furthermore, this invention also introduces a category quota management and scarce category data synthesis method. Considering the potential scarcity of data in emerging themes or specific policy texts within the solid waste field, such as "circular economy" and "artificial intelligence applications," this invention designs an adaptive batch sampling strategy. Within a single sampling period, it increases the exposure frequency of scarce category subsets during training by performing targeted enhancement and repeated sampling. This strategy effectively prevents the model from developing a bias towards data-rich categories during training, ensuring balanced learning of the eight core themes throughout the entire lifecycle of solid waste management. Simultaneously, to avoid the potential for "catastrophic forgetting" during deep specialization, this invention consciously retains and integrates mixed text from general retrieval datasets and non-solid waste domains during the fine-tuning stage. This achieves a perfect balance between domain specialization and general semantic understanding capabilities, ensuring that the model maintains top-tier ranking quality and recall accuracy when handling complex cross-domain, multilingual retrieval tasks.

[0040] Specifically, such as Figure 2 As shown, a pre-defined Chinese retrieval dataset is first mixed into the triplet dataset to construct training and test training sample sets. In this embodiment, general-purpose Chinese retrieval datasets such as T2Retrieval and MMarco Retrieval can be mixed into the triplet dataset. By setting a reasonable mixing ratio, it is ensured that while the model improves its knowledge in the solid waste domain, it still maintains robust understanding of general-purpose language, ensuring that the model can still provide logically coherent and semantically accurate retrieval results when faced with fuzzy queries from non-professional users.

[0041] Furthermore, in this embodiment of the invention, BAAI / BGE-M3 is selected as the initial base model. It not only has powerful Transformer architecture support, but also has the natural advantage of multilingual and long text modeling, which enables it to adapt to the complex professional terminology environment in the field of solid waste.

[0042] Supervised fine-tuning is key to improving the model's task adaptability in this embodiment of the invention. This embodiment employs a full-parameter fine-tuning strategy at this stage, utilizing the previously constructed high-quality question-answering training sample set (Query-Positives-Negatives) to deeply reshape the initial base model, thereby obtaining the initial solid waste-specific text embedding model.

[0043] During supervised fine-tuning, the initial pedestal model is guided to generate specialized questions with deep reasoning properties using a triplet dataset. By constructing queries with high semantic complexity, the spatial distance between similar corpora in the vector space is reduced, improving the model's ability to distinguish specialized text corpora. Subsequently, keyword mining techniques are used to retrieve paragraphs with high lexical overlap but subtle semantic differences from the retrieval pool as difficult negative examples, forcing the model to learn more granular feature discrimination capabilities during fine-tuning.

[0044] To ensure the model's optimal performance on specialized tasks, this embodiment of the invention introduces a specialized contrastive learning loss function during the fine-tuning stage, and combines it with a cosine annealing learning rate scheduling strategy. By setting a very small temperature coefficient, the model is forced to construct a steeper classification boundary in the vector space, thereby obtaining a text embedding model specifically for solid waste (i.e., the WuYu-E model).

[0045] In some embodiments, it also includes: A pre-set Chinese retrieval dataset is mixed into the benchmark dataset to construct a test training sample set; The test training sample set is input into the solid waste-specific text embedding model to obtain the actual embedding results; The test training sample set is input into a pre-built large language model to construct a multi-dimensional stress test benchmark set; The effectiveness is evaluated by comparing actual embedding results with a multi-dimensional stress test benchmark set.

[0046] In practical implementation, regarding the construction of the data quality verification and performance evaluation system, this embodiment of the invention also constructs a comprehensive benchmark evaluation system specifically for embedded models in the field of solid waste management, namely a solid waste management-specific evaluation benchmark (SWM Benchmark) that far exceeds general standards. This benchmark not only includes traditional retrieval accuracy testing but also specifically designs classification tasks and semantic similarity measurement tasks. In the classification error correction task, seemingly reasonable but misclassified "pseudo-labels" are deliberately generated using a large model to test whether the model can identify the conflict between the text and the incorrect category, thereby verifying the model's accurate perception of the classification boundary. In other words, this embodiment of the invention breaks through the limitations of traditional evaluation that only focuses on single retrieval accuracy, and conducts extreme stress tests on the model from three dimensions: semantic similarity measurement, professional category classification, and complex retrieval. The introduction of "adversarial samples" forces the model to have extremely high resolution in the vector space, thereby providing a quantitative basis for evaluating the reliability of the model in vertical domain practice.

[0047] Specifically, such as Figure 2As shown, a pre-set Chinese retrieval dataset can be mixed into the benchmark dataset to construct a test training sample set. The test training sample set is then input into the initial solid waste-specific text embedding model to obtain the actual embedding results. The test training sample set is then input into a pre-built large language model to construct a multi-dimensional stress test benchmark set. Finally, the performance is compared and evaluated based on the actual embedding results and the multi-dimensional stress test benchmark set, which verifies the model's accurate perception of the classification boundary.

[0048] The following section provides a detailed explanation of the construction method for the dedicated text embedding model for solid waste management proposed in this invention through a specific retrieval task.

[0049] In the evaluation of the retrieval task, this embodiment uses NDCG@10 and Top-K Accuracy as core metrics. By mixing random negative examples, difficult negative examples, and adversarial negative examples among tens of thousands of interference texts, a "stress test" environment simulating a complex real-world application scenario is constructed. Experimental data shows that on a custom retrieval benchmark in the solid waste field, the WuYu-E model not only significantly outperforms the original base model, but its NDCG metric is also significantly improved. Furthermore, in the evaluation of the semantic text similarity task, this invention also uses Pearson correlation coefficients to deeply quantify the scientific validity of the embedding vector space structure. By comparing the consistency between the vector distance ranking predicted by the model and the semantic similarity ranking given by human experts, the superiority of the WuYu-E model at the vector representation level is demonstrated.

[0050] This invention comprehensively incorporates the general retrieval set DuRetrieval, the classification standard set STSB, and the aforementioned constructed vertical domain dataset to form a multi-dimensional evaluation benchmark, ensuring the robustness of the model under different tasks. The performance of the WuYu-E model on the SWM benchmark and other general embedding models on general benchmarks is shown in Table 1 below: Table 1 Overall accuracy metrics for different general models

[0051] In specific engineering deployment embodiments, this invention utilizes the sentence-transformers framework as the underlying engine, coupled with the AdamW optimizer and TF32 / BF16 accuracy support, to achieve high efficiency and robustness in the model training process. All hyperparameters used in the training process have undergone rigorous screening through multiple rounds of ablation experiments. Through real-time monitoring of the NDCG curve, Loss curve, and Top-k Accuracy curve, this invention verifies that the model exhibits excellent convergence stability during fine-tuning. This standardized technical implementation paradigm can not only be applied to the field of solid waste management but also provides a replicable and scalable engineering reference for building high-performance embedded models in other highly specialized vertical fields.

[0052] In summary, the method for constructing a dedicated text embedding model for solid waste management proposed in this embodiment of the invention effectively solves the problems of inaccurate representation and weak anti-interference ability of general embedding models when dealing with low-frequency terms and complex policy logic in the solid waste vertical domain by constructing a supervised fine-tuning framework based on a professional corpus of eight core dimensions of solid waste management and a large language model LLM-driven hard-to-bear sample generation technology. It significantly improves the discriminativeness of the semantic space in the vertical domain. The developed solid waste-specific embedding WuYu-E model, while maintaining the ability to understand general language, greatly enhances the model's retrieval accuracy, classification accuracy, and semantic representation robustness in the real-world environment through the application of full-parameter supervised fine-tuning, adaptive sampling, and specialized loss functions. It provides quantifiable underlying support for the intelligent transformation of solid waste management combined with large language description and the high-performance retrieval enhancement generation system.

[0053] Next, with reference to the accompanying drawings, an apparatus for constructing a dedicated text embedding model for solid waste management according to an embodiment of the present invention is described.

[0054] Figure 3 This is a block diagram illustrating a construction apparatus for a dedicated text embedding model for solid waste management, provided in an embodiment of the present invention.

[0055] like Figure 3 As shown, the construction device 30 for the dedicated text embedding model for solid waste management includes: a document conversion module 301, a semantic segmentation module 302, a dual verification module 303, and a supervised fine-tuning module 304.

[0056] The document conversion module 301 converts multi-format documents from a pre-built underlying corpus for solid waste management into structured text. The semantic segmentation module 302 performs semantically sensitive segmentation on the structured text to obtain segmented structured text. The dual verification module 303 performs rule-based filtering and prompting engineering dual verification on the segmented structured text to obtain a triplet dataset and a benchmark dataset. The supervised fine-tuning module 304 uses the triplet dataset and benchmark dataset to perform fully parameter-supervised fine-tuning of the pre-built initial base model to obtain a text embedding model specifically for solid waste.

[0057] In some embodiments, the underlying corpus for solid waste management is constructed based on a classification system encompassing eight dimensions: basic concepts of solid waste, laws and policies, management systems, transfer and collection, disposal technologies, risk assessment, zero-waste cities and circular economy, and artificial intelligence applications.

[0058] In some embodiments, the dual verification module 303 includes: The removal unit is used to remove residual information such as hyperlinks, non-text symbols, bulk references, and formatting errors from the segmented structured text using filters, so as to obtain the filtered structured text. The decision unit is used to determine the solid waste relevance of the filtered structured text through the prompting process to obtain a triplet dataset and a benchmark dataset. Both the triplet dataset and the benchmark dataset include multiple query-positive-negative triplets.

[0059] In some embodiments, the monitoring and fine-tuning module 304 includes: The training sample set unit is used to mix a pre-defined Chinese retrieval dataset into the triplet dataset to construct the training sample set; The supervised fine-tuning unit is used to perform full-parameter supervised fine-tuning of the pre-built initial pedestal model based on the contrastive learning loss function and cosine annealing scheduling strategy, using the training sample set to obtain a text embedding model specifically for solid waste.

[0060] In some embodiments, it also includes: The test sample set construction module is used to mix a pre-defined Chinese search dataset into the benchmark dataset to construct a test training sample set; The actual embedding module is used to input the test training sample set into the solid waste-specific text embedding model to obtain the actual embedding results; The test embedding module is used to input the test training sample set into a pre-built large language model to construct a multi-dimensional stress test benchmark set; The evaluation module is used to compare and evaluate the effectiveness based on the actual embedding results and a multi-dimensional stress test benchmark set.

[0061] It should be noted that the foregoing explanation of the embodiment of the method for constructing a dedicated text embedding model for solid waste management also applies to the apparatus for constructing a dedicated text embedding model for solid waste management in this embodiment, and will not be repeated here.

[0062] The device for constructing a dedicated text embedding model for solid waste management proposed in this embodiment of the invention effectively solves the problems of inaccurate representation and weak anti-interference ability of general embedding models when dealing with low-frequency terms and complex policy logic in the solid waste vertical domain by constructing a supervised fine-tuning framework based on a professional corpus of eight core dimensions of solid waste management and a large language model LLM-driven hard-to-bear sample generation technology. It significantly improves the distinguishability of the semantic space in the vertical domain. The developed solid waste-specific embedding WuYu-E model, while maintaining the ability to understand general language, greatly enhances the model's retrieval accuracy, classification accuracy and semantic representation robustness in the real-world environment through the application of full-parameter supervised fine-tuning, adaptive sampling and specialized loss functions. It provides quantifiable underlying support for the intelligent transformation of solid waste management combined with large language description and high-performance retrieval enhancement generation system.

[0063] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0064] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0065] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of the invention pertain.

[0066] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus or device (such as a computer-based system, a processor-included system or other system that can fetch and execute instructions from, an instruction execution system, apparatus or device).

Claims

1. A method for constructing a specialized text embedding model for solid waste management, the method comprising: Includes the following steps: Convert multi-format documents from a pre-built corpus for solid waste management into structured text; The structured text is subjected to semantically sensitive segmentation to obtain the segmented structured text; The segmented structured text is subjected to rule-based filtering and prompting engineering for dual verification to obtain a triplet dataset and a benchmark dataset. The pre-built initial pedestal model was fine-tuned with full parameter supervision using the triple dataset to obtain a text embedding model specifically for solid waste.

2. The method of claim 1, wherein the method is a method of constructing a solid waste management oriented specialized text embedding model, characterized by, The underlying corpus for solid waste management is constructed based on a classification system encompassing eight dimensions: basic concepts of solid waste, laws and policies, management systems, transfer and collection, disposal technologies, risk assessment, zero-waste cities and circular economy, and artificial intelligence applications. 3.The method of claim 1, wherein, The segmented structured text undergoes dual validation using rule-based filtering and prompting engineering to obtain a triplet dataset and a benchmark dataset, including: The filtered structured text is obtained by removing hyperlinks, non-text symbols, bulk references, and residual information with formatting errors from the segmented structured text. The filter-based structured text is relevance to solid waste by a prompting process to obtain the triplet dataset and the benchmark dataset. Both the triplet dataset and the benchmark dataset include multiple query-positive-negative triplets. 4.The method of claim 1, wherein, The step of performing fully parameter-supervised fine-tuning of the pre-built initial base model using the triplet dataset and the benchmark dataset to obtain a text embedding model specifically for solid waste includes: A preset Chinese retrieval dataset is mixed into the triplet dataset to construct a training sample set; Based on the contrastive learning loss function and cosine annealing scheduling strategy, the pre-constructed initial base model is fine-tuned with full parameter supervision using the training sample set to obtain the solid waste-specific text embedding model. 5.The method of claim 1, wherein, Also includes: A pre-set Chinese retrieval dataset is mixed into the benchmark dataset to construct a test training sample set; The test training sample set is input into the solid waste-specific text embedding model to obtain the actual embedding results; The test training sample set is input into a pre-built large language model to construct a multi-dimensional stress test benchmark set; The effectiveness is evaluated by comparing the actual embedding results with the multi-dimensional stress test benchmark set. 6.A device for constructing a specialized text embedding model for solid waste management, characterized in that, include: The document conversion module is used to convert multi-format documents in a pre-built underlying corpus for solid waste management into structured text; The semantic segmentation module is used to perform semantically sensitive segmentation on the structured text to obtain the segmented structured text; The dual verification module is used to perform rule-based filtering and prompting engineering dual verification on the segmented structured text to obtain the triple dataset and the evaluation benchmark dataset. The supervised fine-tuning module is used to perform full-parameter supervised fine-tuning of the pre-built initial pedestal model using the triplet dataset and the benchmark dataset to obtain a text embedding model specifically for solid waste.

7. The apparatus for constructing a dedicated text embedding model for solid waste management according to claim 6, characterized in that, The underlying corpus for solid waste management is constructed based on a classification system encompassing eight dimensions: basic concepts of solid waste, laws and policies, management systems, transfer and collection, disposal technologies, risk assessment, zero-waste cities and circular economy, and artificial intelligence applications.

8. The apparatus for constructing a dedicated text embedding model for solid waste management according to claim 6, characterized in that, The dual verification module includes: The removal unit is used to remove residual information such as hyperlinks, non-text symbols, bulk references, and format errors from the segmented structured text using a filter, so as to obtain the filtered structured text. The determination unit is used to determine the solid waste relevance of the filtered structured text through the prompting process to obtain the triplet dataset and the benchmark dataset, wherein both the triplet dataset and the benchmark dataset include multiple query-positive-negative triplets.

9. The apparatus for constructing a dedicated text embedding model for solid waste management according to claim 6, characterized in that, The monitoring and fine-tuning module includes: The training sample set unit is used to mix a preset Chinese retrieval dataset into the triplet dataset to construct a training sample set; The supervised fine-tuning unit is used to perform full-parameter supervised fine-tuning of the pre-constructed initial pedestal model using the training sample set based on the contrastive learning loss function and the cosine annealing scheduling strategy, so as to obtain the solid waste-specific text embedding model.

10. The apparatus for constructing a dedicated text embedding model for solid waste management according to claim 6, characterized in that, Also includes: The test sample set construction module is used to mix a preset Chinese retrieval dataset into the benchmark dataset to construct a test training sample set; The actual embedding module is used to input the test training sample set into the solid waste-specific text embedding model to obtain the actual embedding result; The test embedding module is used to input the test training sample set into a pre-built large language model to construct a multi-dimensional stress test benchmark set. The evaluation module is used to compare and evaluate the effects based on the actual embedding results and the multi-dimensional stress test benchmark set.