A method for constructing a chitosan oligosaccharide preparation process multi-index comprehensive evaluation database
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-19
- Publication Date
- 2026-08-11
AI Technical Summary
[0007]本发明的目的在于提供一种壳寡糖制备工艺多指标综合评价数据库的构建方法,以解决上述背景技术中提出现有壳寡糖制备研究中存在的数据来源分散、工艺表达不统一、指标口径不一致、深层次数据冲突难以识别、不同工艺路线之间难以横向比较、现有数据库缺乏内置评价支持能力的问题
1、本发明实现了壳寡糖制备工艺相关数据的系统归集与结构化组织,解决了现有壳寡糖制备工艺数据分散、来源异构和难以统一整理的问题,为后续工艺研究提供了集中、规范的数据基础。
Smart Images

Figure CN122551969A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of process database construction and multi-index evaluation technology, specifically a method for constructing a comprehensive evaluation database of multiple indicators for chitosan oligosaccharide preparation processes. Background Technology
[0002] Chitosan oligosaccharides are a class of oligosaccharide products obtained from the degradation of chitosan. They possess good water solubility, biocompatibility, and bioactivity, and have significant application value in food, medicine, agriculture, and environmental protection. In recent years, with the increasing demand for resource utilization of crustacean waste such as shrimp and crab shells, the preparation of chitosan oligosaccharides from shrimp and crab shells has gradually become a research hotspot. Existing methods for preparing chitosan oligosaccharides mainly include acid degradation, oxidative degradation, enzymatic hydrolysis, physical-assisted methods, and chemical-enzymatic coupling methods. Since different preparation methods vary significantly in terms of raw material applicability, reaction conditions, product composition, resource consumption, and environmental impact, how to systematically and uniformly evaluate different processes has become a problem that needs to be solved in chitosan oligosaccharide preparation research. In existing studies, the judgment of process quality is mostly based on single indicators such as yield, reaction time, average molecular weight, or the proportion of the target degree of polymerization range. The evaluation methods are relatively scattered, making it difficult to comprehensively reflect the overall level of the process and hindering horizontal comparisons between different process routes.
[0003] In recent years, database technology and data-driven methods have been widely applied in the fields of biomanufacturing and process engineering. Process databases have become an important technical means to support the evaluation, optimization analysis, and model building of complex processes. To date, although research on chitosan oligosaccharide preparation has accumulated a wealth of experimental results and process parameters, the relevant data are mostly scattered across different literature, patents, and experimental records, lacking a systematic chitosan oligosaccharide preparation process database oriented towards comprehensive evaluation of multiple indicators. Furthermore, due to differences in raw material sources, process routes, indicator definitions, data units, and expression methods among different studies, directly integrating existing data can easily lead to problems such as information gaps, indicator conflicts, and incomparable results. In addition, some studies also suffer from incomplete indicators, delayed data updates, and unclear process boundaries.
[0004] More importantly, the heterogeneity of chitosan oligosaccharide preparation process data is not limited to terminology and units. Deep-seated "implicit conflicts" often exist between data from different sources. For example, a higher yield in the same literature may be primarily attributed to the lower molecular weight or higher degree of deacetylation of the raw chitosan, rather than the efficiency of the process itself. Different studies may use completely different system boundaries when calculating "energy consumption per unit product," some including energy consumption for raw material drying and pulverization, while others only calculate energy consumption for the reaction stage. These deep-seated data conflicts cannot be effectively identified and resolved using only general data cleaning and standardization methods; they require the introduction of knowledge of the process mechanism of chitosan oligosaccharide degradation for judgment and correction.
[0005] Furthermore, most existing process databases remain at the "data storage" level, lacking the capability for process analysis and evaluation. After querying data, users still need to manually extract data, construct evaluation matrices, and select evaluation methods to complete process comparison and optimization analysis. This data preparation process is time-consuming, inefficient, and the evaluation results are significantly influenced by subjective factors.
[0006] Therefore, existing scattered data cannot directly support a unified quantitative evaluation of chitosan oligosaccharide preparation processes, nor can it provide a reliable basis for comparative analysis of different process routes and subsequent process optimization. It is necessary to construct a method for building a multi-index comprehensive evaluation database for chitosan oligosaccharide preparation processes. This method involves systematically collecting process data, performing mechanism-driven conflict detection and consistency repair, standardization, classification coding, and correlation integration, and pre-configuring evaluation support capabilities within the database. This will establish a dedicated database for multi-index comprehensive evaluation, providing a data foundation for the comprehensive evaluation and subsequent application of chitosan oligosaccharide preparation processes. Summary of the Invention
[0007] The purpose of this invention is to provide a method for constructing a comprehensive evaluation database of multiple indicators for chitosan oligosaccharide preparation processes, in order to solve the problems mentioned in the background art that exist in existing chitosan oligosaccharide preparation research, such as scattered data sources, inconsistent process expressions, inconsistent indicator standards, difficulty in identifying deep-seated data conflicts, difficulty in making horizontal comparisons between different process routes, and lack of built-in evaluation support capabilities in existing databases.
[0008] To achieve the above objectives, the present invention provides the following technical solution: a method for constructing a comprehensive evaluation database of multiple indicators for chitosan oligosaccharide preparation processes, the method comprising: Step 1: Collect multi-source data related to the preparation process of chitosan oligosaccharide and construct the original process dataset. The multi-source data includes at least literature data, experimental record data and process data. Step 2: Extract fields, filter and integrate data from the original process dataset. During the integration process, based on the predefined chitosan oligosaccharide degradation kinetic rules and the unified full life cycle system boundary, perform process mechanism conflict detection and consistency repair on process data from different sources to form an effective set of process entries. Step 3: Standardize terminology, units, perform anomaly checks and missing value processing on the set of valid process entries to form a standardized process dataset; Step 4: Based on the comprehensive evaluation requirements of multiple indicators in the preparation process of chitosan oligosaccharide, the standardized process dataset is classified, coded, mapped, and hierarchically organized to construct a database schema structure; wherein, the relation mapping includes establishing the association between process entities, parameter entities, indicator entities, and source entities, and pre-calculating and storing sensitivity coefficients reflecting the degree of influence of process parameters on the comprehensive evaluation results in the association; Step 5: Based on the database schema structure, store the standardized process dataset as a dedicated database for comprehensive evaluation of multiple indicators of chitosan oligosaccharide preparation process; Step 6: Extract the evaluation object set and index set based on the dedicated database, generate a process evaluation matrix to support subsequent horizontal process comparison, unified quantitative evaluation and key parameter analysis, call the pre-stored sensitivity coefficients, and generate the process evaluation matrix and process optimization suggestions.
[0009] Further, step 1 includes: Step 11: Search publicly published literature and collect data on raw material sources, preparation routes, process parameters, and product performance in the preparation process of chitosan oligosaccharides; Step 12: Collect experimental data and extract reaction conditions, result indicators, and resource consumption data from the chitosan oligosaccharide preparation experiment; Step 13: Collect process data, including energy consumption, water consumption, reagent consumption, waste liquid discharge and economic data related to the preparation of chitosan oligosaccharides; Step 14: Perform deduplication, initial screening, and source verification on the collected data to construct the original process dataset.
[0010] This invention does not simply accumulate data during the data acquisition phase. Instead, it proactively collects data in advance, considering dimensions such as process route, process parameters, product quality, resource consumption, environmental impact, and economic efficiency, all in preparation for subsequent comprehensive evaluation based on multiple indicators. The reason for this is that if only a few result indicators such as yield and molecular weight are collected, the database, while usable as a reference, cannot truly support comprehensive evaluation based on multiple indicators. By pre-designing the data acquisition scope to align with the evaluation objectives, it ensures that the database possesses unified quantitative comparison capabilities in the future.
[0011] Furthermore, in step 2, the field extraction includes extracting the fields of raw material type, pretreatment method, preparation route, reaction temperature, reaction pH, reaction time, substrate concentration, liquid-solid ratio, enzyme addition amount, yield, percentage of target degree of polymerization range, average molecular weight, energy consumption, water consumption, waste liquid volume, and unit cost. The data filtering includes removing duplicate process entries, entries with missing key fields, and entries with unclear process boundaries; The process mechanism conflict detection and consistency repair specifically includes: when the ratio of yield to average molecular weight of the same process object in data from different sources exceeds the preset degradation kinetic threshold range, the raw material parameters are traced back and data alignment or weight reduction is performed; and for indicators involving unit resource consumption, the original data is traced back and corrected according to a unified full life cycle system boundary list, wherein the system boundary list defines whether it includes the energy and material consumption of raw material pretreatment, drying and pulverization processes.
[0012] The key step of "process mechanism conflict detection and consistency repair" in step 2 of this invention addresses the unique challenges in the field of chitosan oligosaccharide preparation process data. Conflicts between data from different sources in this field are often not merely superficial inconsistencies in format or units, but rather stem from information asymmetry at the process mechanism level. For example, the degradation process of chitosan follows certain kinetic laws, and there is an inherent constraint relationship between the product yield and its average molecular weight. When a record shows a high yield but the average molecular weight does not decrease accordingly, it may mean that the raw material chitosan itself has a low molecular weight or a high degree of deacetylation, rather than an efficient degradation process. If such implicit conflicts are ignored and data is directly entered into the database, subsequent evaluation results will be systematically biased. Therefore, this invention introduces a conflict detection mechanism based on degradation kinetic rules to backtrack, align, or downweight abnormal data, ensuring the consistency of the mechanism in the data entering the database from the data source.
[0013] Meanwhile, different studies often use inconsistent system boundaries when calculating resource consumption indicators such as energy consumption and water consumption per unit product. Some studies only measure the energy consumption in the reactor stage, while others include the energy consumption of pretreatment processes such as raw material drying and crushing. Direct comparison without boundary unification will lead to incomparable evaluation results. This invention predefines a unified full life cycle system boundary list and backtracks and corrects the original data to ensure that all resource consumption indicators are compared under the same boundary conditions.
[0014] Further, step 3 includes: Step 31: Map synonymous process terms from different data sources to unified standard terms; Step 32: Convert temperature, time, concentration, yield, molecular weight, energy consumption, and water consumption into a unified unit of measurement; Step 33: Identify anomalous data using rule-based validation, range validation, and source cross-validation; Step 34: Complete and mark missing data in non-critical fields, and remove or reduce the weight of missing data in critical fields to form a standardized process dataset.
[0015] Step 3, following the mechanistic conflict resolution in Step 2, further standardizes the data representation, unifying the terminology and units of data from different sources, thus laying the foundation for subsequent database construction and evaluation analysis.
[0016] Furthermore, the classification coding in step 4 includes establishing a unified coding system for raw material type, preparation route type, process parameter type, evaluation index type, and data source type; The hierarchical organization includes constructing a multi-index comprehensive evaluation hierarchical structure consisting of a target layer, a criterion layer, and an indicator layer; wherein, the target layer is a comprehensive evaluation of the chitosan oligosaccharide preparation process, and the criterion layer includes at least product quality, reaction efficiency, resource consumption, environmental impact, and economic efficiency.
[0017] The reason for using a database schema structure instead of a single table is that the database of this invention is not an ordinary data table, but a database specifically designed for subsequent multi-indicator comprehensive evaluation and process analysis. While a simple table format can store data, it cannot efficiently express the hierarchical and relational relationships between process objects and evaluation indicators, nor is it conducive to subsequent evaluation matrix extraction and database expansion and updates. Therefore, this invention introduces classification coding, relational mapping, and hierarchical organization mechanisms, giving the database a structured feature that separates the schema layer and the data layer, thereby improving database maintainability and application capabilities.
[0018] Furthermore, the sensitivity coefficient pre-calculated in step 4 is obtained by calculating the impact of changes in each process parameter on the comprehensive evaluation value based on historical data in the parameter table and index table, and storing the calculation results in the relational mapping for direct retrieval during database queries.
[0019] This invention pre-defines sensitivity coefficients within its database schema structure, a core feature distinguishing it from existing process databases. Existing process databases, after storing data, require users to manually export data, select analysis methods (such as regression analysis or sensitivity analysis), and perform calculations to understand which process parameters have the greatest impact on product quality or overall cost. This process is time-consuming and requires a certain level of data analysis expertise. This invention, however, pre-calculates the sensitivity coefficients of each process parameter to the overall evaluation results using data from existing parameter and indicator tables during the database construction phase, storing these coefficients as part of the database. Thus, when users query process data, the database simultaneously provides a sensitivity ranking of key parameters, directly supporting process optimization decisions and upgrading the database from a "passive storage" type to an "active analysis" type.
[0020] Furthermore, the dedicated database constructed in step 5 covers five types of chitosan oligosaccharide preparation routes: acid degradation, oxidative degradation, enzymatic hydrolysis, physical-assisted method, and chemical-enzymatic coupling method.
[0021] Furthermore, the indicator layer sets 16 evaluation indicators, which include: yield, percentage of the target degree of polymerization range, average molecular weight, purity, reaction time, substrate conversion rate, yield per unit time, energy consumption per unit product, water consumption per unit product, acid consumption per unit product, alkali consumption per unit product, enzyme consumption per unit product, waste liquid discharge, chemical oxygen demand, safety risk coefficient, and cost per unit product.
[0022] These 16 indicators correspond to five criteria levels: product quality, reaction efficiency, resource consumption, environmental impact, and economic efficiency. They can comprehensively characterize the overall performance of different chitosan oligosaccharide preparation processes from multiple dimensions.
[0023] Furthermore, the dedicated database includes a process master table, a parameter table, an indicator table, a source table, and a relationship mapping table; The process master table stores unique identifiers for process items and process route information; the parameter table stores process parameter information; the index table stores multi-index evaluation data; the source table stores data source information; and the relationship mapping table stores the relationships between entities and the pre-calculated sensitivity coefficients.
[0024] Furthermore, the method for generating the process evaluation matrix in step 6 includes: Based on the dedicated database, the evaluation object set and indicator set are extracted to construct the original evaluation matrix; The original evaluation matrix is then normalized and standardized according to the index attributes to generate a standardized evaluation matrix that can be used for horizontal comparison and unified quantitative evaluation of the process, so as to be used for subsequent unified quantitative evaluation, horizontal comparison and key parameter analysis of the chitosan oligosaccharide preparation process.
[0025] This invention considers the evaluation support interface during the database construction stage because if the database only stays at the "data storage" level, it is closer to a data management system. However, this invention designs the database structure directly for the evaluation matrix and pre-sets sensitivity coefficients in the database, so that the database can not only store process information, but also directly serve subsequent comprehensive evaluation and process optimization, thereby improving the practical application value of the database.
[0026] Compared with existing technologies, the construction method of this comprehensive evaluation database for the preparation process of chitosan oligosaccharides is not a simple data aggregation, but rather a structured database that can directly serve subsequent process evaluation matrix extraction, process comparison, and key parameter analysis through multi-source data acquisition, effective process item formation, data standardization, classification coding, relationship mapping, and evaluation support interface construction. Specific beneficial effects are as follows: 1. This invention realizes the systematic collection and structured organization of data related to the preparation process of chitosan oligosaccharides, solving the problems of scattered, heterogeneous sources and difficulty in unified organization of existing chitosan oligosaccharide preparation process data, and providing a centralized and standardized data foundation for subsequent process research.
[0027] 2. This invention solves the problem of deep-seated data conflicts that cannot be identified by existing general data cleaning methods by introducing a conflict detection and consistency repair mechanism based on process mechanisms. By verifying the kinetic constraints between yield and molecular weight, and unifying the system boundaries of resource consumption indicators throughout their entire lifecycle, the mechanism consistency and comparability of the data entering the database are guaranteed from the data source, effectively avoiding evaluation bias caused by implicit data conflicts.
[0028] 3. This invention improves the standardization, comparability, and calculability of process data from different sources. By extracting, screening, and integrating raw material sources, preparation routes, process parameters, product quality, resource consumption, environmental impact, and economic indicators from the original data, and by employing standardization steps such as terminology standardization, unit standardization, anomaly verification, and missing value handling, the differences in expression, units, and record completeness of data from different sources are effectively reduced, enabling process data in the database to be retrieved, compared, and analyzed under a unified standard.
[0029] 4. This invention constructs a dedicated database for comprehensive evaluation based on multiple indicators, capable of comprehensively characterizing the overall performance of different chitosan oligosaccharide preparation processes. The database covers five preparation routes: acid degradation, oxidative degradation, enzymatic hydrolysis, physical-assisted methods, and chemical-enzymatic coupling methods, and sets 16 evaluation indicators. It can uniformly characterize chitosan oligosaccharide preparation processes from multiple dimensions, including product quality, reaction efficiency, resource consumption, environmental impact, and economic efficiency. Compared with existing methods that rely solely on a single indicator to judge the quality of a process, this invention can more comprehensively and objectively reflect the overall level of different process routes.
[0030] 5. This invention pre-calculates and stores the sensitivity coefficients of process parameters to the comprehensive evaluation results within the database, upgrading the database from "passive storage" to "active analysis." When querying data, users can directly obtain the impact weight ranking of key parameters without having to export data and perform analysis modeling themselves, significantly reducing the data preparation threshold for process optimization analysis and improving the practical application efficiency of the database.
[0031] 6. The database constructed by this invention can directly support subsequent process evaluation, process comparison, and key parameter analysis, and has good application value. This invention considers the need for subsequent process evaluation matrix extraction during the database construction stage, enabling the database to not only have data storage and management functions, but also directly serve the horizontal comparison, unified quantitative evaluation, and key parameter analysis of chitosan oligosaccharide preparation processes. This reduces the data preparation work before constructing subsequent evaluation models, providing reliable support for the screening, optimization, and database management of chitosan oligosaccharide preparation processes, and has good application prospects. Attached Figure Description
[0032] Figure 1 This is a flowchart illustrating the overall construction process of the multi-index comprehensive evaluation database for the preparation process of chitosan oligosaccharides in this embodiment of the invention.
[0033] Figure 2 This is a distribution diagram of effective process items for the five types of chitosan oligosaccharide preparation routes in the embodiments of the present invention.
[0034] Figure 3 This is a schematic diagram of the structural model of the multi-index comprehensive evaluation database for the preparation process of chitosan oligosaccharides constructed in this embodiment of the invention.
[0035] Figure 4 This is a database hierarchy diagram of the multi-index comprehensive evaluation system in this embodiment of the invention.
[0036] Figure 5 This is a flowchart illustrating the transformation of raw data into valid process entries and structured data records in an embodiment of the present invention.
[0037] Figure 6 This is a flowchart of field integrity and confidence scoring in an embodiment of the present invention.
[0038] Figure 7 This is a schematic diagram of the evaluation matrix construction in an embodiment of the present invention.
[0039] Figure 8 This is a schematic diagram illustrating the calculation of index weights using the entropy weight method in an embodiment of the present invention.
[0040] Figure 9 This is a visualization of the criterion layer weight vector in an embodiment of the present invention.
[0041] Figure 10 This is a distribution diagram of the average route-level scores for the five types of chitosan oligosaccharide preparation routes in the embodiments of the present invention.
[0042] Figure 11 This is a schematic diagram of the sensitivity analysis of the multi-index comprehensive evaluation results in an embodiment of the present invention.
[0043] Figure 12This is a statistical chart showing the coverage of the database for the comprehensive evaluation of multiple indicators of the chitosan oligosaccharide preparation process in this embodiment of the invention.
[0044] Figure 13 This is a statistical chart showing the indicator coverage of structured data records in an embodiment of the present invention.
[0045] Figure 14 This is an overview diagram of the database and evaluation interface in an embodiment of the present invention. Detailed Implementation
[0046] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0047] Please see Figure 1-14 The present invention provides the following technical solutions: Example 1: Raw data collection and construction of effective process entries for a comprehensive evaluation database of chitosan oligosaccharide preparation process.
[0048] This embodiment provides a method for data acquisition and effective process entry construction for a multi-index comprehensive evaluation database of chitosan oligosaccharide preparation process. The method focuses on the comprehensive evaluation of multiple indicators of chitosan oligosaccharide preparation process, collecting literature data, experimental records, and process data related to the chitosan oligosaccharide preparation process. The raw data is then deduplicated, filtered, and integrated. The core of this method includes introducing a conflict detection and consistency repair mechanism based on the process mechanism during the integration process, resolving deep-seated data conflicts that traditional data cleaning methods cannot identify. Ultimately, this results in a set of effective process entries that can be directly used for subsequent database construction.
[0049] The current mainstream preparation methods for chitosan oligosaccharides mainly include chemical degradation, physical methods and enzymatic methods. Among them, oxidation methods and combined routes are also common research directions in recent years. Therefore, in this embodiment, acid degradation, oxidative degradation, enzymatic hydrolysis, physical-assisted methods and chemical-enzyme coupling methods are set as the five primary routes of the database.
[0050] In this embodiment, the literature data is preferably derived from high-quality publicly available papers related to chitosan oligosaccharide preparation, chitosan degradation, and process optimization, especially from relevant studies in journals such as *Carbohydrate Polymers*, *International Journal of Biological Macromolecules*, and *Bioresource Technology*. Experimental record data is derived from original laboratory experimental records. Process data is derived from process records related to unit energy consumption, unit water consumption, reagent consumption, wastewater discharge, and unit cost. Data quality rules preferably refer to the definitions of completeness, reliability, temporal representativeness, geographical representativeness, and technological representativeness in the ILCD data quality index system.
[0051] In this embodiment, a total of 100 articles and 396 sets of experimental records were screened and included, forming 863 valid process entries, which were further analyzed to obtain 4008 structured data records. The distribution of entries for the five preparation routes is as follows: acid degradation method 238 entries, oxidative degradation method 176 entries, enzymatic hydrolysis method 214 entries, physical-assisted method 91 entries, and chemical-enzymatic coupling method 144 entries.
[0052] Let the original source dataset be:
[0053] Where n=496, representing the total number of sources for the 100 articles and 396 sets of experimental records.
[0054] Let the set of effective process items after deduplication, screening, and integration be:
[0055] Where m=863.
[0056] Further define the source utilization coefficient:
[0057] In this embodiment:
[0058] The results indicate that a single article or a single set of experimental records can be broken down into multiple process objects with independent parameter boundaries and result indicators. Therefore, the core statistical caliber of the database should preferably be "the number of effective process entries" rather than "the number of source documents".
[0059] Let k be the number of structured field records corresponding to the i-th valid process entry.i Then the total number of structured data records N r Defined as:
[0060] In this embodiment:
[0061] The average number of structured field records per valid process entry is:
[0062] The 4008 structured records were divided into five categories: 612 raw material and source records, 1508 process parameter records, 876 product quality records, 621 resource and environmental records, and 391 economic and supplementary records. This partitioning method allows the database to retain the integrity of process objects while also providing field-level access capabilities.
[0063] For route r, let the number of its entries be n. r Then, the proportion of this route in the total effective process items is:
[0064] Substituting the number of route entries of the five categories in this embodiment, we can obtain:
[0065] The statistical results above show that acid degradation, oxidative degradation, and enzymatic hydrolysis account for a relatively high proportion of the effective process items included in this embodiment, indicating that current research on chitosan oligosaccharide preparation mainly focuses on chemical degradation and enzymatic routes. Physically assisted methods and chemical-enzymatic coupling methods account for a lower proportion, but still have a certain number of representative process items, which can be used for subsequent horizontal comparisons with mainstream process routes. This distribution result shows that the database constructed in this embodiment takes into account both mainstream technical routes and novel combined routes in the selection of data sources and item screening, and can comprehensively cover the main research directions of current chitosan oligosaccharide preparation processes. Furthermore, from the perspective of route coverage, all five preparation routes have independent item sizes, and there is no situation where the proportion of a single route item is too high, leading to a serious imbalance in the database structure. This is beneficial for the comparability analysis between different routes in subsequent multi-index comprehensive evaluation. Based on the above route distribution results, a route-level distribution map and a database coverage statistical table can be further formed, providing a data foundation for subsequent process evaluation matrix extraction, route-level comparative analysis, and key parameter identification.
[0066] Specific Implementation Method of Process Mechanism Conflict Detection and Consistency Repair in this Embodiment In the process of forming valid process entries, this embodiment introduces a process mechanism conflict detection and consistency repair mechanism, which is a key step that distinguishes it from conventional data integration. The specific implementation method is as follows: (i) Yield-molecular weight conflict detection based on degradation kinetics rules The degradation of chitosan into chitosan oligosaccharides follows certain kinetic laws. Generally speaking, under the same raw materials and reaction conditions, the higher the degree of degradation (i.e., the lower the average molecular weight of the product), the yield will vary within a certain range. If a process record shows a high yield but the average molecular weight does not decrease accordingly, or a low yield but the average molecular weight is extremely low, there may be errors in data recording or abnormal raw material parameters.
[0067] This embodiment pre-determines corresponding degradation kinetic threshold rules for different preparation routes. These threshold rules are determined based on publicly available chitosan degradation kinetics literature and generally accepted degradation patterns in the field, and are used to define reasonable ranges for the ratio of yield to average molecular weight under different routes. When the ratio of yield to average molecular weight in a process record exceeds the pre-determined threshold range for the corresponding route, the system automatically marks the entry as "suspected mechanism conflict." Subsequently, the raw material parameters corresponding to this entry are traced back, including the degree of deacetylation, initial molecular weight, and batch information of the raw chitosan. If the traceback reveals that the raw material parameters significantly deviate from typical values for similar processes, it is determined that the abnormal product indicators of this entry are mainly due to raw material characteristics rather than process advantages, and the entry is downweighted. If the raw material parameter information cannot be traced back, the entry is marked and listed as "low confidence data" in subsequent evaluations.
[0068] (ii) Consistency Restoration of Resource Consumption Indicators Based on the System Boundary Throughout the Entire Lifecycle For resource consumption indicators such as energy consumption per unit product and water consumption per unit product, different literature and experimental records often use different system boundaries in their calculations, which is an important source of data incomparability.
[0069] This embodiment predefines a unified system boundary for the entire lifecycle, dividing the chitosan oligosaccharide preparation process into three stages: raw material pretreatment, degradation reaction, and product separation and purification. This embodiment specifies that the resource consumption indicators in the database cover the entire process of these three stages by default, forming a unified measurement benchmark.
[0070] During the data integration process, for each record involving unit resource consumption indicators, the system boundary range declared in the original literature or experimental records is first identified and compared with the unified system boundary of this embodiment: if the original boundary is consistent with the unified boundary, the original value is directly adopted; if the original boundary only covers part of the process segment, it is supplemented by estimation based on the average resource consumption ratio of the missing process segment in similar processes; if the original boundary declaration is unclear, it is marked and used for sensitivity testing or as reference data in subsequent evaluations.
[0071] Through the conflict detection and consistency repair driven by the above two mechanisms, this embodiment ensures the mechanistic consistency and system boundary comparability of the data entering the database from the data source, laying a reliable foundation for subsequent standardization processing and database construction. Example 2: Normalization, Data Quality Control, and Confidence Calculation
[0072] This embodiment provides a standardized method for constructing process datasets and controlling data quality. Based on the valid process entries obtained in Embodiment 1, the data undergoes field extraction, terminology standardization, unit standardization, anomaly detection, missing value handling, classification coding, and relation mapping. Publicly available research is used to provide reference anchors for typical parameter fields in the database; for example, oxidation methods can record H2O2 concentration, room temperature / heating conditions, and two-step reaction structures; enzymatic methods can record optimal pH, optimal temperature, E / S ratio, and reaction time, etc.
[0073] In this embodiment, 16 core fields are preferably set, namely: raw material type, preparation route, pretreatment method, reaction temperature, reaction pH, reaction time, substrate concentration, liquid-solid ratio, enzyme addition amount, yield, percentage of target degree of polymerization range, average molecular weight, unit energy consumption, unit water consumption, waste liquid discharge, and unit product cost.
[0074] Among them, the 16 core fields can not only support database construction, but also form a mappable interface with the subsequent 16 comprehensive evaluation indicators.
[0075] Let q be the number of core fields that have been filled in the i-th valid process entry. i Then define the integrity coefficient C. i for:
[0076] And set the following thresholds: When C i An entry with a value ≥0.75 is considered a high integrity entry. When 0.625 ≤ C < 0.75, it is an entry that can be added to the database; When 0.50≤C i When <0.625, it is a reference storage entry only; When Ci When the value is less than 0.50, the item is removed.
[0077] After screening, among the 863 valid process entries in this embodiment, 412 entries were of high integrity, 311 entries were eligible for storage, and 140 entries were for reference only.
[0078] Therefore, the proportion of high integrity entries is:
[0079] The number of items in the formal evaluation set is:
[0080] The proportion of the formal evaluation set is:
[0081] In this embodiment, drawing on common data quality evaluation approaches used in process databases and LCA databases, five data quality dimensions are set: reliability R i Completeness C i Time representative T i Regional representativeness G i Technical Representative M i Each dimension is normalized to [0,1]. The overall confidence score Q is defined. i for:
[0082] And set the following hierarchical rules: When Q i A value ≥ 0.80 indicates high confidence level data. When 0.65≤Q i When the value is less than 0.80, the data is considered to have a medium confidence level. When 0.50≤Q i When the confidence level is less than 0.65, the data is considered low but should be retained. When Q i If the value is less than 0.50, it will not be included in the formal evaluation database.
[0083] In this embodiment, anomaly detection employs a dual mechanism of "range rule + source cross-validation". The following reasonable ranges for process parameters are set: reaction temperature: 20–95 ℃, reaction pH: 0.5–8.5, reaction time: 0.1–72 h, substrate concentration: 0.1–150 g / L, liquid-to-solid ratio: 2–100 mL / g, hydrogen peroxide concentration: 0.1%–10%, enzyme dosage: 1–200 U / g.
[0084] When a record exceeds the above range, it is automatically marked as suspected abnormal data and undergoes manual review. This range covers some mainstream process conditions in the reviewed papers, encompassing common experimental windows and enabling anomaly screening. For data confirmed as abnormal through manual review, if it is due to an entry error, it is corrected and retained; if it represents experimental conditions that deviate extremely from the mainstream range and lack a reasonable explanation, it is discarded.
[0085] In this embodiment, a database schema structure is further established. The database includes a Process_Main table, a Parameter_Table, an Indicator_Table, a Source_Table, and a Relation_Map table. The Process_Main table stores the process number, raw material type, and route type; the Parameter_Table stores parameters such as temperature, pH, time, substrate concentration, and E / S ratio; the Indicator_Table stores evaluation information such as yield, average molecular weight, unit energy consumption, unit water consumption, and unit cost; the Source_Table stores literature or experimental sources; and the Relation_Map table stores the relationships between process objects and parameters, indicators, and sources.
[0086] In this embodiment, a three-layer evaluation structure is constructed: the target layer is a comprehensive evaluation of the chitosan oligosaccharide preparation process; the criteria layer includes product quality, reaction efficiency, resource consumption, environmental impact, and economic efficiency; and the indicator layer contains 16 evaluation indicators. The number of structured records under the five criteria layers in the database are as follows: 1288 records for product quality, 722 for reaction efficiency, 841 for resource consumption, 549 for environmental impact, and 608 for economic efficiency. This distribution can be used to check the coverage balance of data across different dimensions during subsequent evaluations.
[0087] The method of pre-calculation and storage of sensitivity coefficients in this embodiment In this embodiment, during the database schema construction process, the sensitivity coefficients of each process parameter to the comprehensive evaluation result are pre-calculated and stored in the relational mapping table. The specific implementation method is as follows: First, historical data is extracted from the existing parameter and indicator tables to construct a parameter-indicator association dataset. For each process parameter (such as reaction temperature, reaction time, enzyme dosage, hydrogen peroxide concentration, etc.) and each evaluation indicator (such as yield, average molecular weight, unit energy consumption, etc.), the impact of parameter changes on indicator values is calculated, obtaining the sensitivity coefficient of each parameter to each indicator. The sensitivity coefficient is defined as the ratio of the indicator change rate to the parameter change rate; the larger the absolute value of the coefficient, the more significant the effect of the parameter on the indicator.
[0088] Furthermore, in order to obtain the overall sensitivity of the parameters to the comprehensive evaluation value, the sensitivity coefficients of each indicator can be weighted and aggregated according to the criterion layer weights to obtain the comprehensive sensitivity coefficient of each parameter.
[0089] After calculation, the comprehensive sensitivity coefficient of each parameter is stored as an additional field in a relational mapping table and associated with the corresponding process parameter number. When a user retrieves a certain type of process route through the database query interface, the system can synchronously return the sensitivity ranking of each process parameter under that route for the user's reference, eliminating the need for the user to export data and perform analysis and modeling themselves. Subsequently, as process data in the database continues to accumulate and update, the sensitivity coefficients can be recalculated and updated periodically to continuously improve their accuracy and representativeness. Example 3: Evaluation matrix construction, weight calculation, comprehensive score and sensitivity analysis
[0090] This embodiment provides a method for constructing an evaluation matrix and supporting evaluation based on a comprehensive evaluation database of multiple indicators for chitosan oligosaccharide preparation processes. Building upon the database established in Embodiment 2, this method extracts the set of evaluation objects and the set of indicators to generate a process evaluation matrix, which is then used for route-level comparison and key parameter analysis and verification.
[0091] Suppose the formal evaluation set contains m=723 process items, and the index set contains n=16 evaluation indicators, then the original evaluation matrix is defined as follows:
[0092] Right now:
[0093] Where, x ij This represents the original value of the i-th process item on the j-th evaluation index. For benefit-type indicators in the original evaluation matrix, minimum-maximum standardization is adopted:
[0094] For cost-related indicators, inverse standardization is used:
[0095] Thus, the standardized matrix is obtained:
[0096] In this embodiment, the entropy weight method is used to calculate the objective weight of each indicator. First, the weight of the i-th process under the j-th indicator is defined:
[0097] Next, calculate the entropy value of the j-th index:
[0098] The coefficient of difference is:
[0099] The weight of the j-th indicator is:
[0100] To facilitate route-level interpretation, this embodiment further consolidates the 16 indicators into 5 criterion layers. To avoid subjective setting of criterion layer weights, this embodiment further employs expert interviews to determine the weights of the five criterion layers. The five criterion layers are: product quality C1, reaction efficiency C2, resource consumption C3, environmental impact C4, and economic efficiency C5. 5。
[0101] In this embodiment, an expert advisory group of m=8 experts is preferably invited. These experts come from fields including chitosan oligosaccharide preparation, chemical process optimization, green manufacturing, environmental assessment, and industrialization cost accounting. Each expert scores the importance of the five criteria, ranging from 1 to 10, where 1 represents the lowest importance and 10 represents the highest importance. Let the score of the k-th expert for the i-th criterion be:
[0102] This results in an expert rating matrix:
[0103] The scoring results of the 8 experts are shown in the table below.
[0104]
[0105] The average of the expert scores for each criterion is calculated:
[0106] Substituting the above scoring data, we can obtain:
[0107] To eliminate differences in scoring scales among different experts, the average score was normalized to obtain the initial weights:
[0108] in:
[0109] Therefore, the initial weights are calculated as follows:
[0110]
[0111] To further consider the consistency of expert opinions, this embodiment calculates the standard deviation of each criterion score:
[0112] And calculate the coefficient of variation:
[0113] When the CV of a certain criterion i A large CV indicates a significant disagreement among experts regarding the importance of the criterion; when the CV is large... i A lower value indicates a more unified consensus among experts. To reduce the impact of high-disagreement criteria on the final weights, a consensus correction factor is introduced:
[0114] The corrected weights can then be expressed as:
[0115] In this embodiment, the differences in expert scores across the criteria are small, and the calculated coefficients of variation are all at low levels. Therefore, the weights after consistency correction are only slightly different from the initial weights. To facilitate engineering applications and route-level interpretation, the corrected weights are rounded to two decimal places, resulting in the final criterion-level weight vector:
[0116] verify:
[0117] The final weights were determined as follows: product quality 0.30, reaction efficiency 0.22, resource consumption 0.18, environmental impact 0.15, and economic efficiency 0.15. This weighting structure reflects the actual evaluation logic in the chitosan oligosaccharide preparation scenario, which prioritizes product results and efficiency while also considering greenness and cost.
[0118] Define the overall score for the i-th process item as:
[0119] The scores for each route category are averaged. Let n be the number of routes in category r that make it into the formal evaluation set. r (f) The average score at the route level is:
[0120] The number of entries for the five routes included in the formal evaluation set were set as follows: 196 for acid degradation, 149 for oxidative degradation, 182 for enzymatic hydrolysis, 72 for physical-assisted methods, and 124 for chemical-enzymatic coupling methods, totaling 723 entries.
[0121] This also allows us to calculate the percentage of each route in the formal evaluation set:
[0122] In this embodiment, the sensitivity coefficient of the key parameter is further defined:
[0123] When |K j The larger the | value, the more significant the impact of the j-th parameter on the overall score. Sensitivity analysis was performed on the following parameters: reaction temperature, reaction time, enzyme dosage, hydrogen peroxide concentration, and unit energy consumption. These parameters were chosen because publicly available studies typically optimize oxidation methods around H2O2 concentration and time, while enzymatic methods typically optimize around pH, temperature, E / S ratio, and time. These parameters have a strong impact on both product results and process efficiency and should be given priority in process optimization. Single-parameter perturbation analysis was further performed on the above key parameters.
[0124] This embodiment demonstrates that the multi-index comprehensive evaluation database for chitosan oligosaccharide preparation process constructed by this invention can not only complete the entire process from multi-source heterogeneous data acquisition, effective process item formation, standardization processing, pattern structure establishment to evaluation matrix output, but also directly support process comparison, comprehensive evaluation and key parameter analysis, verifying the feasibility of this invention in chitosan oligosaccharide preparation process research and subsequent process optimization.
[0125] Contents not described in detail in this specification are prior art known to those skilled in the art. Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for constructing a multi-index comprehensive evaluation database for chitosan oligosaccharide preparation process, characterized in that: The method includes: Step 1: Collect multi-source data related to the preparation process of chitosan oligosaccharide and construct the original process dataset. The multi-source data includes at least literature data, experimental record data and process data. Step 2: Extract fields, filter and integrate data from the original process dataset. During the integration process, based on the predefined chitosan oligosaccharide degradation kinetic rules and the unified full life cycle system boundary, perform process mechanism conflict detection and consistency repair on process data from different sources to form an effective set of process entries. Step 3: Standardize terminology, units, perform anomaly checks and missing value processing on the set of valid process entries to form a standardized process dataset; Step 4: Based on the comprehensive evaluation requirements of multiple indicators in the preparation process of chitosan oligosaccharide, the standardized process dataset is classified, coded, mapped, and hierarchically organized to construct a database schema structure; wherein, the relation mapping includes establishing the association between process entities, parameter entities, indicator entities, and source entities, and pre-calculating and storing sensitivity coefficients reflecting the degree of influence of process parameters on the comprehensive evaluation results in the association; Step 5: Based on the database schema structure, store the standardized process dataset as a dedicated database for comprehensive evaluation of multiple indicators of chitosan oligosaccharide preparation process; Step 6: Extract the evaluation object set and index set based on the dedicated database, generate a process evaluation matrix to support subsequent horizontal process comparison, unified quantitative evaluation and key parameter analysis, call the pre-stored sensitivity coefficients, and generate the process evaluation matrix and process optimization suggestions.
2. The method according to claim 1, wherein the method is characterized by: Step 1 includes: Step 11: Search publicly published literature and collect data on raw material sources, preparation routes, process parameters, and product performance in the preparation process of chitosan oligosaccharides; Step 12: Collect experimental data and extract reaction conditions, result indicators, and resource consumption data from the chitosan oligosaccharide preparation experiment; Step 13: Collect process data, including energy consumption, water consumption, reagent consumption, waste liquid discharge and economic data related to the preparation of chitosan oligosaccharides; Step 14: Perform deduplication, initial screening, and source verification on the collected data to construct the original process dataset.
3. The method according to claim 2, wherein the method is characterized by: In step 2, the field extraction includes extracting the fields of raw material type, pretreatment method, preparation route, reaction temperature, reaction pH, reaction time, substrate concentration, liquid-solid ratio, enzyme addition amount, yield, percentage of target degree of polymerization range, average molecular weight, energy consumption, water consumption, waste liquid volume and unit cost. The data filtering includes removing duplicate process entries, entries with missing key fields, and entries with unclear process boundaries; The process mechanism conflict detection and consistency repair specifically includes: when the ratio of yield to average molecular weight of the same process object in data from different sources exceeds the preset degradation kinetic threshold range, the raw material parameters are traced back and data alignment or weight reduction is performed; and for indicators involving unit resource consumption, the original data is traced back and corrected according to a unified full life cycle system boundary list, wherein the system boundary list defines whether it includes the energy and material consumption of raw material pretreatment, drying and pulverization processes.
4. The method according to claim 3, wherein the method is characterized by: Step 3 includes: Step 31: Map synonymous process terms from different data sources to unified standard terms; Step 32: Convert temperature, time, concentration, yield, molecular weight, energy consumption, and water consumption into a unified unit of measurement; Step 33: Identify anomalous data using rule-based validation, range validation, and source cross-validation; Step 34: Complete and mark missing data in non-critical fields, and remove or reduce the weight of missing data in critical fields to form a standardized process dataset.
5. The method according to claim 4, wherein the method is characterized by: The classification coding in step 4 includes establishing a unified coding system for raw material type, preparation route type, process parameter type, evaluation index type, and data source type; The hierarchical organization includes constructing a multi-index comprehensive evaluation hierarchical structure consisting of a target layer, a criterion layer, and an indicator layer; wherein, the target layer is a comprehensive evaluation of the chitosan oligosaccharide preparation process, and the criterion layer includes at least product quality, reaction efficiency, resource consumption, environmental impact, and economic efficiency.
6. The method according to claim 5, wherein the method is characterized by: The sensitivity coefficients pre-calculated in step 4 are obtained by: calculating the impact of changes in each process parameter on the comprehensive evaluation value based on historical data in the parameter table and index table, and storing the calculation results in the relational mapping for direct retrieval during database queries.
7. The method for constructing a multi-index comprehensive evaluation database for chitosan oligosaccharide preparation process according to claim 5, characterized in that: The dedicated database constructed in step 5 covers five types of chitosan oligosaccharide preparation routes: acid degradation, oxidative degradation, enzymatic hydrolysis, physical-assisted method, and chemical-enzymatic coupling method.
8. The method according to claim 5, wherein the method is characterized by: The indicator layer sets 16 evaluation indicators, which include: Yield, percentage of target degree of polymerization range, average molecular weight, purity, reaction time, substrate conversion, yield per unit time, energy consumption per unit product, water consumption per unit product, acid consumption per unit product, alkali consumption per unit product, enzyme consumption per unit product, waste liquid discharge, chemical oxygen demand, safety risk coefficient, and cost per unit product.
9. The method according to claim 7, wherein the method is characterized by: The dedicated database includes a process master table, a parameter table, an indicator table, a source table, and a relationship mapping table; The process master table stores unique identifiers for process items and process route information; the parameter table stores process parameter information; the index table stores multi-index evaluation data; the source table stores data source information; and the relationship mapping table stores the relationships between entities and the pre-calculated sensitivity coefficients.
10. The method according to claim 1, wherein the method is characterized by: The method for generating the process evaluation matrix in step 6 includes: Based on the dedicated database, the evaluation object set and indicator set are extracted to construct the original evaluation matrix; The original evaluation matrix is then normalized and standardized according to the index attributes to generate a standardized evaluation matrix that can be used for horizontal comparison and unified quantitative evaluation of the process, so as to be used for subsequent unified quantitative evaluation, horizontal comparison and key parameter analysis of the chitosan oligosaccharide preparation process.