Traditional Chinese medicine decoction piece production tracing method based on Internet of Things
Through the Internet of Things production traceability method, the Black Hawk optimization algorithm is used to construct a random forest network model, and outliers in the production data of Chinese herbal medicine are traced, which solves the problem that the existing technology cannot effectively trace the source of quality problems, and achieves the effect of rapidly improving production control processes and improving product quality.
Patent Information
- Application Number
- CN202510107905.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-06-06
AI Technical Summary
The existing Chinese herbal medicine production traceability methods cannot effectively trace the source of quality problems, which makes it difficult to quickly improve the production control process and increase the testing cost.
The Internet of Things-based Chinese herbal medicine production traceability method is adopted, and the medicinal material planting area is distinguished and marked, the Internet of Things platform is established to enter production data, generate QR codes, and a random forest network model is constructed using the Black Hawk optimization algorithm to trace outliers in the production data, and the source of outliers is analyzed to improve the production control process.
It can quickly locate the source of problems, and by adjusting and optimizing production processes in a targeted manner, timely discovering and solving problems in the production process, and improving production efficiency and product quality.
Smart Images

Figure CN120106644A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of Chinese medicine traceability, and in particular to a Chinese medicine decoction piece production traceability method based on the Internet of Things. Background Art
[0002] Chinese herbal medicine slices refer to Chinese medicinal materials that have been processed and prepared according to the theories of traditional Chinese medicine and Chinese medicine preparation methods and can be directly used in clinical Chinese medicine. There is no absolute boundary between Chinese medicinal materials and Chinese herbal medicine slices. Chinese herbal medicine slices include some Chinese herbal medicine slices processed at the place of origin, which are made into slices when formulated and prepared according to the theories of traditional Chinese medicine, original medicinal material slices, and slices that have been cut and processed.
[0003] Chinese herbal medicine slices are one of the three pillars of my country's traditional Chinese medicine industry. They are traditional weapons necessary for clinical diagnosis and treatment of traditional Chinese medicine, and are also important raw materials for Chinese patent medicines. With the continuous improvement and maturity of its processing theory, it has become an important means of clinical disease prevention and treatment in traditional Chinese medicine. At the same time, as an important form of clinical use of traditional Chinese medicine, the quality and safety of Chinese herbal medicine slices are directly related to the health and well-being of the people. With the continuous expansion of the traditional Chinese medicine market, the Chinese herbal medicine slice industry has exposed many problems, such as the unstable quality of Chinese herbal medicines due to the influence of factors such as the environment and the use of pesticides and fertilizers in the planting of Chinese herbal medicines, the lack of unified and strict standards and effective supervision in the processing and processing process, and the frequent appearance of counterfeit and inferior products in the circulation link. These problems have seriously restricted the healthy development of the traditional Chinese medicine industry and also caused consumers to worry about the quality and safety of Chinese herbal medicine slices. Therefore, it is urgent to establish a complete set of traceability methods for the production of Chinese herbal medicine slices.
[0004] Moreover, the existing production traceability method for Chinese herbal medicines can track the production process of each product by testing and recording the environmental data and quality data of Juncao and setting labels for the Chinese herbal medicines, but it cannot trace the source of the quality problems that have occurred, and thus cannot quickly and effectively improve the production control process, which greatly increases the testing cost.
[0005] Therefore, how to trace the production quality of Chinese herbal medicines is a technical problem that technicians need to solve at present. Summary of the invention
[0006] Based on the above purpose, the present invention provides a method for tracing the production of Chinese herbal medicine slices based on the Internet of Things
[0007] A method for tracing the production of Chinese herbal medicine pieces based on the Internet of Things, comprising:
[0008] S1: Differentiate and mark the medicinal material planting areas of Chinese herbal medicine pieces to obtain marking information.
[0009] S2: Establish an Internet of Things platform, enter the marking information and the data of each production link into the Internet of Things platform to form production data, generate a corresponding QR code, and paste the QR code on the packaging of the corresponding Chinese herbal medicine product. The production link data includes processing time, processing personnel, and equipment status.
[0010] S3: Acquire sources based on historical production data and corresponding to abnormal values in the historical production data as a data set, normalize the data set, and divide the data values into a training set and a test set according to a preset ratio.
[0011] S4: Construct a random forest network model based on the Black Hawk optimization algorithm, use the training data to train the network model, obtain the optimal parameter group, and test the optimal parameter group through the test set. If the test result is correct, the optimal network model is obtained. Otherwise, the network model is repeatedly trained until the test result is correct. Use the optimal network model to trace the abnormal values in the production data in the Internet of Things platform to obtain the source of the abnormal values.
[0012] S5: Analyze the sources and improve production control processes.
[0013] Preferably, in step S1, the specific steps include:
[0014] S11: Collect and record the planting information of medicinal materials, wherein the planting information includes geographical location, climate conditions and soil type.
[0015] S12: Performing a quality check on the medicinal materials, removing medicinal materials with damaged appearance and peculiar smell, marking the remaining medicinal materials according to corresponding planting information, and obtaining marking information.
[0016] Preferably, in step S3, the specific steps of obtaining the optimal prediction model include:
[0017] S31: normalizing the data set using maximum value normalization.
[0018] S32: The data is divided in chronological order, with the first 80% of the data used as the training set for training the random forest model, and the last 20% of the data used as the test set.
[0019] Preferably, in step S4, the specific steps include:
[0020] S41: Build a random forest prediction model.
[0021] S42: Optimizing the forest prediction model using the Black Hawk optimization algorithm and the training set to obtain an optimal parameter set.
[0022] Preferably, in step S41, the specific steps include:
[0023] S411: setting the number and depth of decision trees according to the characteristics of the production link data.
[0024] S412: Set the proportion of features randomly selected by each decision tree during the training process.
[0025] S413: Setting the minimum number of samples required for splitting a decision tree node and the maximum number of leaf nodes of the decision tree.
[0026] According to a method for tracing the production of Chinese herbal medicine pieces based on the Internet of Things provided by the present invention, in step S42, the specific steps include:
[0027] S421: Use the F1 score as the fitness function of the network model.
[0028] S422: Initialize the population, randomly generate the positions of individuals in the black hawk population, and each individual represents a parameter combination of the random forest network model.
[0029] S423: The individuals of the initialized population are close to the position of the current best individual to obtain population one.
[0030] S424: The individuals in the population one find the global optimal position in the population one through snatching behavior.
[0031] S425: Use a random sampling mechanism to extract some individuals from population one to move closer to the global optimal position, and obtain population two.
[0032] S426: Calculate the fitness value of population 2. If the output fitness is higher than the threshold, output the optimal parameter set. Otherwise, return to step S422 until the output fitness value is higher than the preset fitness threshold or the model training reaches the maximum number of iterations to obtain the optimal parameter set.
[0033] According to a method for tracing the production of Chinese herbal medicine pieces based on the Internet of Things provided by the present invention, in step S424, the snatching behavior expression formula is:
[0034]
[0035] in, is the position after the jump, is the current position, δ is the randomly generated jump distance coefficient, It is the position of the best individual in the contemporary era.
[0036] Preferably, in step S423, the specific steps include:
[0037] S4231: The individuals in the initialized population use hovering behavior to find the individual local optimal position, and then obtain the optimal individual position through comparison.
[0038] S4232: Adjusting the positions of individuals in the initialized population through capturing behavior.
[0039] S4233: The individuals in the initialized population approach the optimal individual position through round-up behavior to obtain population one.
[0040] Preferably, in step S4233, the encirclement behavior expression formula is:
[0041]
[0042] in, is the position after iteration, is a random position, is the current best position, α and r 1 is a random factor, t 1 It is a random number generated by the Tent mapping.
[0043] Preferably, in step S5, the specific steps include:
[0044] S51: Use a scatter plot to visualize the outliers and their sources.
[0045] S52: Using the RCA analysis method, analyze the potential causes of the abnormal value sources layer by layer, and improve the production control process based on the potential causes.
[0046] Beneficial effects of the present invention:
[0047] By building a random forest network model based on the Black Hawk optimization algorithm, and using historical production data and data sets constructed from the sources of outliers in historical production data to train the network model, the optimal network model is obtained. The outliers in the production data in the networking platform are traced using the optimal network model, which can quickly locate the source of the problem. By analyzing the source of outliers, the production process can be adjusted and optimized in a targeted manner, problems in the production process can be discovered and solved in a timely manner, and production efficiency and product quality can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are only for the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0049] Figure 1This is one of the flow charts of the method for tracing the production of Chinese herbal medicine pieces based on the Internet of Things provided by an embodiment of the present invention;
[0050] Figure 2 This is a second flow chart of a method for tracing the production of Chinese herbal medicine pieces based on the Internet of Things provided by an embodiment of the present invention;
[0051] Figure 3 This is the third flow chart of the method for tracing the production of Chinese herbal medicine pieces based on the Internet of Things provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0052] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. At the same time, it is explained here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments, and those skilled in the art may also adopt other alternatives to implement some known technologies; and the accompanying drawings are only for more specific description of the embodiments, and are not intended to specifically limit the present invention.
[0053] Example 1, please refer to Figure 1-Figure 3 This embodiment provides a method for tracing the production of Chinese herbal medicine pieces based on the Internet of Things, including:
[0054] S1: Differentiate and mark the medicinal material planting areas of Chinese herbal medicine pieces to obtain marking information.
[0055] S11: Collect and record the planting information of medicinal materials, including geographical location, climate conditions and soil type.
[0056] In this embodiment, in addition to the geographical location, climate conditions and soil type, detailed information such as the name of the grower or company name, planting techniques such as organic planting, traditional planting, etc., seed sources, irrigation methods, types of fertilizers and frequencies, etc. can also be collected and recorded. This information will help to more fully understand the growth environment and planting process of medicinal materials, and provide a basis for subsequent quality control and traceability. Using the Internet of Things sensor network and GIS technology, the environmental parameters of the planting area, such as temperature, humidity, light, etc., as well as soil nutrient content and pH, can be monitored and recorded in real time to ensure the accuracy and real-time nature of the data.
[0057] S12: Perform quality inspection on the medicinal materials, remove the medicinal materials with damaged appearance and peculiar smell, mark the remaining medicinal materials according to the corresponding planting information, and obtain marking information.
[0058] In terms of odor inspection, this embodiment can use the sense of smell to determine whether the smell of the medicinal material is normal, whether there is any peculiar smell or signs of deterioration. In terms of taste inspection, it can be checked by tasting whether there is an abnormal taste. In terms of appearance inspection, the shape, color, size, surface texture, etc. of the medicinal material can be used to ensure that the appearance of the medicinal material meets the specified standards. For qualified medicinal materials, unique marking information, such as a QR code or a barcode, is generated based on the planting information. The marking information should contain key information such as the planting area, planting batch, and harvest time of the medicinal material. Modern analysis technologies such as hyperspectral imaging and near-infrared spectroscopy can also be used to quickly and accurately detect the quality indicators of medicinal materials. At the same time, wireless radio frequency identification or near-field communication technology embeds the marking information into the medicinal material packaging or label to facilitate subsequent traceability.
[0059] S2: Establish an IoT platform, enter the marking information and data of each production link into the IoT platform to form production data, generate a corresponding QR code, and affix the QR code to the packaging of the corresponding Chinese herbal medicine products. Production link data includes processing time, processing personnel, and equipment status.
[0060] In this embodiment, intelligent warehousing management can also be established through the Internet of Things cloud platform. According to the real-time monitoring data, the platform can automatically adjust the storage conditions to maintain the optimal storage state of Chinese herbal medicines. When the storage environment is abnormal, the platform can issue an early warning in time so that managers can take corresponding measures. This method can improve the storage quality and safety of Chinese herbal medicines and reduce storage costs. Anti-counterfeiting functions can also be added through the Internet of Things cloud platform. Not only can product information be understood through the QR code, but the authenticity of the product can also be identified. In this embodiment, a risk matrix is used to assess the risk of the product. This includes determining risk levels, such as high, medium, low and corresponding warning thresholds. Risk assessment should not be just a one-time process, but should be continuously updated as the product life cycle progresses and new quality data is collected. Machine learning algorithms can be used to learn historical data to predict possible risk trends in the future, and the warning thresholds can be adjusted accordingly. On the basis of the three levels of high, medium and low, the risk levels can be further subdivided, such as extremely high risk, high risk and low high risk in high risk, so as to formulate early warning strategies more accurately. Warning information is published on the Internet of Things platform, such as red, yellow, green and other warning lights or text prompts. At the same time, the early warning information can be sent to the mobile phones or mailboxes of relevant personnel in real time through SMS, email, etc. This includes suspending production, strengthening inspection, etc. An emergency response mechanism can also be established to ensure that countermeasures can be taken quickly after the early warning prompt is issued. This includes setting up an emergency team, formulating emergency plans, preparing emergency supplies, etc. The message push function or API interface of the Internet of Things platform can also be used to realize the real-time transmission of early warning information. Project management tools or custom process management systems are also used to formulate and handle the implementation of strategies. Early warning information can also be visualized, and intuitive early warning information visualization tools such as trend charts and bar charts can be provided to enable relevant personnel to quickly understand the risk status. Allow relevant personnel to view detailed risk assessment reports and handling suggestions by clicking or hovering on the early warning information. Early warning levels can be set to push early warning information to managers and technicians at different levels. High-risk early warnings should be pushed to senior managers and key technical personnel first so that decisions can be made and actions can be taken quickly.
[0061] S3: Obtain sources based on historical production data and corresponding to abnormal values in the historical production data as a data set, normalize the data set, and divide the data values into a training set and a test set according to a preset ratio.
[0062] In this embodiment, these production data can be received from various production links, quality inspection departments, and sales feedback. In the process of integrating these data and collecting data, attention should be paid to ensuring the integrity and consistency of the data. By verifying the integrity and consistency of the data, we can identify and fix any possible data omissions or inconsistencies, laying a solid foundation for subsequent data processing and analysis. In order to improve data quality, duplicate data records and invalid or redundant data can be checked and removed. This helps to reduce noise in the data processing process and improve the accuracy and efficiency of the model. For missing values in the data, they can be filled or deleted according to the distribution of the data and business logic. For example, if the missing values are caused by random errors in the data recording process, we may choose to delete these records; if the missing values are caused by certain specific conditions, such as equipment failure, they can be filled according to the context. For outliers in the data, detailed inspection and analysis should be carried out. Statistical methods and business knowledge can be used to determine whether these outliers are caused by real abnormal events, such as sudden failures in the production process or data input errors. For real outliers, we will retain them as part of the model training; for outliers caused by data input errors, they can be corrected or deleted. In order to eliminate the dimensional differences between different features and improve the training efficiency and accuracy of the model, tools such as correlation matrix can also be used to analyze the correlation between features. This helps to identify highly correlated or redundant features for feature selection or dimensionality reduction. Mutual information is a metric used to evaluate the dependency between two variables. Mutual information can also be used to evaluate the correlation between features and target variables.
[0063] S31: normalize the data set using maximum value normalization.
[0064] In this embodiment, the maximum value normalization expression formula is:
[0065]
[0066] Among them, X is the normalized data, X 1 is the original data, X a is the maximum value of the data, X i is the minimum value of the data.
[0067] It can also be processed by decimal scaling normalization, which scales the data by moving the decimal point position of the data. The specific method is to find the maximum value in the data and determine the position of its decimal point. Then, move the decimal point of all the data to the same position. In some cases, you may need to adjust the format of the data after moving the decimal point to ensure that they meet the input requirements of the model. For example, if the model requires the input data to be floating point numbers, you need to ensure that the data after decimal scaling normalization is also in floating point format. The normalized data can be verified to ensure that they meet the input requirements of the model and that no new data quality issues are introduced.
[0068] S32: Use time order to divide the data, use the first 80% of the data as the training set to train the random forest model, and use the last 20% of the data as the test set.
[0069] In this embodiment, 70% of the data can be used as a training set, 20% of the data can be used as a validation set to adjust model parameters, and 10% of the data can be used as a test set. This method can ensure the generalization ability of the model on different data sets. A cross-validation method can also be used. For example, the data set is divided into k parts, and k-1 parts are selected as training sets each time, and the remaining part is used as a test set. Repeat this process k times, and select a different part as the test set each time. Finally, the average performance of the k experiments is calculated as the final performance of the model.
[0070] S4: Construct a random forest network model based on the Black Hawk optimization algorithm, use the training data to train the network model, obtain the optimal parameter combination, and test the optimal parameter combination through the test set. If the test result is correct, the optimal network model is obtained. Otherwise, the network model is repeatedly trained until the test result is correct. Use the optimal network model to trace the outliers in the production data in the IoT platform and obtain the source of the outliers. It should be noted that the Black Eagle Optimizer (BEO) is a new type of meta-heuristic algorithm, which combines the biological behavior of the black hawk and mathematical transformation to guide the search behavior of particles; the principle of the Black Eagle Optimizer is based on the various biological behaviors of the black hawk, including hunting, hovering, capturing, snatching, warning, migration, courtship and hatching, etc. These behaviors are mathematized in the algorithm to guide the search process; specifically, the algorithm achieves optimization through the following steps: population initialization: initialize the population by randomly generating search agents; hunting behavior: simulate the process of the black hawk looking for prey, and improve the search density and global optimization ability by adapting the distribution group position and increasing the population size; hovering behavior: refine the range of the global optimal position through rotation search; capturing, snatching, warning, migration, courtship and hatching behaviors: these behaviors are implemented in the algorithm through different mathematical transformations to simulate different survival strategies of the black hawk.
[0071] In this example, biological laws and mathematical transformations are combined: the search process is guided by simulating the biological behavior of black hawks. High-density search method: the search density is improved by adaptively distributing the group positions and increasing the population size, enhancing the global optimization capability; rotation search: the rotation search is performed through hovering operations to further refine the range of the global optimal position.
[0072] S41: Build a random forest prediction model.
[0073] In this embodiment, a comprehensive and in-depth exploratory analysis can be performed on the input data. This step is crucial for understanding the overall structure of the data and revealing its potential laws. By examining the distribution of the data in detail, it is possible to gain insight into the distribution of each eigenvalue in the overall data set, which provides an important reference for subsequent data processing and model building. At the same time, special attention can be paid to the correlation between features, because the interaction between features often contains rich information. Using visualization tools such as correlation matrices or heat maps, the degree of association and change trends between features can be intuitively seen. These tools not only help identify strongly correlated feature pairs, but also provide strong support for subsequent feature selection. Based on the results of the exploratory analysis, a series of feature engineering operations can be performed according to the specific needs of the problem and the characteristics of the data itself. This includes but is not limited to feature selection, feature extraction, and feature scaling. Feature selection aims to select features that have a significant impact on the prediction target from the original feature set to improve the accuracy and efficiency of the model. Feature extraction is to construct new and more explanatory features by converting or combining original features, thereby further enriching the information input of the model. Feature scaling is to ensure that all features are comparable in value, to avoid adverse effects on model training due to large differences in value ranges. In the process of feature engineering, you can also make full use of the information provided by tools such as correlation matrices or heat maps to assist in feature selection and optimization. These tools not only help identify features that have a significant impact on the prediction target, but also provide useful guidance for feature extraction and scaling. Through this series of feature engineering operations, a more refined and effective feature set can be constructed, laying a solid foundation for subsequent model training and prediction.
[0074] S411: Set the number and depth of decision trees according to the characteristics of production process data.
[0075] In this embodiment, according to the characteristics of the data and the complexity of the problem, a reasonable number and depth of decision trees are set. Generally speaking, the more decision trees there are, the more stable the performance of the model and the more accurate the prediction results. However, too many decision trees will increase the computing cost and time. Therefore, it is necessary to set it reasonably according to the characteristics of the data and the computing resources. In practical applications, the number of decision trees is usually set to more than 100 to achieve better performance and error rate. The number of decision trees is also affected by factors such as data scale and number of features. For large-scale data sets and complex problems, more decision trees may be needed to improve the performance of the model. The depth of the decision tree determines the complexity of the tree and the fitting ability of the model. Shallow trees may cause underfitting of the model, while too deep trees may cause overfitting. Therefore, it is necessary to set the depth of the decision tree reasonably. In practical applications, the optimal depth can be determined by methods such as cross-validation. The depth of the decision tree is affected by factors such as data characteristics and sample size. For data sets with complex feature relationships and a large number of samples, deeper decision trees may be needed to capture the subtle differences in the data.
[0076] S412: Set the proportion of features randomly selected by each decision tree during the training process.
[0077] In this embodiment, a reasonable proportion of randomly selected features can be set according to the characteristics of the data and the complexity of the problem. When training each decision tree, the random forest model will randomly select a part of the features for splitting. This randomness helps to increase the diversity and robustness of the model. The proportion of randomly selected features is usually set to the square root or logarithm of the total number of features. This setting can reduce the computational cost and time while ensuring the performance of the model. The proportion of randomly selected features is also affected by factors such as the importance and relevance of the data features. For data sets with a large number of redundant features, the proportion of randomly selected features can be appropriately increased to reduce the impact of redundant features on model performance. For classification problems, if the number of categories is large or the boundaries between categories are not obvious, the proportion of randomly selected features can be appropriately increased to increase the model's ability to distinguish different categories. For regression problems, if the target variable has a large range of variation or there is a nonlinear relationship, the proportion of randomly selected features also needs to be adjusted accordingly to ensure that the model can capture the complex patterns in the data. According to the importance score of the feature, the proportion of randomly selected features of each decision tree during the training process can be dynamically adjusted. Features with higher importance scores are more likely to be selected. When setting the proportion of randomly selected features, you can make judgments and adjustments based on domain knowledge and experience to ensure that the model can better adapt to the needs of specific application scenarios.
[0078] S413: Setting the minimum number of samples required for splitting a decision tree node and the maximum number of leaf nodes of the decision tree.
[0079] In this embodiment, a reasonable minimum number of samples and a maximum number of leaf nodes of the decision tree can be set according to the characteristics of the data and the complexity of the problem. A smaller minimum number of samples may cause the tree to grow too fast and easily overfit. A larger minimum number of samples may cause the tree to grow too slowly and easily underfit. Therefore, it is necessary to set it reasonably according to the characteristics of the data and the complexity of the problem. The minimum number of samples is also affected by factors such as the data scale and the number of features. For large-scale data sets and complex problems, it may be necessary to set a larger minimum number of samples to avoid overfitting. The maximum number of leaf nodes of the decision tree limits the growth range of the tree and helps to prevent overfitting. A larger maximum number of leaf nodes may cause the tree to grow excessively and increase the complexity of the model. A smaller maximum number of leaf nodes may limit the growth of the tree and reduce the performance of the model. Therefore, it is necessary to set it reasonably according to the characteristics of the data and the complexity of the problem. The maximum number of leaf nodes is also affected by factors such as the data scale and the number of features. For data sets with complex feature relationships and a large number of samples, it may be necessary to set a larger maximum number of leaf nodes to capture the subtle differences in the data.
[0080] S42: Use the Black Hawk optimization algorithm and training set to optimize the forest prediction model and obtain the optimal parameter set.
[0081] S421: Use the F1 score as the fitness function of the network model.
[0082] In this embodiment, the F1 score is an indicator used to measure the accuracy of the binary classification model. It takes into account both the precision and recall of the classification model and is the harmonic mean of the precision and recall. The maximum value of the F1 score is 1, indicating that the model is perfect in both precision and recall. The minimum value is 0, indicating that the model performs poorly in both indicators.
[0083] The specific expression of F1 score is:
[0084]
[0085] Among them, F1 is the evaluation index, P is the precision, R is the recall, TP is the number of samples in the positive class that are predicted as positive, FP is the number of samples in the negative class that are predicted as positive, and FN is the number of samples in the positive class that are predicted as negative.
[0086] S422: Initialize the population, randomly generate the positions of individuals in the black hawk population, and each individual represents a parameter combination of the random forest network model.
[0087] In this embodiment, when initializing the population, it is necessary to determine the size of the population, that is, the number of initial solutions. The larger the population size, the wider the search space, but the higher the computational cost. In practical applications, the population size can be reasonably set according to computing resources and time constraints. When randomly generating the initial solution, it is necessary to set a reasonable range for each parameter of the random forest. These ranges can be set based on experience, data characteristics or suggestions in the literature. In order to increase the diversity of the population, the generation of the initial solution should have a certain degree of randomness. This can be achieved by randomly selecting values within the set parameter range. Before building the random forest prediction model, feature selection and preprocessing operations can also be performed. This includes removing redundant features, processing missing values, standardizing or normalizing features, etc. These operations help to improve the performance and stability of the model.
[0088] S423: Initialize the individuals of the population close to the position of the current best individual to obtain population 1. The specific steps include:
[0089] S4231: Initialize the individuals in the population and use the hovering behavior to find the individual's local optimal position, and then obtain the optimal individual position through comparison.
[0090] In this embodiment, each individual in the initialization population first performs a hovering behavior, which simulates the habit of the black hawk to pause and carefully observe the surrounding environment when exploring the environment. Mathematically, it can also be achieved by allowing the individual to perform a small-scale random search near its current position to find the local optimal position. The scope and step length of the search can be adjusted according to the specific problem. By comparing the fitness value after each individual hovering behavior, the local optimal individual position in the current population can be determined. This optimal position will serve as a reference point for individual movement and snatching behavior in subsequent steps.
[0091] S4232: adjusting the positions of individuals in the initialized population through capturing behavior;
[0092] S4233: The individuals in the initialized population approach the optimal individual position through the round-up behavior to obtain population 1, wherein the round-up behavior expression formula is:
[0093]
[0094] in, is the position after iteration, is a random position, is the current best position, α and r 1 is a random factor, t 1 It is a random number generated by Tent mapping; Tent mapping, also known as tent map, refers to a piecewise linear mapping.
[0095] S424: Individuals in population one find the global optimal position in population one through snatching behavior.
[0096] In this embodiment, in step S424, the specific steps include: the snatching behavior expression formula is:
[0097]
[0098] in, is the position after the jump, is the current position, δ is the randomly generated jump distance coefficient, It is the position of the best individual in the contemporary era.
[0099] S425: Use a random sampling mechanism to extract some individuals from population one to move closer to the global optimal position, and obtain population two.
[0100] In this embodiment, some individuals are made close to the current optimal position, while other individuals spread outward to prevent falling into the local optimum. In the random sampling process, we can introduce the idea of stratified sampling and stratify according to the fitness value of the individual or its relative distance from the global optimal position. In this way, it can ensure that a part of the individuals with high fitness or close to the optimal position are selected for fine search, and a part of the individuals far from the optimal position can be retained for extensive exploration. This differentiation strategy helps to maintain the diversity and exploration ability of the population while accelerating convergence. The distance of the selected individuals to the global optimal position can be designed as a certain adaptive function based on the distance between the current position of the individual and the global optimal position. For example, when the individual is far away from the global optimal position, a larger step size can be taken for rapid approximation; and when the individual is close to the global optimal position, the step size is reduced for fine search. This adaptive adjustment strategy helps to improve search efficiency and accuracy. In the process of approaching the global optimal position, in order to increase the diversity of the search and avoid falling into the local optimal solution, we can introduce certain random perturbations to the movement of the individual. This random perturbation can be manifested as a small random change in the position vector, or a random adjustment in the step size and direction. By introducing random perturbations, we can increase the flexibility and exploration ability of the search while maintaining the directionality of the search. In the process of random sampling and individual movement, special attention should be paid to retaining those individuals that are at the edge of the population or far away from the current optimal position. Although these individuals may have low current fitness, they contain rich search information and potential exploration directions. By retaining these marginal individuals, new search directions and possibilities can be continuously introduced during the search process, thereby enhancing the diversity of the search and avoiding premature convergence. In addition to guiding some individuals to approach the global optimal position, we can also design diverse exploration strategies for other individuals. We can also use highly exploratory movement methods such as random walks and Levy flights to allow these individuals to conduct extensive exploration in the search space. By combining the two strategies of fine search and extensive exploration, we can cover the search space more comprehensively and increase the possibility of finding the global optimal solution. At the same time, we can also use the roulette method to select which individual to approach the optimal position by comparing the fitness value of each individual.
[0101] S426: Calculate the fitness value of population 2. If the output fitness is higher than the threshold, output the optimal parameter set. Otherwise, return to step S422 until the output fitness value is higher than the preset fitness threshold or the model training reaches the maximum number of iterations to obtain the optimal parameter set.
[0102] In this embodiment, in order to speed up the iterative training process, parallel computing technology can also be used. This includes the use of hardware resources such as multi-core processors and distributed computing clusters. Modern computers are usually equipped with multi-core processors, and each core can perform computing tasks independently. When building a random forest model, the training of each decision tree is independent, so it can be executed on multiple cores in parallel. The distributed computing cluster consists of multiple computers, which are connected through a network and perform tasks together. The training task of the random forest can be split into multiple subtasks and executed in parallel on different nodes of the distributed computing cluster. Although parallel computing may require more hardware resources, the overall computing cost can be reduced by improving computing efficiency and optimizing resource utilization. With the development of cloud computing and virtualization technology, computing resources can be dynamically allocated on demand to further reduce computing costs.
[0103] S5: Analyze sources and improve production control processes.
[0104] S51: Use scatter plots to visualize outliers and their sources.
[0105] In this embodiment, various abnormal values may be encountered during the production process of Chinese herbal medicine slices, such as abnormal processing time, processing personnel operating errors, abnormal equipment status, etc. In order to intuitively display these abnormal values and their sources, a scatter plot can be used for visualization. Assume that there is a data set that contains data from multiple production links in the production process of Chinese herbal medicine slices, as well as abnormal values corresponding to each production link. These data can also be imported into the data analysis software and visualized using the scatter plot function. In the scatter plot, we can represent the data points of each production link with different colors or shapes to distinguish different production links. At the same time, the abnormal values can be represented by special marks to highlight them. In addition, labels or annotations can be added to the scatter plot to illustrate the source of each data point or abnormal value. For example, a note can be added next to the abnormal value to indicate whether it is caused by too long processing time, processing personnel operating errors, or abnormal equipment status. Through the visualization of the scatter plot, we can intuitively see the abnormal values and their sources in the production process of Chinese herbal medicine slices, so that it is easier to find problems and take corresponding measures to improve them.
[0106] S52: Use the RCA analysis method to analyze the potential causes of abnormal values layer by layer, and improve the production control process based on the potential causes.
[0107] In this embodiment, a random forest network model based on the Black Hawk optimization algorithm is constructed, and the network model is trained using a data set constructed from historical production data and the source of outliers therein to obtain an optimal network model. The optimal network model is used to trace outliers in the production data in the networking platform, and the source of the problem can be quickly located. By analyzing the source of outliers, the production process can be adjusted and optimized in a targeted manner, problems in the production process can be discovered and solved in a timely manner, and production efficiency and product quality can be improved.
[0108] Among them, root cause analysis (RCA) is a structured problem-solving method, which is used to gradually find the root cause of the problem and solve it, rather than just focusing on the symptoms of the problem; among them, root cause analysis is a systematic problem-solving process, including determining and analyzing the cause of the problem, finding solutions to the problem, and formulating preventive measures for the problem; in the field of organizational management, root cause analysis can help stakeholders discover the crux of organizational problems and find fundamental solutions.
[0109] The present invention covers any substitution, modification, equivalent method and scheme made on the essence and scope of the present invention. In order to make the public have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention, but those skilled in the art can fully understand the present invention without the description of these details. In addition, in order to avoid unnecessary confusion about the essence of the present invention, well-known methods, processes, procedures, components and circuits are not described in detail.
[0110] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A method for tracing the production of Chinese herbal medicine pieces based on the Internet of Things, characterized in that: The method comprises: S1: Differentiate and mark the medicinal material planting areas of Chinese herbal medicine pieces to obtain marking information; S2: Establish an Internet of Things platform, enter the marking information and the data of each production link into the Internet of Things platform to form production data, and generate a corresponding QR code; the production link data includes processing time, processing personnel, and equipment status; S3: obtaining sources based on historical production data and corresponding to abnormal values in the historical production data as a data set, normalizing the data set and dividing the data set into a training set and a test set according to a preset ratio; S4: constructing a random forest network model based on the Black Hawk optimization algorithm, inputting the training set into the random forest network model for training, obtaining an optimal parameter group, and using the test set to test the optimal parameter group; if the test result is correct, the optimal network model is obtained, otherwise, the network model is repeatedly trained until the test result is correct; S5: Using the optimal network model to trace the abnormal values in the production data in the Internet of Things platform, obtain the source of the abnormal values, analyze the source, and improve the production control process.
2. The method for tracing the production of Chinese herbal medicine slices based on the Internet of Things according to claim 1, characterized in that: In step S1, after the medicinal material planting areas of the Chinese herbal medicine pieces are respectively distinguished and marked, and before the marking information is obtained, the following steps are further included: Obtaining the planting information of the medicinal materials to be recorded, and performing quality inspection on the medicinal materials to be recorded, removing medicinal materials with damaged appearance and odor, wherein the planting information includes geographical location, climate conditions and soil type; The remaining medicinal materials are marked according to the corresponding planting information to obtain marking information.
3. The method for tracing the production of Chinese herbal medicine slices based on the Internet of Things according to claim 1, characterized in that: In step S3, the data set is normalized and divided into a training set and a test set according to a preset ratio, including: The data set was normalized by using maximum normalization and divided by time sequence. The first 80% of the data was used as a training set for training the random forest model; and the last 20% of the data was used as a test set.
4. The method for tracing the production of Chinese herbal medicine slices based on the Internet of Things according to claim 1, characterized in that: In step S4, the construction of a random forest network model based on the Black Hawk optimization algorithm includes: S41: Establish a random forest prediction model; S42: Optimizing the forest prediction model using the Black Hawk optimization algorithm and the training set to obtain an optimal parameter set.
5. The method for tracing the production of Chinese herbal medicine slices based on the Internet of Things according to claim 4, characterized in that: In step S41, the specific steps include: S411: setting the number and depth of decision trees according to the characteristics of the production link data; S412: Set the proportion of randomly selected features for each decision tree during the training process; S413: Setting the minimum number of samples required for splitting a decision tree node and the maximum number of leaf nodes of the decision tree.
6. The method for tracing the production of Chinese herbal medicine slices based on the Internet of Things according to claim 4, characterized in that: In step S42, the specific steps include: S421: Using the F1 score as the fitness function of the network model; S422: Initialize the population, randomly generate the positions of individuals in the black hawk population, and each individual represents the parameter combination of the random forest network model; S423: moving the individuals of the initialized population close to the position of the current best individual to obtain population 1; S424: The individuals in the population one find the global optimal position in the population one through snatching behavior; S425: Use a random sampling mechanism to extract some individuals from population one to move closer to the global optimal position, and obtain population two; S426: Calculate the fitness value of population 2. If the output fitness is higher than the threshold, output the optimal parameter set. Otherwise, return to step S422 until the output fitness value is higher than the preset fitness threshold or the model training reaches the maximum number of iterations to obtain the optimal parameter set.
7. The method for tracing the production of Chinese herbal medicine slices based on the Internet of Things according to claim 6, characterized in that: In step S424, the snatching behavior expression formula is: in, is the position after the jump, is the current position, δ is the randomly generated jump distance coefficient, It is the position of the best individual in the contemporary era.
8. The method for tracing the production of Chinese herbal medicine slices based on the Internet of Things according to claim 6, characterized in that: In step S423, the specific steps include: S4231: The individuals in the initialized population use hovering behavior to find the best local position of the individuals, and obtain the best individual position through comparison; S4232: adjusting the positions of individuals in the initialized population through capturing behavior; S4233: The individuals in the initialized population approach the optimal individual position through round-up behavior to obtain population one.
9. The method for tracing the production of Chinese herbal medicine slices based on the Internet of Things according to claim 7, characterized in that: In step S4233, the encirclement behavior expression formula is: in, is the position after iteration, is a random position, is the current best position, α and r1 are random factors, and t1 is a random number generated by Tent mapping.
10. The method for tracing the production of Chinese herbal medicine slices based on the Internet of Things according to claim 1, characterized in that: In step S5, the specific steps include: S51: Visualizing the outliers and their sources using a scatter plot; S52: Using the RCA analysis method, analyze the potential causes of the abnormal value sources layer by layer, and improve the production control process based on the potential causes.
Citation Information
Patent Citations
Traditional Chinese medicine decoction piece production data acquisition system for quality tracing
CN113657907A
Behavior baseline targeted capturing method based on intrusion data traceability
CN115118505A
Traditional Chinese medicinal material quality information tracing method and system
CN117636081A
Distributed driving vehicle trajectory tracking control method considering energy consumption factor
CN119078883A
Medium and long term hydrological probability forecasting method and system considering interpretable deep learning model nested combination
CN119204350A
Cited By
Traditional Chinese medicine beverage raw material traceability analysis method and system based on neural network
CN120765272A
Traditional Chinese medicine decoction piece production data management method and system based on quality tracing
CN121119444A