Food production safety multi-source heterogeneous data acquisition and processing method based on data mining
By collecting, storing, transforming, and analyzing multi-source heterogeneous data from the food production process, a multi-factor analysis system is constructed, solving the problem of identifying and analyzing data with common influences that cannot be found in existing technologies, and realizing real-time monitoring and accurate supervision of food production safety.
Patent Information
- Application Number
- CN202511470835.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-15
- Publication Date
- 2026-01-09
AI Technical Summary
Existing food production safety data processing technologies are unable to effectively identify and analyze multi-source heterogeneous data with common influence, leading to inaccurate food safety supervision and an inability to guarantee the safety of produced food.
By collecting and storing multi-source heterogeneous data in the food production safety process, converting it into effective data, dividing it into variable groups and dependent groups, constructing a multi-factor analysis complex, exploring the influence relationship between dependent groups and variable groups, and conducting real-time monitoring.
It improves the accuracy and rationality of food production safety supervision, and can take into account the combined effects of different influencing factors, enabling real-time monitoring and accurate analysis of food production safety.
Smart Images

Figure CN121302199A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of food production safety data processing, in particular to a food production safety multi-source heterogeneous data collection and processing method based on data mining. BACKGROUND
[0002] Food production safety data processing technology refers to the technology of organizing and analyzing data generated during food production to monitor food safety. The core is to find the impact of different data on food safety, so as to discover or predict production anomalies in the production process.
[0003] The existing food production safety data processing technology usually adopts the traditional threshold comparison or index calculation method to analyze and process food production safety data. The traditional threshold comparison method has low accuracy, and there is a certain common effect between different food production safety data, that is, a certain safety index may be affected by multiple food production safety data. The index calculation method cannot identify and analyze such data with common influence, but treats all data equally. This will lead to a large deviation between the analysis result and the actual result, which is not conducive to the control of food production safety data. For example, in the patent application with the publication number CN120235497A, a food processing production intelligent control method and system based on the Internet of Things is disclosed. This scheme processes food production safety data by index calculation and threshold comparison to monitor food production safety. This method cannot identify and analyze data with common influence, so the monitoring of food production safety is not accurate enough. The existing food production safety data processing technology cannot identify and analyze data with common influence, which leads to the problem that the safety of the produced food cannot be guaranteed. SUMMARY
[0004] This invention aims to at least partially address one of the technical problems in the existing technology. It involves collecting and storing multi-source heterogeneous data from the food production safety process, extracting data from this data, converting it into valid data, integrating the valid data, dividing it into variable groups and dependent groups, converting the valid data in the variable groups into a unified format, and simultaneously converting the valid data in the dependent groups into qualified or unqualified categories. Then, monitoring personnel divide the influencing factors into different factor analysis groups, construct a multi-factor analysis complex, input the factor analysis groups into the multi-factor analysis complex to obtain factor analysis data, analyze the factor analysis data to determine the relationship between influencing factors and safety items, and finally conduct real-time monitoring of food production safety based on the influence relationships. This addresses the problem that existing food production safety data processing technologies cannot identify and analyze data with common influences, leading to compromised food safety.
[0005] To achieve the above objectives, this application provides a method for collecting and processing multi-source heterogeneous data on food production safety based on data mining, comprising the following steps: Collect and store multi-source heterogeneous data during the food production safety process; Extract data from multi-source heterogeneous data and transform it into effective data; The effective data is integrated and divided into variable groups and dependent variable groups; Convert the valid data in the variable group into a uniform format, and at the same time convert the valid data in the dependent variable group into qualified or unqualified; A multi-factor analysis complex is constructed based on the variable group and the dependent group, and the effective data is analyzed to explore the influence relationship between the dependent group and the variable group. Real-time monitoring of food production safety based on influencing relationships.
[0006] Furthermore, the collection and storage of multi-source heterogeneous data in the food production safety process includes the following sub-steps: Collect multi-source heterogeneous data during the food production safety process, including equipment sensor data, environmental monitoring data, and quality testing data. Store heterogeneous data from multiple sources.
[0007] Furthermore, data extraction from multi-source heterogeneous data, transforming it into effective data, includes the following sub-steps: Construct an effective thesaurus, which records all words related to food production safety; By using keyword extraction technology and semantic recognition technology, words that are in the effective word library are extracted from the equipment sensing data, environmental monitoring data and quality inspection data, and the equipment data, environmental data and quality data are obtained. The equipment data, environmental data and quality data all contain timestamps, and the quality data also contains batch numbers. Combine device data, environmental data, and quality data with the same timestamp into a single valid data entry.
[0008] Further, the effective data is integrated and divided into variable groups and dependent groups, including the following sub-steps: The valid data is stored in a database, which is named the basic database; The column headings in the base database are divided into variable groups and dependent groups based on nouns and adjectives, respectively. Nouns belong to the variable group, and adjectives belong to the dependent group.
[0009] Furthermore, converting the valid data in the variable group into a uniform format, and simultaneously converting the valid data in the dependent variable group into qualified or unqualified data, includes the following sub-steps: Name the nouns in the variable group as influencing factors, normalize the influencing factors in the variable group, and convert the values of the influencing factors in each valid data point into normalized values. The adjectives in the dependent variable group are named safety items. Based on the production safety standards of the safety items, the content of the safety items in each valid data is uniformly converted into qualified or unqualified expressions.
[0010] Furthermore, constructing a multifactorial analysis framework based on the variable group and dependent variable group, and analyzing the effective data to explore the influence relationship between the dependent variable group and the variable group includes the following sub-steps: The monitoring personnel divided the influencing factors into different factor analysis groups; Construct a multi-factor analysis complex, input the factor analysis groups into the multi-factor analysis complex, and obtain the factor analysis data; Analyze the factor analysis data to determine the relationship between influencing factors and safety items.
[0011] Furthermore, the process of dividing influencing factors into different factor analysis groups by monitoring personnel includes the following sub-steps: The monitoring personnel grouped the influencing factors. When grouping, the number of influencing factors in each factor analysis group was named the number of factors. The number of factors varied from group to group. The factor analysis group obtained by the monitoring personnel is named the priority analysis group. Different combinations of influencing factors are made, and all existing combinations of influencing factors are listed and named the all-round combination. All-round combinations that are the same as those in the priority analysis group are removed, and the remaining all-round combinations are named the lag analysis group. Both the priority analysis group and the lag analysis group are factor analysis groups, and the priority analysis group is analyzed first in subsequent analyses.
[0012] Furthermore, a multi-factor analysis complex is constructed, and the factor analysis groups are entered into the multi-factor analysis complex to obtain factor analysis data, including the following sub-steps: Obtain the total number of influencing factors, named the total number of parameters, construct a regular polygon with the total number of parameters, and define each edge of the regular polygon as representing an influencing factor. This yields a multi-factor analysis complex. Name the edges of the multi-factor analysis complex as parameter lines, and name the influencing factors corresponding to the parameter lines as representative parameters. The values on each parameter line represent 0 to 1 in a counter-clockwise order. The minimum value on the parameter line is 0, and the maximum value is 1, which are the two endpoints of the parameter line. When analyzing any factor analysis group, it is named the target analysis group, and the influencing factors in the target analysis group are named target analysis factors. The target analysis factors are then numbered and labeled with the symbol TAF. n This indicates that n is a positive integer and n is the index of the TAF. n The corresponding parameter line is marked as PL n ; When entering a valid piece of data, use TAF. n The value of PL n The corresponding point is marked as PP. n , through PP n PL n The perpendicular line, marked RL n The RL n Consider it to be infinitely long; Different RL nAn intersection point will appear between them, which is named the parameter junction point. The leftmost parameter junction point is marked as the initial loop point. A two-dimensional coordinate system is established with the horizontal direction in the multi-factor analysis complex as the X-axis and the vertical direction as the Y-axis, named the quadrant auxiliary diagram. The quadrant auxiliary diagram is constructed with the initial loop point as the origin. The parameter junction point closest to the initial loop point in the fourth quadrant is found. If there is no parameter junction point in the fourth quadrant, the parameter junction point closest to the initial loop point in the first quadrant is found. If there is no parameter junction point in the first quadrant, the parameter junction point closest to the initial loop point in the second quadrant is found. If there is no parameter junction point in the second quadrant, the parameter junction point closest to the initial loop point in the third quadrant is found and named the adjacent junction point. The initial loop point is connected to the adjacent junction point to obtain the parameter influence line. Rename the initial loop point to a connected loop point, and rename the adjacent connection points to initial loop points. Repeat the analysis of parameter influence lines until all parameter connection points become connected loop points. Use the midpoint of all parameter influence lines as the new parameter connection points, and repeat the analysis of parameter influence lines until only one parameter connection point exists. Name the final parameter connection point as the multi-factor influence point. Each valid data point corresponds to a multi-factor influence point. All valid data are entered into a multi-factor analysis system to obtain different multi-factor influence points, which are the factor analysis data.
[0013] Further analysis of the factor analysis data to determine the relationship between influencing factors and safety items includes the following sub-steps: Cluster analysis was performed on the points affected by multiple factors to obtain different clusters, which were named multi-factor influence clusters. When analyzing whether there is a correlation between the target analysis group and any safety item, the corresponding safety item is named the target item. The number of target items in the multi-factor influence points of the multi-factor influence cluster is counted as the number of non-compliant items, and marked as G1. At the same time, the number of multi-factor influence points in the multi-factor influence cluster is counted and marked as G2. Calculate G1 / G2 and name the result as the anomaly rate. Calculate the anomaly rate of each multi-factor cluster. If there is a multi-factor cluster with an anomaly rate of 1, mark that there is an influence relationship between the target analysis group and the target item. Then, form a safety reference group with the influence relationship between the target analysis group and the target item. The range of anomalies in the statistical safety reference group is named the anomaly range. The anomaly range is then divided into three sub-ranges, which are named the anomaly sub-ranges. These sub-ranges are named the normal value range, the risk value range, and the anomaly value range in ascending order.
[0014] Furthermore, real-time monitoring of food production safety based on influencing relationships includes the following sub-steps: When analyzing any safety reference group, the multi-factor influence points are analyzed based on the influencing factors in the safety reference group and named as real-time judgment points; Find the cluster of multi-factor influences that is closest to the real-time judgment point, name it the real-time judgment cluster, and obtain the anomaly rate of the real-time judgment cluster, name it the real-time anomaly parameter; The system determines the sub-range of an anomaly parameter. If the real-time anomaly parameter is within the risk value range, a risk signal is sent to the control center. If the real-time anomaly parameter is within the anomaly value range, an anomaly signal is sent to the control center. The anomaly signal has a higher priority than the risk signal.
[0015] The beneficial effects of this invention are as follows: This invention collects and stores multi-source heterogeneous data from the food production safety process, then extracts the data from the multi-source heterogeneous data, converts it into effective data, integrates the effective data, divides it into variable groups and dependent groups, converts the effective data in the variable groups into a unified format, and converts the effective data in the dependent groups into qualified or unqualified. Then, monitoring personnel divide the influencing factors into different factor analysis groups, and construct a multi-factor analysis complex. The factor analysis groups are entered into the multi-factor analysis complex. The advantage is that it converts multi-source heterogeneous food production safety data into normalized data with a unified dimension, and then constructs a multi-factor analysis complex. The multi-factor analysis complex can reflect the characteristics under the combined effect of different influencing factors, improving the accuracy and rationality of food production safety supervision. This invention analyzes factor analysis data to determine the relationship between influencing factors and safety items, and finally monitors food production safety in real time based on the relationship. Its advantage lies in considering the combination of different influencing factors and the relationship between different safety items. When judging whether food production safety is qualified, it can combine multiple influencing factors for simultaneous analysis, rather than analyzing each influencing factor independently, thus improving the accuracy and effectiveness of food production safety supervision. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating the steps of the method of the present invention; Figure 2 This is a schematic diagram of the multi-factor analysis complex of the present invention; Figure 3 The RL of the present invention n A schematic diagram; Figure 4 This is a schematic diagram of the parameter influence lines of the present invention; Figure 5 This is a schematic diagram of the factor analysis data of the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] Example 1, please refer to Figure 1 As shown, this application provides a method for collecting and processing multi-source heterogeneous data on food production safety based on data mining, including the following steps: Step S1 involves collecting and storing multi-source heterogeneous data from the food production safety process; Step S1 includes the following sub-steps: Step S101: Collect multi-source heterogeneous data in the food production safety process. The multi-source heterogeneous data includes equipment sensor data, environmental monitoring data, and quality testing data. Step S102: Store the multi-source heterogeneous data; In practice, for example, a piece of equipment sensor data might be {production line number: "PL_01", timestamp: "2025-05-16 15:36:28", temperature (°C): "68.2"}. Equipment sensor data refers to the parameters of different production equipment in the production workshop during operation. Environmental monitoring data is mainly used to monitor environmental data in the production workshop, such as indoor temperature, humidity, and bacterial count. Quality inspection data refers to the results obtained from quality inspection of products, such as whether there are any abnormalities in the product's color, taste, or texture. For example, a piece of quality inspection data might describe the food as "milky white, odorless, with normal viscosity and normal ingredient content".
[0019] Step S2 involves extracting data from the multi-source heterogeneous data and converting it into valid data. Step S2 includes the following sub-steps: Step S201: Construct an effective vocabulary database, which contains all words related to food production safety. Step S202: Using keyword extraction technology and semantic recognition technology, extract words from the device sensing data, environmental monitoring data and quality inspection data that are in the effective word library by comparing them with the effective word library, and obtain device data, environmental data and quality data. The device data, environmental data and quality data all contain timestamps, and the quality data also contains batch numbers. Step S203: Combine device data, environmental data, and quality data with the same timestamp into a single valid data entry; In specific implementation, the effective term library records all words related to food production safety. The effective term library can be obtained through big data or set by monitoring personnel according to the product; this embodiment will not elaborate further. The effective term library is mainly used to extract effective data from multi-source heterogeneous data. The processing of multi-source heterogeneous data can also be completed using existing technologies. This embodiment does not impose detailed restrictions on the extraction of multi-source heterogeneous data. Because there are too many factors affecting food safety, this embodiment cannot list them all. Therefore, this embodiment only uses a few factors as examples to illustrate the specific analysis process of this embodiment. For example, the extracted equipment data is {Equipment A temperature: "68.2℃", timestamp: "2025-05-16 15:36:28"}, environmental data is {Workshop temperature: "26.5℃", workshop humidity: "45%", workshop colony count: "2" / dish, timestamp: "2025-05-16 15:36:28"}, and quality data is {Normal color, normal taste, normal viscosity, timestamp: "2025-05-16"}. The three data entries, "15:36:28", "Batch No.: G20250516001", have the same timestamp and are therefore merged into one valid data entry. The valid data is: {Equipment A temperature: 68.2℃, Workshop temperature: 26.5℃, Workshop humidity: 45%, Workshop colony count: 2 / plate, normal color, normal taste, normal viscosity, timestamp: "2025-05-16 15:36:28", batch number: "G20250516001"}.
[0020] Step S3 involves integrating the valid data and dividing it into variable groups and dependent variable groups; Step S3 includes the following sub-steps: Step S301: Store the valid data in the database and name it the basic database; Step S302: Divide the column headings in the basic database into variable groups and dependent groups according to nouns and adjectives, respectively. Nouns belong to the variable group and adjectives belong to the dependent group. In practice, most of the monitoring results for food safety and quality are adjectives, such as "normal color," "normal taste," and "normal viscosity" mentioned in this embodiment, or "abnormal color," "abnormal taste," and "abnormal viscosity." However, there are also a few parameters that are not adjectives, such as the content of a certain element in food. This type of data can be classified by the monitoring personnel themselves. The variable group is actually a grouping of parameters that producers can directly monitor and control during the food production process, while the dependent group is actually a parameter related to food safety assessment that producers cannot directly monitor and control during the food production process. This embodiment does not make specific requirements on the classification method of variable group and dependent group; existing classification standards can be referred to.
[0021] Step S4 involves converting the valid data in the variable group into a uniform format, and simultaneously converting the valid data in the dependent variable group into qualified or unqualified data. Step S4 includes the following sub-steps: Step S401: Name the nouns in the variable group as influencing factors, normalize the influencing factors in the variable group, and convert the values of the influencing factors in each valid data point into normalized values. Step S402: Name the adjectives in the dependent variable group as safety items, and based on the production safety standards of the safety items, uniformly convert the content of the safety items in each valid data into qualified or unqualified expressions; In specific implementation, normalization is performed using existing normalization techniques, which will not be described in detail in this embodiment. All data in the dependent variable group are safety items, and each safety item has a corresponding safety standard. The safety standard can refer to existing standards. For example, "normal color" can be transformed into "qualified" color, where color is a safety item and "qualified" is a specific attribute of color. The processing and extraction of multi-source heterogeneous data can refer to existing technologies. The extraction of safety items and influencing factors also uses existing technologies. The focus of this embodiment is on the subsequent analysis of different influencing factors to explore whether different safety items are affected by the combined effect of different influencing factors. Data preprocessing is not important in this embodiment, so it will not be described in detail in this embodiment.
[0022] Step S5 involves analyzing the effective data based on the variable group and the dependent variable group to uncover the influence relationship between the dependent variable group and the variable group. Step S5 includes the following sub-steps: Step S501: The monitoring personnel divide the influencing factors into different factor analysis groups; Step S501 includes the following sub-steps: Step S5011: The monitoring personnel group the influencing factors. When grouping, the number of influencing factors in each factor analysis group is named the number of factors. The number of factors in each factor analysis group may differ. Step S5012: The factor analysis group obtained by the monitoring personnel is named the priority analysis group. Different combinations of influencing factors are made, and all existing combinations of influencing factors are listed and named the all-round combination. The all-round combination that is the same as the priority analysis group is removed, and the remaining all-round combination is named the lag analysis group. Both the priority analysis group and the lag analysis group are factor analysis groups. In the subsequent analysis, the priority analysis group is analyzed first. In practice, for example, the influencing factors in one priority analysis group include the temperature of equipment A and the workshop temperature, while the influencing factors in another priority analysis group include the workshop temperature, workshop humidity and workshop colony count. The priority analysis group is only used to determine the priority of the analysis. In this embodiment, all combinations of influencing factors will be analyzed in the end. Therefore, the division of priorities is not important. Moreover, the number of influencing factors in the factor analysis group is not fixed when dividing the factor analysis group.
[0023] Step S502: Construct a multi-factor analysis complex, input the factor analysis groups into the multi-factor analysis complex, and obtain factor analysis data; Step S502 includes the following sub-steps: Please see Figure 2 As shown, in step S5021, the total number of influencing factors is obtained and named as the total number of parameters. A regular polygon with the total number of parameters is constructed. Each edge of the regular polygon is defined, that is, each edge represents an influencing factor, and a multi-factor analysis complex is obtained. The edges of the multi-factor analysis complex are named as parameter lines, and the influencing factors corresponding to the parameter lines are named as representative parameters. The values on each parameter line are in counterclockwise order, representing 0 to 1. The minimum value on the parameter line is 0 and the maximum value is 1, which are the two endpoints of the parameter line. In specific implementation, taking the influencing factors listed in this embodiment as an example, the total number of parameters is 4. Therefore, a regular quadrilateral, i.e., a square, is constructed to obtain the multi-factor analysis complex, as shown below. Figure 2 As shown, one side of the square is a parameter line, and there are a total of 4 parameter lines, each corresponding to one of the 4 influencing factors. Step S5022: When analyzing any factor analysis group, name it the target analysis group, name the influencing factors in the target analysis group the target analysis factors, number the target analysis factors, and use the symbol TAF. n This indicates that n is a positive integer and n is the index of the TAF. n The corresponding parameter line is marked as PL n ; Please see Figure 3 As shown, in step S5023, when entering a valid piece of data, the TAF is... n The value of PL n The corresponding point is marked as PP. n , through PP n PL n The perpendicular line, marked RL n RL n Consider it to be infinitely long; In practice, taking the factor analysis group consisting of workshop temperature, workshop humidity, and workshop colony count as the target analysis group as an example, workshop temperature, workshop humidity, and workshop colony count are the target analysis factors, which are numbered to obtain TAF. n and PL n 1 ≤ n ≤ 3. For example, a valid data point is {Equipment A temperature: 0.82, workshop temperature: 0.89, workshop humidity: 0.71, workshop colony count: 0.4, color: "qualified", taste: "qualified", viscosity: "qualified", timestamp: "2025-05-16 15:36:28", batch number: "G20250516001"}. PP1 is 0.89, PP2 is 0.71, and PP3 is 0.4. Draw a perpendicular line from the point on PL1 at 0.89 to PL1, from the point on PL2 at 0.71 to PL2, and from the point on PL3 at 0.4 to PL3. This will ultimately construct the RL. n like Figure 3 As shown, Figure 3 The dashed line in the middle is RL. n .
[0024] Please see Figure 4 As shown, in step S5024, different RL n An intersection point will appear between them, which is named the parameter junction point. The leftmost parameter junction point is marked as the initial loop point. A two-dimensional coordinate system is established with the horizontal direction in the multi-factor analysis complex as the X-axis and the vertical direction as the Y-axis, named the quadrant auxiliary diagram. The quadrant auxiliary diagram is constructed with the initial loop point as the origin. The parameter junction point closest to the initial loop point in the fourth quadrant is found. If there is no parameter junction point in the fourth quadrant, the parameter junction point closest to the initial loop point in the first quadrant is found. If there is no parameter junction point in the first quadrant, the parameter junction point closest to the initial loop point in the second quadrant is found. If there is no parameter junction point in the second quadrant, the parameter junction point closest to the initial loop point in the third quadrant is found and named the adjacent junction point. The initial loop point is connected to the adjacent junction point to obtain the parameter influence line. Step S5025: Rename the initial loop point to a connected loop point, rename the adjacent connection point to an initial loop point, and repeat the analysis of the parameter influence line until all parameter connection points become connected loop points; take the midpoint of all parameter influence lines as the new parameter connection point, and repeat the analysis of the parameter influence line until there is only one single parameter connection point, and name the final parameter connection point as the multi-factor influence point. Please see Figure 5 As shown in step S5026, each valid data point corresponds to a multi-factor influence point. All valid data are entered into the multi-factor analysis complex to obtain different multi-factor influence points. The multi-factor influence points are the factor analysis data. In practice Figure 3 There are two parameter binding points in total. In this embodiment, the upper parameter binding point is named the first binding point, and the lower parameter binding point is named the second binding point. Since both the first and second binding points are on the far left, the lowermost parameter binding point is selected as the initial loop point, i.e., the second binding point is named the initial loop point. The constructed quadrant auxiliary diagram overlaps with the cross line formed by RL3 and RL2, so it will not be specifically shown in this embodiment. At this time, there are no parameter binding points in the fourth quadrant. Therefore, the parameter binding point closest to the initial loop point in the first quadrant is searched. Points in the positive Y-axis direction are considered to be in the first and second quadrants. The first binding point is found as the adjacent binding point, and the parameter influence line is obtained by connecting them as shown in the figure. Figure 4 As shown, searching for adjacent connection points in the order of the fourth quadrant, first quadrant, second quadrant, and third quadrant is to ensure that the parameter connection points form a loop, thus gradually narrowing the search inwards to ultimately obtain the multi-factor influence point. Without this process, the formation of the multi-factor influence point would be uncontrollable and lack reference value. The second connection point is marked as the connected loop point, and the first connection point is marked as the initial loop point. At this point, there are no parameter connection points, so the analysis of the parameter influence line is stopped. The midpoint of each parameter influence line is taken as a new parameter connection point. Now, only one parameter connection point exists; this parameter connection point is the multi-factor influence point, and the final factor analysis data is as follows: Figure 5 As shown.
[0025] Step S503: Analyze the factor analysis data to determine the relationship between influencing factors and safety items; Step S503 includes the following sub-steps: Step S5031: Perform cluster analysis on the points affected by multiple factors to obtain different clusters, which are named multi-factor influence clusters; Step S5032: When analyzing whether there is a correlation between the target analysis group and any safety item, the corresponding safety item is named the target item. The number of target items that are not qualified in the multi-factor influence points of the multi-factor influence cluster is counted and marked as G1. At the same time, the number of multi-factor influence points in the multi-factor influence cluster is counted and marked as G2. Step S5033: Calculate G1 / G2, name the calculation result as the anomaly rate, calculate the anomaly rate of each multi-factor influence cluster, if there is a multi-factor influence cluster with an anomaly rate of 1, mark that there is an influence relationship between the target analysis group and the target item, and form a safety reference group by combining the target analysis group and the target item with the influence relationship. Step S5034: Calculate the range of abnormality rates in the safety reference group and name it as the abnormal range. Divide the abnormal range into three sub-ranges on average and name them as abnormal sub-ranges. Name the abnormal sub-ranges as normal value range, risk value range and abnormal value range in ascending order. In practice, cluster analysis employs existing clustering algorithms, resulting in nine multi-factor influence clusters. Taking one of these clusters as an example, when analyzing the safety item color, color is the target item. A multi-factor influence point is obtained from the analysis of one valid data point. If the color in this valid data point is unqualified, then the target item of the multi-factor influence point is unqualified. Statistically, G1 is 84 and G2 is 132. Finally, the anomaly rate of the current multi-factor influence cluster for the target item is calculated to be 7 / 11. The anomaly rate of each multi-factor influence cluster is calculated; if there are... A clustering with an anomaly rate of 1 indicates that when the multi-factor influence point is in this cluster under the combined effect of these three influencing factors, the color of the produced food will be unqualified regardless of how other influencing factors change. This means that there is an influence relationship between the target analysis group and the target item. The target analysis group and the target item are combined into a safety reference group. Since the color of the valid data is qualified not only depends on the influencing factors in the target analysis group, but may also be interfered with by other influencing factors, it is necessary to divide the abnormal range. Finally, the normal value range, risk value range, and abnormal value range are obtained.
[0026] Step S6 involves real-time monitoring of food production safety based on the influencing relationships. Step S6 includes the following sub-steps: Step S601: When analyzing any safety reference group, analyze the multi-factor influence points based on the influencing factors in the safety reference group and name them as real-time judgment points; Step S602: Find the multi-factor influence cluster that is closest to the real-time judgment point, name it the real-time judgment cluster, and obtain the anomaly rate of the real-time judgment cluster, name it the real-time anomaly parameter. Step S603: Determine the abnormal sub-range to which the real-time abnormal parameter belongs. If the real-time abnormal parameter is within the risk value range, send a risk signal to the control center. If the real-time abnormal parameter is within the abnormal value range, send an abnormal signal to the control center. The abnormal signal has a higher priority than the risk signal. In practice, the real-time judgment point can be analyzed according to the process of analyzing the multi-factor influence points. This embodiment will not be explained in detail. Finding the real-time judgment cluster is actually finding the multi-factor influence cluster to which the multi-factor influence point closest to the real-time judgment point belongs. If the real-time abnormal parameter is within the normal range, and the abnormality rate within the normal range is low, it is easily affected by other influencing factors. Therefore, it is usually considered that the abnormal parameter is within the normal range at this time, which is a normal situation and there is no need to send an early warning.
[0027] Example 2: This application provides an electronic device, which may include a processor, a communication interface, a memory, and a communication bus. The processor, communication interface, and memory communicate with each other via the communication bus. The memory stores computer-readable instructions, and the processor can call these instructions. When the processor executes a computer-readable instruction, it performs steps such as those in the "Multi-Source Heterogeneous Data Acquisition and Processing Method for Food Production Safety Based on Data Mining," to achieve the following functions: acquiring and storing multi-source heterogeneous data from the food production safety process; converting the multi-source heterogeneous data into effective data; integrating the effective data and dividing it into variable groups and dependent groups; converting the effective data in the variable groups into a unified format; constructing a multi-factor analysis complex based on the variable groups and dependent groups and analyzing the effective data to mine the influence relationships between the dependent and variable groups; and performing real-time monitoring of food production safety based on these influence relationships.
[0028] Furthermore, when the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0029] Example 3: This application also provides a computer program product, which includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the data mining-based method for collecting and processing multi-source heterogeneous data on food production safety provided by the above methods. This method includes: collecting and storing multi-source heterogeneous data in the food production safety process; converting the multi-source heterogeneous data into effective data; integrating the effective data and dividing it into variable groups and dependent groups; converting the effective data in the variable groups into a unified format; constructing a multi-factor analysis complex based on the variable groups and dependent groups and analyzing the effective data to mine the influence relationship between the dependent groups and variable groups; and monitoring food production safety in real time based on the influence relationship.
[0030] Example 4: This application also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it performs the steps of the above-described method for collecting and processing multi-source heterogeneous data on food production safety based on data mining, to achieve the following functions: collecting and storing multi-source heterogeneous data in the food production safety process; converting multi-source heterogeneous data into effective data; integrating the effective data and dividing it into variable groups and dependent groups; converting the effective data in the variable groups into a unified format; constructing a multi-factor analysis complex based on the variable groups and dependent groups and analyzing the effective data to mine the influence relationship between the dependent groups and variable groups; and monitoring food production safety in real time based on the influence relationship.
[0031] Based on the above description of the embodiments, the embodiments of the present invention can be provided as methods, systems, or computer program products. Based on this understanding, the above technical solutions, in essence or in terms of their contribution to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or certain parts of the embodiments.
[0032] In the embodiments provided in this application, it should be understood that the disclosed system or method can be implemented in other ways. The embodiments described above are merely illustrative. For example, the division of modules or units is only a logical functional division, and there may be other division methods in actual implementation. Furthermore, multiple modules or units may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the coupling or direct coupling or communication connection shown or discussed may be through some communication interfaces. The indirect coupling or communication connection between systems, modules, and units may be electrical, mechanical, or other forms.
[0033] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for collecting and processing multi-source heterogeneous data on food production safety based on data mining, characterized in that, Includes the following steps: Collect and store multi-source heterogeneous data during the food production safety process; Extract data from multi-source heterogeneous data and transform it into effective data; The effective data is integrated and divided into variable groups and dependent variable groups; Convert the valid data in the variable group into a uniform format, and at the same time convert the valid data in the dependent variable group into qualified or unqualified; A multi-factor analysis complex is constructed based on the variable group and the dependent group, and the effective data is analyzed to explore the influence relationship between the dependent group and the variable group. Real-time monitoring of food production safety based on influencing relationships.
2. The method for acquiring and processing multi-source heterogeneous data on food production safety based on data mining according to claim 1, characterized in that, The collection and storage of multi-source heterogeneous data in the food production safety process includes the following sub-steps: Collect multi-source heterogeneous data during the food production safety process, including equipment sensor data, environmental monitoring data, and quality testing data. Store heterogeneous data from multiple sources.
3. The method for acquiring and processing multi-source heterogeneous data on food production safety based on data mining according to claim 2, characterized in that, Extracting data from multi-source heterogeneous data and converting it into effective data includes the following sub-steps: Construct an effective thesaurus, which records all words related to food production safety; By using keyword extraction technology and semantic recognition technology, words that are in the effective word library are extracted from the equipment sensing data, environmental monitoring data and quality inspection data, and the equipment data, environmental data and quality data are obtained. The equipment data, environmental data and quality data all contain timestamps, and the quality data also contains batch numbers. Combine device data, environmental data, and quality data with the same timestamp into a single valid data entry.
4. The method for acquiring and processing multi-source heterogeneous data on food production safety based on data mining according to claim 3, characterized in that, Integrating valid data and dividing it into variable groups and dependent groups includes the following sub-steps: The valid data is stored in a database, which is named the basic database; The column headings in the base database are divided into variable groups and dependent groups based on nouns and adjectives, respectively. Nouns belong to the variable group, and adjectives belong to the dependent group.
5. The method for acquiring and processing multi-source heterogeneous data on food production safety based on data mining according to claim 4, characterized in that, Converting valid data in the variable group to a uniform format, and simultaneously converting valid data in the dependent variable group to qualified or unqualified, includes the following sub-steps: Name the nouns in the variable group as influencing factors, normalize the influencing factors in the variable group, and convert the values of the influencing factors in each valid data point into normalized values. The adjectives in the dependent variable group are named safety items. Based on the production safety standards of the safety items, the content of the safety items in each valid data is uniformly converted into qualified or unqualified expressions.
6. The method for acquiring and processing multi-source heterogeneous data on food production safety based on data mining according to claim 5, characterized in that, Constructing a multifactorial analysis framework based on variable groups and dependent variable groups, and analyzing the effective data to explore the influence relationship between the dependent variable group and the variable group, includes the following sub-steps: The monitoring personnel divided the influencing factors into different factor analysis groups; Construct a multi-factor analysis complex, input the factor analysis groups into the multi-factor analysis complex, and obtain the factor analysis data; Analyze the factor analysis data to determine the relationship between influencing factors and safety items.
7. The method for acquiring and processing multi-source heterogeneous data on food production safety based on data mining according to claim 6, characterized in that, The process of dividing influencing factors into different factor analysis groups by monitoring personnel includes the following sub-steps: The monitoring personnel grouped the influencing factors. When grouping, the number of influencing factors in each factor analysis group was named the number of factors. The number of factors varied from group to group. The factor analysis group obtained by the monitoring personnel is named the priority analysis group. Different combinations of influencing factors are made, and all existing combinations of influencing factors are listed and named the all-round combination. All-round combinations that are the same as those in the priority analysis group are removed, and the remaining all-round combinations are named the lag analysis group. Both the priority analysis group and the lag analysis group are factor analysis groups, and the priority analysis group is analyzed first in subsequent analyses.
8. The method for acquiring and processing multi-source heterogeneous data on food production safety based on data mining according to claim 7, characterized in that, Constructing a multi-factor analysis complex, inputting factor analysis groups into the multi-factor analysis complex, and obtaining factor analysis data includes the following sub-steps: Obtain the total number of influencing factors, named the total number of parameters, construct a regular polygon with the total number of parameters, and define each edge of the regular polygon as representing an influencing factor. This yields a multi-factor analysis complex. Name the edges of the multi-factor analysis complex as parameter lines, and name the influencing factors corresponding to the parameter lines as representative parameters. The values on each parameter line represent 0 to 1 in a counter-clockwise order. The minimum value on the parameter line is 0, and the maximum value is 1, which are the two endpoints of the parameter line. When analyzing any factor analysis group, it is named the target analysis group, and the influencing factors in the target analysis group are named target analysis factors. The target analysis factors are then numbered and labeled with the symbol TAF. n This indicates that n is a positive integer and n is the index of the TAF. n The corresponding parameter line is marked as PL n ; When entering a valid piece of data, use TAF. n The value of PL n The corresponding point is marked as PP. n , through PP n PL n The perpendicular line, marked RL n The RL n Consider it to be infinitely long; Different RL n An intersection point will appear between them, which is named the parameter junction point. The leftmost parameter junction point is marked as the initial loop point. A two-dimensional coordinate system is established with the horizontal direction in the multi-factor analysis complex as the X-axis and the vertical direction as the Y-axis, named the quadrant auxiliary diagram. The quadrant auxiliary diagram is constructed with the initial loop point as the origin. The parameter junction point closest to the initial loop point in the fourth quadrant is found. If there is no parameter junction point in the fourth quadrant, the parameter junction point closest to the initial loop point in the first quadrant is found. If there is no parameter junction point in the first quadrant, the parameter junction point closest to the initial loop point in the second quadrant is found. If there is no parameter junction point in the second quadrant, the parameter junction point closest to the initial loop point in the third quadrant is found and named the adjacent junction point. The initial loop point is connected to the adjacent junction point to obtain the parameter influence line. Rename the initial loop point to a connected loop point, and rename the adjacent connection points to initial loop points. Repeat the analysis of parameter influence lines until all parameter connection points become connected loop points. Use the midpoint of all parameter influence lines as the new parameter connection points, and repeat the analysis of parameter influence lines until only one parameter connection point exists. Name the final parameter connection point as the multi-factor influence point. Each valid data point corresponds to a multi-factor influence point. All valid data are entered into a multi-factor analysis system to obtain different multi-factor influence points, which are the factor analysis data.
9. The method for acquiring and processing multi-source heterogeneous data on food production safety based on data mining according to claim 8, characterized in that, Analyzing factor analysis data to determine the relationship between influencing factors and safety items includes the following sub-steps: Cluster analysis was performed on the points affected by multiple factors to obtain different clusters, which were named multi-factor influence clusters. When analyzing whether there is a correlation between the target analysis group and any safety item, the corresponding safety item is named the target item. The number of target items in the multi-factor influence points of the multi-factor influence cluster is counted as the number of non-compliant items, and marked as G1. At the same time, the number of multi-factor influence points in the multi-factor influence cluster is counted and marked as G2. Calculate G1 / G2 and name the result as the anomaly rate. Calculate the anomaly rate of each multi-factor cluster. If there is a multi-factor cluster with an anomaly rate of 1, mark that there is an influence relationship between the target analysis group and the target item. Then, form a safety reference group with the influence relationship between the target analysis group and the target item. The range of anomalies in the statistical safety reference group is named the anomaly range. The anomaly range is then divided into three sub-ranges, named the anomaly sub-ranges. These sub-ranges are named the normal value range, the risk value range, and the anomaly value range in ascending order.
10. The method for acquiring and processing multi-source heterogeneous data on food production safety based on data mining according to claim 9, characterized in that, Real-time monitoring of food production safety based on influencing relationships includes the following sub-steps: When analyzing any safety reference group, the multi-factor influence points are analyzed based on the influencing factors in the safety reference group and named as real-time judgment points; Find the cluster of multi-factor influences that is closest to the real-time judgment point, name it the real-time judgment cluster, and obtain the anomaly rate of the real-time judgment cluster, name it the real-time anomaly parameter; The system determines the sub-range of an anomaly parameter. If the real-time anomaly parameter is within the risk value range, a risk signal is sent to the control center. If the real-time anomaly parameter is within the anomaly value range, an anomaly signal is sent to the control center. The anomaly signal has a higher priority than the risk signal.
Citation Information
Patent Citations
Food processing production intelligent control method and system based on Internet of Things
CN120235497A