Big data processing system applied to Internet of Things
By designing an IoT big data processing system that includes data preprocessing, rules engines and data storage modules, the problems of traditional systems in data formats, high computing pressure, high operation and maintenance costs and data silos are solved, and efficient real-time and offline data processing is achieved, reducing system environment requirements.
Patent Information
- Application Number
- CN202510207508.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
AI Technical Summary
Traditional big data processing systems have problems such as inconsistent data formats, high computing pressure, high operation and maintenance costs, low code reusability, and data island problems in the field of the Internet of Things, making it difficult to effectively process real-time and offline data.
A big data processing system applied to the Internet of Things is designed, including a acquisition module, a transmission module and a computing center in the cloud. The computing center includes a data preprocessing module, a rule engine module and a data storage module. The system can process real-time and offline data simultaneously, and reduce system environment requirements and solve data silos through a common rule engine and data storage algorithm.
Real-time computing and offline computing share a set of rules, with higher calculation completeness, low omission rate of calculation results, lower system operating environment requirements, and reduced server CPU frequency and memory usage.
Smart Images

Figure CN120144538A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data processing, and particularly to a big data processing system applied to the Internet of Things. Background Art
[0002] The Internet of Things is a technology that transmits device data to a network through information sensing devices according to agreed protocols for information exchange and communication. Big data technology is a new generation of architecture and technology designed to obtain value from high-frequency, large-capacity, and different-structured and -typed data more economically.
[0003] Traditional cloud computing rule engines are too single. In the face of fields with highly variable specific data structures, such as the Internet of Things field, the data formats of different product devices are different, so the corresponding rule engines for analyzing data also need to be different. With the continuous increase in products, the number of data engines is also increasing, the computing pressure on the server surges, and the operation and maintenance costs also increase accordingly. Secondly, when traditional data is processed, some data continuously enters the system, and the relationship between the data is not strong. The system requires a set of streaming processing logics to analyze each piece of data received in real time; some data arrives in batches, with an obvious time interval between batches, and the data within a batch has a strong relationship. The system needs to analyze all the data in this batch macroscopically and as a whole. Such a data processing system must have two sets of logics, one for offline processing and analyzing a batch of data, and one for real-time processing and analyzing each piece of data. This results in low code reusability, high demand iteration costs, and complex task handover and project management. Then, as the business grows, more and more storage engines are introduced, and data is stored in different ways, which causes poor data circulation and the problem of data islands. Finally, there are multiple data source links at the application layer, and in the case of heterogeneous processing logics, there are data inconsistency problems, with high problem troubleshooting costs, long cycles, and difficult traceability.
[0004] At present, for big data processing, traditional solutions generally develop two sets of computing logics, which are respectively used for real-time data processing and offline data processing; for each additional data processing rule, a rule engine needs to be added; a system can only process one type of data singly. For example, a system for processing image data cannot process text data. In this case, a big data processing system for the corresponding data type needs to be additionally added, and there is no good solution to the data island problem between systems; for the intermediate values and result values generated by data processing, they can only be stored in a log file in the way of programmers manually inserting log points, and then a log analysis system is additionally introduced to analyze the logs to establish a complete data link relationship. However, there are some deficiencies in traditional solutions. For example: 1. In traditional architecture solutions, because the same business solution often requires developing two sets of logics, there is logical redundancy; 2. A large number of rule engines need to be created for the numerous rules generated by real-world requirements, which sharply increases the operation and maintenance costs of the system and the requirements for the system environment; 3. A data processing system cannot process different types of data simultaneously; 4. It is difficult to obtain and establish an association relationship for the intermediate values and result values generated by data processing, and an additional log analysis system is required. Therefore, improvements are needed in this regard. Summary of the Invention
[0005] The purpose of the present invention is to provide a big data processing system applied to the Internet of Things to solve the problems raised in the above background technology.
[0006] To achieve the above purpose, the present invention provides the following technical solution: A big data processing system applied to the Internet of Things includes an acquisition module, a transmission module, and a computing center in the cloud. The computing center includes a data preprocessing module, a rule engine module, and a data storage module;
[0007] The acquisition module is used to collect the required data through sensors;
[0008] The transmission module is used to send the data to the computing center in the cloud by selecting the method of real-time uploading data or batch uploading data;
[0009] The computing center is used to perform complex analysis and processing on the original data, and then save the obtained result data to a data warehouse;
[0010] The data preprocessing module is used to receive data through different types of network protocols; then parse the received data and classify and split these data according to data processing algorithms; the data preprocessing module is also used to receive real-time data and offline data simultaneously, and label the data with "stream processing" or "batch processing" according to the needs of real-world requirements to facilitate subsequent distinction;
[0011] The rule engine module needs to first define the processing logic of the data, and provides a general code template class for common processing logics, which is used to implement operations such as label classification of data, filtering data by value, and regularizing and merging data. The rule engine module includes a primary data buffer pool and a secondary buffer pool;
[0012] The data storage module needs to first define the data storage rules, and then through a general storage engine, the data will be structured and logicalized according to the defined rules, and stored in the data warehouse in a unified instantiation manner. The data storage module has two functions. One is that there is a "capture point" set in the program. This algorithm will locate to this "capture point" to automatically track and capture the process values generated by each module, and establish the link relationship between the process data and the result data, so as to facilitate which key guiding data for subsequent analysis results are. At the same time, when data is abnormal, the entire data link can be checked, and the link relationship will be stored in a relational data. The other is that according to the different types of original data, the corresponding storage method will be automatically matched according to the data storage rules. For example, image data will be stored in a vector database, and text data will be stored in a document database or a relational database.
[0013] Preferably, the data preprocessing module includes a formatting preprocessing algorithm. The algorithm will automatically match the processing logic according to the different data types. The final effect is to convert the image data into a pixel matrix, and each element in the matrix is the RGB value of the pixel. Then, the matrix is processed by the preprocessing algorithm to obtain the eigenvalue of the image; the data preprocessing module splits and matches the text type data according to the physical model, and the final preprocessed result data will be temporarily stored in the primary data buffer pool for the rule engine module to use.
[0014] Preferably, the formatting preprocessing algorithm includes an image preprocessing algorithm, and the algorithm formula is:
[0015]
[0016] where P ij is the value of the i-th color channel of the j-th pixel, N is the total number of pixels in the image, and μ i is the color mean value of the image;
[0017] In order to be able to more accurately describe the image data in subsequent processing, additional correction calculations need to be performed through μ i to obtain the color class variance σ i and the class skewness S i , and the formulas are as follows:
[0018]
[0019] The formatting preprocessing algorithm also includes a text preprocessing algorithm for splitting the text according to the physical model fields.
[0020] Preferably, the primary data buffer pool filters out the data marked with "batch processing" in the data preprocessing part, and finds a suitable "batch node" to divide the data into multiple data segments, indicating that this segment of data is to be processed together as the same batch; for the data marked with "stream processing" in the data preprocessing part, one piece of data is one data segment, and each piece of data is processed separately; during data processing, the rule engine continuously reads the data segments in the primary buffer pool and finds the corresponding data processing template according to the rule matching algorithm to complete the calculation and analysis of the data.
[0021] The secondary data buffer pool is used to temporarily store the calculation results of the rule engine and establish an association relationship between the calculation results and the data segments in the primary buffer pool, which is conducive to troubleshooting which data causes the abnormal results during subsequent analysis of abnormal results.
[0022] Preferably, the rule matching algorithm uses the architecture style of an interpreter. This algorithm first reads the data processing configuration file, which defines processing rules such as label classification, value screening, and regularized merging for different data types, and can define the granularity of different types of processing. "Batch processing" and "stream processing" can concurrently process N pieces, and the value N is positively correlated with the configuration level of the operating environment. Finally, the preprocessed result data will be temporarily stored in the secondary data buffer pool for use by the data storage module.
[0023] When classifying, screening, and merging data of the image type, it is necessary to compare the similarity between the current data and the target data. The similarity formula is:
[0024] d = (μ 1 - μ 2 ) 2 + (σ 1 - σ 2 ) 2 + (s 1 - s 2 ) 2
[0025] Where μ 1 is the color mean of the current data, σ 1 is the color class variance of the current data, s 1 is the color class skewness of the current data, μ 2 is the color mean of the target data, σ 2 is the color class variance of the target data, s 2 is the color class skewness of the target data, and d is the similarity value. The larger this value is, the smaller the similarity between the two image data.
[0026] When classifying, filtering, and merging text-type data, it is necessary to compare the similarity between the current data and the target data. The similarity formula is as follows:
[0027]
[0028] Among them, A is the text set obtained after splitting the current data by the data preprocessing algorithm, B is the text set obtained after splitting the target data by the data preprocessing algorithm, R(A×B) represents the number of elements after finding the Cartesian product of A and B, 2R(C) represents the number of elements of the physical model field C, and J(C) represents the text similarity of these two text data with respect to the physical model field C.
[0029] Preferably, the data storage algorithm in the data storage module is an improved traditional consistent hashing algorithm. The specific implementation of this algorithm is as follows:
[0030] First, calculate the total weight of all nodes. The formula is as follows:
[0031]
[0032] Among them, W represents the comprehensive weight, and w i represents the weight of each node, which is usually positively correlated with the configuration of the node;
[0033] Then, calculate the virtual nodes of all nodes. The purpose is to enable nodes with higher configurations to have more virtual nodes. The formula is as follows:
[0034]
[0035] Among them, V is the pre-set virtual node base number, min is the minimum weight number, and v i is the number of virtual nodes owned by the i-th node;
[0036] Finally, calculate the position of the virtual node on the hash ring. The formula is as follows:
[0037] h(n ij )=(n i +j)mod M
[0038] Among them, n ij represents the position of the j-th virtual node of the node, and M is a fixed positive integer 5. When data is allocated, the physical node corresponding to the first virtual node found by the data in the specified direction is the actual storage location of the data.
[0039] Compared with the prior art, the beneficial effects of the present invention are:
[0040] A big data processing system applied to the Internet of Things proposed by the present invention can use a set of rules for both real-time calculation and offline calculation, and has a higher calculation integrity rate, with the calculation result omission rate <0.1%; data is stored in a common data warehouse in the same way, and outliers can be traced according to the calculation results; the environmental requirements for the system to run are lower, with low requirements for the server CPU frequency and low memory occupancy rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a schematic structural diagram of the system of the present invention.
[0042] Figure 2 It is a system flow chart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0044] Please refer to Figures 1 to 2 , the present invention provides a technical solution: a big data processing system applied to the Internet of Things, which performs formatted preprocessing, logical rule processing, and structured storage processing on data one by one through a progressive algorithm. The formatted preprocessing algorithm solves the compatibility problem of different types of data; the logical rule processing algorithm improves the traditional architecture method of "one rule, one engine" to a configuration through multiple rules, using a general engine to reduce the development, operation and maintenance costs of the system and reduce the system environment requirements; the structured storage processing algorithm establishes an association relationship between the intermediate values and result values generated in each stage according to the rules, and automatically stores the data in a suitable way into the data warehouse to solve the data island problem. This system includes an acquisition module, a transmission module, and a computing center in the cloud. The computing center includes a data preprocessing module, a rule engine module, and a data storage module;
[0045] The acquisition module is used to collect the required data through sensors, including environmental parameters, infrared induction, video images, etc.; the transmission module is used to select the method of real-time uploading data or batch uploading data and send the data to the computing center in the cloud; the computing center is used to perform complex analysis and processing on the original data, and then save the obtained result data into the data warehouse;
[0046] 2. The data preprocessing module is used to receive data through different types of network protocols, such as HTTP protocol, TCP protocol, UDP protocol, MQTT protocol; then parse the received data, and the parsable data types include image type and text type, and classify and split these data according to the data processing algorithm; the data preprocessing module is also used to receive real-time data and offline data simultaneously, and label the data with "stream processing" or "batch processing" according to the needs of actual requirements for subsequent differentiation; the data preprocessing module contains a formatting preprocessing algorithm, and the algorithm will automatically match the processing logic according to different data types. The final effect is to convert the image data into a pixel matrix, and each element in the matrix is the RGB value of the pixel, and then obtain the feature value of the image through the preprocessing algorithm; the data preprocessing module splits and matches the text type data according to the object model, and the final preprocessed result data will be temporarily stored in the first-level data buffer pool for the rule engine module to use;
[0047] The formatting preprocessing algorithm includes an image preprocessing algorithm, and the algorithm formula is:
[0048]
[0049] where P ij is the value of the i-th color channel of the j-th pixel, N is the total number of pixels in the image, and μ i is the color mean value of the image;
[0050] In order to be able to more accurately describe the image data in subsequent processing, additional correction calculations need to be performed through μ i to obtain the color class variance σ i and the class skewness S i , and the formulas are as follows:
[0051]
[0052]
[0053] The formatting preprocessing algorithm also includes a text preprocessing algorithm, which is used to split the text according to the object model fields. For example, the text data is "The monitoring data of sensor 1 shows that the humidity is normal, sensor 2 detects a temperature change, and sensor 3 detects an increase in wind speed", and the specified object model fields are "sensor 1" and "sensor 3", then the data obtained after splitting this text data is "The monitoring data of sensor 1 shows that the humidity is normal, and sensor 3 detects an increase in wind speed".
[0054] The rule engine module needs to first define the processing logic of data, and provides a general code template class for common processing logics, which is used to implement operations such as label classification of data, screening data by value, and regularizing and merging data. The rule engine module includes a primary data buffer pool and a secondary buffer pool;
[0055] The primary data buffer pool filters out the data marked with "batch processing" in the data preprocessing part, and finds a suitable "batch node" to divide the data into multiple data segments, indicating that this segment of data is to be processed together as the same batch; for the data marked with "stream processing" in the data preprocessing part, one piece of data is a data segment, and each piece of data is processed separately; during data processing, the rule engine continuously reads the data segments in the primary buffer pool, and finds the corresponding data processing template according to the rule matching algorithm to complete the calculation and analysis of the data; the secondary data buffer pool is used to temporarily store the calculation results of the rule engine, and establish an association relationship between the calculation results and the data segments in the primary buffer pool, which is conducive to troubleshooting which data caused the abnormal results during subsequent analysis of abnormal results;
[0056] The rule matching algorithm uses the architecture style of an interpreter. This algorithm first reads a data processing configuration file, which defines processing rules such as label classification, value screening, and regularizing and merging for different data types, and can define the granularity of different types of processing. "Batch processing" and "stream processing" can concurrently process N pieces, and the value of N is positively correlated with the configuration level of the operating environment. Finally, the preprocessed result data will be temporarily stored in the secondary data buffer pool for use by the data storage module. This configuration file will be loaded when the system starts or dynamically loaded through commands during system operation;
[0057] When classifying, screening, and merging data of the image type, it is necessary to compare the similarity between the current data and the target data. The similarity formula is:
[0058] d = (μ 1 - μ 2 ) 2 + (σ 1 - σ 2 ) 2 + (s 1 - s 2 ) 2
[0059] where μ 1 is the color mean of the current data, σ 1 is the color class variance of the current data, s 1 is the color class skewness of the current data, μ 2 is the color mean of the target data, σ 2 is the color class variance of the target data, s 2It is the color class skewness of the target data, and d is the similarity value. The larger this value is, the smaller the similarity between the two image data;
[0060] When classifying, screening, and merging text-type data, it is necessary to compare the similarity between the current data and the target data. The similarity formula is:
[0061]
[0062] Where A is the text set obtained after the current data is split by the data preprocessing algorithm, B is the text set obtained after the target data is split by the data preprocessing algorithm, R(A×B) represents the number of elements after finding the Cartesian product of A and B, 2R(C) represents the number of elements of the object model field C, and J(C) represents the text similarity of these two text data with respect to the object model field C. For example: The text of A is "The monitoring data of sensor 1 shows that the humidity is normal, and sensor 2 monitors an increase in wind speed", and the text of B is "The monitoring data of sensor 1 shows that the humidity is normal, and sensor 2 monitors a temperature change". Then we can get J(sensor 1) = 11 / 11 = 100%, J(sensor 2) = 3 / 7 = 42%.
[0063] The data storage module needs to first define the data storage rules, and then through a general storage engine, it will structure and logicalize the data according to the defined rules and store it in the data warehouse in a unified instantiation manner. The data storage module has two functions. One is that there is a "capture point" set in the program. This algorithm will locate to this "capture point" to automatically track and capture the process values of each module, and establish the link relationship between the process data and the result data, facilitating the identification of the key guiding data for subsequent analysis results. At the same time, when data is abnormal, it can troubleshoot the entire data link, and the link relationship will be stored as relational data; the other is to automatically match the corresponding storage method according to the type of the original data. For example, image data will be stored in a vector database, and text data will be stored in a document database or a relational database.
[0064] The data storage algorithm in the data storage module is an improved traditional consistent hashing algorithm. The traditional consistent hashing algorithm evenly distributes data and requests to be processed to each node. In reality, each node may have different storage capacities or processing capabilities. The specific implementation of this algorithm is as follows:
[0065] First, calculate the total weight of all nodes. The formula is as follows:
[0066]
[0067] Where W represents the comprehensive weight, and w i represents the weight of each node, which is usually positively correlated with the configuration of the node;
[0068] Then calculate the virtual nodes of all nodes. The purpose is to enable nodes with higher configurations to have more virtual nodes. The formula is as follows:
[0069]
[0070] where v is the pre-set virtual node base number, min is the minimum weight number, and v i is the number of virtual nodes owned by the i-th node;
[0071] Finally, calculate the positions of the virtual nodes on the hash ring. The formula is as follows:
[0072] h(n ij )=(n i +j)mod M
[0073] where n ij represents the position of the j-th virtual node of the node, and M is a fixed positive integer 5. When data is allocated, the physical node corresponding to the first virtual node found by the data in the specified direction is the actual storage location of the data.
[0074] This system first automatically identifies the data type through a data processing algorithm, and parses and preprocesses the data in different ways according to different data types to solve the data compatibility problem; uses a general rule engine to dynamically load the rule configuration file through a rule matching algorithm to implement different processing logics, reducing the number of instances of the rule engine; automatically captures the process values generated during the data processing process through a data storage algorithm, and establishes the association relationship between the data. Finally, the algorithm automatically stores the data in a suitable manner together to solve the data island problem. The difference is that this technical solution uses a progressive algorithm to process the data one by one, can perform offline and real-time processing on different types of data at the same time, can dynamically adjust the processing rules during the operation of the system, and can establish the link relationship between the result data and the process data. The storage engine can store the data in the data warehouse in a specified manner according to the set storage rules. For example, when mining a large amount of data to find useful data as training samples for AI, it is usually necessary to first screen out the useful data for classification, integration and tagging. The data processing rules for "tagging" can be formulated first, and the tag content is the calculation result of the rule engine. The corresponding relationship between the tag and the data set with the tag is the link relationship between the result value and the process value.
[0075] Illustrate with examples:
[0076] Now, 1000 devices are connected to the system simultaneously using the MQTT and TCP protocols. The devices use the MQTT protocol to cyclically send sensor data to the system at a rate of 1 message per second, and the devices use the TCP protocol to cyclically send image data to the system at a rate of 1 message per second.
[0077] The rules of the system rule engine are as follows:
[0078] 1. The first 30 pieces of data of each device are not processed.
[0079] 2. Text data is subjected to real-time merging calculation, and image data is subjected to offline calculation in batches of 10 pieces of data. Data with a similarity > 10% is grouped into one category.
[0080] Query the values in the data warehouse. Within the first 100 seconds of the system operation, the calculation results are as follows:
[0081]
[0082]
[0083] It can be seen from the results that the real-time calculation was performed 70,000 times, the offline calculation was performed 7,000 times. After merging, the text data with a similarity > 10% was divided into 3,216 categories, and the image data was divided into 1,805 categories after classification. The similarity of the data within the same category is > 10%. The data in this category can be labeled with "10% similarity" and used as samples for AI model training.
[0084] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A big data processing system applied to the Internet of Things, characterized by: It includes an acquisition module, a transmission module and a cloud computing center, wherein the computing center includes a data preprocessing module, a rule engine module and a data storage module; The acquisition module is used to collect required data through sensors; The transmission module is used to select a method of uploading data in real time or in batches to send the data to a computing center in the cloud; The computing center is used to perform complex analysis and processing on the original data, and then save the resulting data into the data warehouse; The data preprocessing module is used to receive data using different types of network protocols; Then parse the received data and classify and split the data according to the data processing algorithm; The data preprocessing module is also used to receive real-time data and offline data at the same time. According to the actual needs, the data is labeled as "stream processing" or "batch processing" to facilitate subsequent distinction; The rule engine module needs to define the data processing logic first, and provides a general code template class for common processing logic, which is used to implement operations such as label classification of data, data screening by value, and regular data merging. The rule engine module includes a primary data buffer pool and a secondary buffer pool; The data storage module needs to define the data storage rules first, and then the general storage engine will structure and logic the data according to the defined rules, and store it in the data warehouse in a unified instantiation method. The data storage module includes two functions. One is that the program is equipped with a "grabbing point". The algorithm will locate the "grabbing point" to automatically track and grab the process value of each module, and establish a link relationship between the process data and the result data, so as to facilitate the post-analysis of the key guiding data of the results. At the same time, when the data is abnormal, the entire data link can be checked, and the link relationship will be stored as relational data; the other is to automatically match the corresponding storage method according to the data storage rules according to the different types of the original data. The image data will be stored in the vector database, and the text data will be stored in the document database or relational database.
2. The big data processing system applied to the Internet of Things according to claim 1, characterized in that: The data preprocessing module includes a formatting preprocessing algorithm, which automatically matches the processing logic according to the different data types. The final effect is to convert the image data into a pixel matrix, each element in the matrix is the RGB value of the pixel, and then the matrix is passed through the preprocessing algorithm to obtain the eigenvalue of the image; The data preprocessing module splits and matches text-type data according to the object model. The final preprocessing result data will be temporarily stored in the primary data buffer pool for use by the rule engine module.
3. The big data processing system applied to the Internet of Things according to claim 2, characterized in that: The formatting preprocessing algorithm includes an image preprocessing algorithm, and the algorithm formula is: Where P ij is the value of the i-th color channel of the j-th pixel, N is the total number of pixels in the image, μ i is the color mean of the image; In order to be able to more accurately describe the image data in subsequent processing, it is also necessary to use μ i Perform additional correction calculation to obtain the color class variance σ i and class skewness S i , the formula is as follows: The formatting preprocessing algorithm also includes a text preprocessing algorithm for splitting the text according to the object model fields.
4. The big data processing system applied to the Internet of Things according to claim 1, characterized in that: The primary data buffer pool will filter out the data with the "batch processing" label in the data preprocessing part, and find a suitable "batch node" to divide the data into multiple data segments, indicating that this segment of data is to be processed together as a batch; for the data with the "stream processing" label in the data preprocessing part, one piece of data is a data segment, and each piece of data is processed separately; during data processing, the rule engine will continuously read the data segments in the primary buffer pool, find the corresponding data processing template according to the rule matching algorithm to complete the data calculation and analysis; The secondary data buffer pool is used to temporarily store the calculation results of the rule engine, and to establish an association relationship between the calculation results and the data segments in the primary buffer pool, which is helpful for troubleshooting which data caused the anomaly when analyzing the abnormal results later.
5. The big data processing system applied to the Internet of Things according to claim 4, characterized in that: The rule matching algorithm uses the interpreter architecture style. The algorithm first reads the data processing configuration file, which defines processing rules such as label classification, value screening, regularization merging, etc. for different data types, and can define the granularity of different types of processing. "Batch processing" and "stream processing" can process N items concurrently. The value N is positively correlated with the configuration level of the operating environment. The final pre-processed result data will be temporarily stored in the secondary data buffer pool for use by the data storage module; When classifying, filtering, and merging image data, the similarity between the current data and the target data needs to be compared. The similarity formula is: d=(μ1-μ2) 2 +(σ1-σ2) 2 +(s1-s2) 2 Where μ1 is the color mean of the current data, σ1 is the color class variance of the current data, S1 is the color class skewness of the current data, μ2 is the color mean of the target data, σ2 is the color class variance of the target data, S2 is the color class skewness of the target data, and d is the similarity value. The larger this value is, the smaller the similarity between the two image data is. When classifying, filtering, and merging text data, the similarity between the current data and the target data needs to be compared. The similarity formula is: Where A is the text set obtained after the current data is split by the data preprocessing algorithm, B is the text set obtained after the target data is split by the data preprocessing algorithm, R(A×B) represents the number of elements after the Cartesian product of A and B is calculated, 2R(C) represents the number of elements in the object model field C, and J(C) represents the text similarity between the two text data with respect to the object model field C.
6. The big data processing system applied to the Internet of Things according to claim 1, characterized in that: The data storage algorithm in the data storage module is an improved traditional consistent hash algorithm, and the specific implementation of the algorithm is as follows: First, calculate the sum of the weights of all nodes. The formula is as follows: Where W represents the weighted combination, w i Represents the weight of each node, which is usually positively correlated with the node configuration; Then calculate the virtual nodes of all nodes, the purpose is to allow nodes with higher configuration to have more virtual nodes, the formula is as follows: Where V is the virtual node base number set in advance, min is the minimum weight number, V i is the number of virtual nodes owned by the i-th node; Finally, the position of the virtual node on the hash ring is calculated using the following formula: h(n ij )=(n i +j)mod M Where n ij Indicates the position of the jth virtual node of the node. M is a fixed positive integer 5. When data is allocated, the physical node corresponding to the first virtual node found by the data in the specified direction is the actual storage location of the data.