Data acquisition method and system for equipment service management platform
By collecting and analyzing the data feature information in the equipment update log database, establishing a data analysis model, and performing logistic regression calculations, the problems of low data acquisition efficiency and waste of computing resources in the existing technology are solved, and efficient and accurate data screening and collection are achieved.
Patent Information
- Application Number
- CN202510324745.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-19
AI Technical Summary
The prior art cannot provide intelligent screening and reduction during data acquisition, resulting in reduced efficiency and old data consumes a large amount of resources during calculation, reducing the accuracy of data processing.
By collecting the data platform feature information and data service feature information in the equipment update log database, performing data processing, establishing a data analysis model, performing logistic regression calculation, obtaining the screening evaluation coefficients, and comparing them with the preset screening threshold, determining the priority screening equipment data paragraphs, and conducting multi-dimensional comprehensive analysis to improve screening accuracy and data acquisition efficiency.
It improves the efficiency and accuracy of data collection, avoids waste of computing resources, and enhances user experience and service efficiency.
Smart Images

Figure CN120196616A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data acquisition, and more specifically, to a data acquisition method and system for a device service management platform. Background Art
[0002] With the rapid development of the Internet of Things technology, the demand for intelligent devices in various industries has increased year by year. As the core tool for realizing remote monitoring, operation and maintenance management, and intelligent decision-making of devices, the device service management platform has been widely used in fields such as industrial manufacturing, energy management, and smart cities. The core of the device service management platform lies in the real-time grasp of the operating status of distributed devices, which is inseparable from an efficient and accurate data acquisition function.
[0003] The existing technology has the following deficiencies:
[0004] Currently, data acquisition refers to obtaining operation parameters and status information from distributed devices and uploading them to the management platform to support subsequent analysis, storage, and decision-making. However, in actual applications, an effective intelligent reduction cannot be provided for the data acquisition of each device, thereby reducing the data acquisition efficiency. At the same time, some old data of device maintenance will consume more computing resources when substituted into comprehensive calculations, and at the same time reduce the accuracy of the platform's data processing. Therefore, a data acquisition method and system for a device service management platform are proposed.
[0005] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present disclosure, and thus it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0006] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a data acquisition method and system for a device service management platform, and solve the problems raised in the above background art by applying different product inspection methods.
[0007] To achieve the above object, the present invention provides the following technical solutions. A data acquisition method for a device service management platform includes:
[0008] S1: According to the device update log database, collect data platform feature information and data service feature information, perform data processing, and obtain the data call frequency, data fluctuation value, data validity period duration, and data area coverage range value;
[0009] S2: Obtain the data call frequency, data fluctuation value, data validity period duration, and data area coverage range value, establish a data analysis model, perform logistic regression calculation, and obtain a screening evaluation coefficient;
[0010] S3: Obtain the screening evaluation coefficient, compare it with the preset screening threshold, obtain the device data paragraphs to be preferentially screened based on the device data paragraph results marked as priority paragraphs in the comparison result, and then determine the data screening paragraphs according to the data similarity of each preferentially screened device data paragraph;
[0011] S4: Obtain the data screening paragraphs, and based on the data screening paragraphs, collect the real-time average upload speed of the user side for the call frequency of the data screening paragraphs and the frequency of data distribution of the data screening paragraphs in different regions;
[0012] S5: According to the real-time average upload speed of the user side for the call of the data screening paragraphs and the frequency of data distribution of the data screening paragraphs in different regions, use fuzzy logic to determine the call sorting result of the data screening paragraphs.
[0013] In a preferred embodiment, the data platform feature information includes the data call frequency and the data fluctuation value; the data service feature information includes the data validity period duration and the data region coverage range value;
[0014] By determining the analysis unit time, recording each data call event, within the set unit time, counting the total number of data calls that occur, and calculating the ratio of the total number of data calls within the unit time to the unit time length to obtain the data call frequency ; where i is the i-th unit time and g is the g-th data label;
[0015] Obtain the time series data of the operating parameters of different data labels from the device operation data record, and calculate the standard deviation of the operating parameter value relative to its mean value to obtain the data fluctuation value ;
[0016] Obtain the data validity period end time and the current time from the data information database, and subtract the current time from the data validity period end time to obtain the data validity period duration ;
[0017] Calculate by providing the longitude and latitude coordinates of each data label point to obtain the geographical location distribution coordinate area of the data label, and calculate the ratio with the area of the coverable area within the time range to obtain the data region coverage range value .
[0018] In a preferred embodiment, substitute the data call frequency, the data fluctuation value, the data validity period duration, and the data region coverage range value into the logical regression calculation. The specific formula is expressed as follows: ;
[0019] In the formula, For the calculation result of logistic regression, that is, the screening evaluation coefficient, e is the natural base, and y is the linear combination term of the logistic regression model. Specifically, y can be set as: ;
[0020] In the formula, is the bias term, , , and are the regression coefficients of the data call frequency, data fluctuation value, data validity period duration, and data area coverage range value respectively.
[0021] In a preferred embodiment, after obtaining the screening evaluation coefficient, the screening evaluation coefficient is compared and analyzed with the continuously iterated screening threshold;
[0022] If the screening evaluation coefficient is greater than or equal to the screening threshold, the current device data paragraph is marked as a priority paragraph and a screening signal is generated;
[0023] If the screening evaluation coefficient is less than the screening threshold, the current device data paragraph is marked as an excluded paragraph and an end signal is generated.
[0024] In a preferred embodiment, the current device data paragraph marked as a priority paragraph is recorded as the device data paragraph to be preferentially screened;
[0025] The values, distributions, and structures in the device data paragraph to be preferentially screened are vectorized, and the numerical similarity, time series similarity, and feature vector similarity are calculated through the cosine similarity formula.
[0026] In a preferred embodiment, the numerical similarity, time series similarity, and feature vector similarity are substituted into the weighting formula to calculate the data similarity of each device data paragraph to be preferentially screened;
[0027] The data similarity of the device data paragraph to be preferentially screened is compared with the preset similarity threshold. If the data similarity of the device data paragraph to be preferentially screened is greater than the similarity threshold, the device data paragraph with the corresponding comparison similarity is excluded; otherwise, the device data paragraph with the corresponding comparison similarity is retained until the comparison ends;
[0028] The device data paragraphs retained in the comparison result are collected to obtain the data screening paragraphs.
[0029] In a preferred embodiment, the data screening paragraphs contain multiple data tags, and each data tag contains multiple data;
[0030] By obtaining the number of users called within the current data screening paragraph, accumulating the real-time upload speeds of all called users, and calculating the ratio with the number of users called within the current data screening paragraph, the real-time average upload speed of the data user side for the data screening paragraph is obtained;
[0031] Extract the geographical location data corresponding to the data screening paragraph from the records of the data source, calculate the ratio of the data quantities corresponding to different locations to the total data quantity of different data screening paragraphs, and then accumulate the data distribution frequencies within each region to obtain the data distribution frequency of the data screening paragraph in different regions.
[0032] In a preferred embodiment, the real-time average upload speed of the data user side for the data screening paragraph and the data distribution frequency of the data screening paragraph in different regions are respectively divided into different fuzzy sets;
[0033] Define the sorting result of the data screening paragraph call as the output variable and divide it into a fuzzy set;
[0034] Formulate fuzzy rules to describe the influence of the real-time average upload speed of the data user side for the data screening paragraph and the data distribution frequency of the data screening paragraph in different regions on the sorting result of the data screening paragraph call;
[0035] Perform fuzzy reasoning according to the fuzzy rules to determine the sorting result of the data screening paragraph call.
[0036] A data acquisition system for a device service management platform includes a data acquisition module, a data processing module, a screening and analysis module, and a paragraph sorting module;
[0037] The data acquisition module is used to collect data platform feature information and data service feature information based on the device update log database, perform data processing to obtain the data call frequency, data fluctuation value, data validity period duration, and data region coverage range value, and send them to the data processing module;
[0038] The data processing module is used to obtain the data call frequency, data fluctuation value, data validity period duration, and data region coverage range value, establish a data analysis model, perform logistic regression calculation to obtain a screening evaluation coefficient, and send it to the screening and analysis module;
[0039] The screening and analysis module is used to obtain the screening evaluation coefficient, compare it with a preset screening threshold, obtain the device data paragraphs to be preferentially screened based on the device data paragraph results marked as priority paragraphs in the comparison result, and then determine the data screening paragraph according to the data similarity of each preferentially screened device data paragraph, and send it to the paragraph sorting module;
[0040] The paragraph sorting module is used to obtain data screening paragraphs. Based on the data screening paragraphs, it collects the call frequency of the data screening paragraphs, the real-time average upload speed of the user side, and the data distribution frequency of the data screening paragraphs in different regions, and uses fuzzy logic to determine the call sorting result of the data screening paragraphs.
[0041] Technical effects and advantages of the present invention:
[0042] 1. By collecting the data call frequency, data fluctuation value, data validity period duration, and data area coverage value, the present invention establishes a data analysis model, performs logistic regression calculation, obtains a screening evaluation coefficient, compares it with a preset screening threshold, and based on the comparison result, obtains the device data paragraphs for priority screening. Then, according to the data similarity of each device data paragraph for priority screening, it determines the data screening paragraphs, conducts multi-dimensional comprehensive analysis, improves the screening accuracy, enhances the data collection efficiency, and avoids waste of computing resources.
[0043] 2. By obtaining the data screening paragraphs, based on the data screening paragraphs, it collects the call frequency of the data screening paragraphs, the real-time average upload speed of the user side, and the data distribution frequency of the data screening paragraphs in different regions, formulates a set of fuzzy rules for fuzzy reasoning, determines the call sorting result of the data screening paragraphs, reduces the computational pressure of subsequent data analysis, improves the service efficiency, and enhances the user experience. Brief Description of the Drawings
[0044] Figure 1 It is a method flow chart of a data collection method for a device service management platform according to the present invention.
[0045] Figure 2 It is a module schematic diagram of a data collection system for a device service management platform according to the present invention. Detailed Embodiments
[0046] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0047] The present invention collects device data feature information and device service feature information based on the device update log database, analyzes from two aspects of data quality and service period characteristics, performs logistic regression calculation through analyzing the data analysis model, outputs to obtain a screening evaluation coefficient, determines the data screening paragraphs through the screening evaluation coefficient, and determines the data collection result through the fuzzy Bayesian model according to the data importance score and data selection frequency within the paragraphs analyzed based on the data screening paragraphs;
[0048] The device service period is a quality guarantee period after the user owns the device. During this quality guarantee period, when the customer scans the code and needs to upload the services they require, data is collected. The collected data needs to analyze in advance the scope collected from the database.
[0049] Embodiment 1
[0050] Please refer to Figure 1 , a data collection method for a device service management platform, and the specific operation process is as follows:
[0051] S1: According to the device update log database, collect data platform feature information and data service feature information, perform data processing to obtain data call frequency, data fluctuation value, data validity period duration, and data area coverage range value;
[0052] Among them, data processing includes data type conversion, missing value processing, data standardization, and feature extraction operations;
[0053] Among them, data type conversion refers to converting different types of data (such as strings, dates, numerical values, etc.) according to the source format of the data and the target application requirements. For example, converting the numerical value saved in the form of a string in the device log or data record to a floating-point type or integer type data, etc.;
[0054] Missing value processing is used to solve the problem of null values in device logs or data records caused by device failures, network transmission problems, or incomplete records. The deletion method or mean filling method is preferably used to process the missing values;
[0055] Data standardization standardizes each piece of collected data according to a unified scale, and normalizes the data call frequency, data fluctuation value, data validity period duration, and data area coverage range value;
[0056] Specifically, the above data operation methods are all existing technologies and will not be elaborated here;
[0057] Among them, the data platform feature information includes data call frequency and data fluctuation value; the data service feature information includes data validity period duration and data area coverage range value;
[0058] The data call frequency refers to the number of times data is called within a set unit time, which measures the data service response ability, stability, and efficiency. Its acquisition logic is to determine the unit time for analysis, record each data call event, and within the set unit time, count the total number of data calls that occur. Calculate the ratio of the total number of data calls within the unit time to the unit time length to obtain the data call frequency ; where i is the i-th unit time and g is the g-th data label;
[0059] It should be noted that the setting of the unit time can be the lengths of "24 hours", "36 hours", and "72 hours". The specific length setting is determined by the experimenter based on the historical data call frequency and the historical service period length, which will not be elaborated here;
[0060] It should be noted that the data label means that the platform pre-classifies all the data in the database in advance, divides them into multiple groups, and attaches corresponding data labels to them. The specific format of the data label is not limited;
[0061] The data fluctuation value refers to the change range of the operating parameters (such as temperature, pressure, current, etc.) of the device within the unit time, which is used to evaluate the screenability of the device data. Its acquisition logic is to obtain the time series data of the operating parameters of different data labels from the device operation data record, and calculate the standard deviation of the operating parameter value relative to its mean value to obtain the data fluctuation value ;
[0062] Among them, the device service management platform preferentially divides the data into multiple labels and labels them as g, that is, g is the gth data label. The specific division rules can be based on the classification of the original data by the experimenter or the classification according to the similarity between the data. The specific classification method is not limited, but is obtained by the experimenter according to the specific implementation method, which will not be elaborated here;
[0063] Specifically, the above-mentioned operating parameters are not limited and can be temperature, pressure, current, etc. In this embodiment, only the most representative operating parameters are selected. For example, for a refrigeration device, the fluctuation value of its refrigeration efficiency data is obtained, and for a heating device, the fluctuation value of its heating surface temperature is obtained, etc. The selection of the specific operating parameters is set by the experimenter according to the specific device application scenario and the device operation characteristics, which will not be elaborated here;
[0064] Among them, the formula for calculating the standard deviation of the operating parameter value relative to its mean value is expressed as follows: ;
[0065] In the formula, represents the standard deviation of the gth type of data in the ith unit time, that is, the data fluctuation value, P is the total number of data labels, is the operating parameter value of the gth type of data, is the mean value of the operating parameter;
[0066] The data validity period duration refers to the remaining time from the current moment to the end of the data validity period. It is an important parameter for measuring the current life cycle state of the data, used to predict the screenability of the data and the validity period duration. Its acquisition logic is to obtain the end time of the data validity period and the current time from the data information database, and subtract the current time from the end time of the data validity period to get the data validity period duration. ;
[0067] Specifically, when selecting the data validity period duration system, it defaults to comparing the end time of the data validity period with the current time. If the end time of the data validity period is greater than the current time, it means that the current data is within the validity period. Otherwise, it means that the end time of the data validity period has passed, and the data will be deleted.
[0068] The data area coverage range value refers to the geographical area range that data with different data tags can cover within the corresponding unit time. It is usually used to measure the availability and coverage of the data platform in different geographical locations. Its acquisition logic is to calculate by providing the latitude and longitude coordinates of each data tag point, obtain the geographical location distribution coordinate area of the data tag, and calculate the ratio with the area of the coverable area within the time range to get the data area coverage range value. ;
[0069] Among them, the latitude and longitude coordinates of each data tag point refer to the specific geographical location associated with each data tag, obtaining the corresponding latitude and longitude points. By calculating the minimum rectangular boundary of these points, the area of the data coverage range is calculated. The range of the bounding box is determined by the maximum and minimum latitude and longitude values, forming a rectangular area. The specific area calculation can use spherical trigonometric formulas, etc., which will not be elaborated here.
[0070] S2: Obtain the data call frequency, data fluctuation value, data validity period duration, and data area coverage range value, establish a data analysis model, perform logistic regression calculation, and obtain the screening evaluation coefficient.
[0071] Among them, the data analysis model refers to the logistic regression calculation model, and the screening evaluation coefficient is generated through logistic regression calculation.
[0072] Substitute the data call frequency, data fluctuation value, data validity period duration, and data area coverage range value into the specific formula of the logistic regression calculation as follows: ;
[0073] In the formula, is the result of the logistic regression calculation, that is, the screening evaluation coefficient. e is the natural base, and y is the linear combination term of the logistic regression model. Specifically, y can be set as: ;
[0074] In the formula, is the bias term, , , as well as They are the regression coefficients of data call frequency, data fluctuation value, data validity period, and data area coverage value;
[0075] Among them, data call frequency, data fluctuation value, data validity period and data area coverage value are all digital manifestations that directly express the current device data segment screening selection;
[0076] It can be seen from the formula that the higher the data call frequency, data fluctuation value and data area coverage value, the more valuable the current device data segment is in terms of demand, dynamics, applicability, etc., and needs to be screened first. The higher the screening evaluation coefficient, the higher the data validity period. Conversely, the higher the data validity period, the more time-effective the current device data segment is, the lower the priority, and the lower the screening evaluation coefficient.
[0077] S3: Obtaining a screening evaluation coefficient and comparing it with a preset screening threshold, obtaining a device data segment for priority screening according to the device data segment results marked as priority segments in the comparison results, and then determining a data screening segment according to the data similarity of each device data segment for priority screening;
[0078] The logic for obtaining the screening threshold is to collect a priority classification set of historical device data segments, divide the data set into a training set and a test set, set the evaluation index and clustering algorithm, and in each round of cross-validation, train the model on the training set and evaluate the model performance on the test set. Then, adjust the screening threshold based on the performance of the validation set. Therefore, the screening threshold is constantly updated.
[0079] In the present invention, clustering algorithm is a kind of unsupervised learning algorithm, which is used to divide the equipment data segment priorities in the data set into marked groups or clusters; the common one is K-means clustering, which divides the weighing data in the data set into K clusters, so that the distance between each equipment data segment priority and the center point (center of mass) of the cluster to which it belongs is minimized, and finally the priority of the equipment data segment distribution is measured by the Euclidean distance, so as to set the screening threshold;
[0080] After obtaining the screening evaluation coefficient, the screening evaluation coefficient is compared and analyzed with the continuously iterated screening threshold;
[0081] If the screening evaluation coefficient is greater than or equal to the screening threshold, the current device data segment is marked as a priority segment and a screening signal is generated;
[0082] If the screening evaluation coefficient is less than the screening threshold, mark the current device data paragraph as the screened paragraph and generate an end signal;
[0083] Mark the current device data paragraph marked as the priority paragraph as the device data paragraph to be preferentially screened;
[0084] Vectorize the values, distributions, and structures in the device data paragraph to be preferentially screened, and calculate the numerical similarity, time series similarity, and feature vector similarity through the cosine similarity formula;
[0085] Specifically, this example analyzes the numerical similarity, time series similarity, and feature vector similarity in three dimensions. In fact, the experimenter can set a relatively dense data similarity according to the actual application to more accurately express the data similarity of each device data paragraph to be preferentially screened, and perform operations to improve the screening accuracy, etc., which will not be elaborated here;
[0086] It should be noted that in this example, the cosine similarity is used to calculate the similarity formula. However, in actual applications, the Euclidean distance can also be used to determine the numerical similarity, etc. For the time series similarity, it can also be determined according to the dynamic time warping method, etc. The method of calculating the similarity formula here is not limited, but is set according to the calculation model preset by the experimenter, which will not be elaborated here;
[0087] Substitute the numerical similarity, time series similarity, and feature vector similarity into the weighted formula to calculate the data similarity of each device data paragraph to be preferentially screened;
[0088] Compare the similarity of the device data paragraph to be preferentially screened with the preset similarity threshold. If the similarity of the device data paragraph to be preferentially screened is greater than the similarity threshold, screen out the device data paragraph with the corresponding comparison similarity. Otherwise, retain the device data paragraph with the corresponding comparison similarity until the comparison ends;
[0089] Collect the device data paragraphs retained in the comparison result to obtain the data screening paragraphs;
[0090] The present invention collects the data call frequency, data fluctuation value, data validity period duration, and data area coverage range value, establishes a data analysis model, performs logistic regression calculation to obtain the screening evaluation coefficient, compares it with the preset screening threshold, and based on the comparison result, obtains the device data paragraphs to be preferentially screened. Then, according to the data similarity of each device data paragraph to be preferentially screened, determines the data screening paragraphs, conducts multi-dimensional comprehensive analysis, improves the screening accuracy, improves the data collection efficiency, and avoids waste of computing resources.
[0091] Embodiment 2
[0092] In Embodiment 1 of the present invention, the call frequency of the collected data, the data fluctuation value, the data validity period duration, and the data area coverage range value are mainly exemplified. A data analysis model is established, logistic regression calculation is performed to obtain a screening evaluation coefficient, and it is compared with a preset screening threshold. According to the comparison result, the device data paragraphs to be preferentially screened are obtained. Then, according to the data similarity of each preferentially screened device data paragraph, the operation strategy of the data screening paragraph is determined. However, in Embodiment 1, only the data screening paragraph is obtained, which provides an accurate method for data screening. However, the calls of the data screening paragraph are not sorted. Obviously, this will cause a large amount of data to still be collected after the user scans the code, causing a computational pressure on subsequent data analysis, delaying the service efficiency, and reducing the user experience. To address the above problems, Embodiment 2 of the present invention is further refined;
[0093] S4: Obtain the data screening paragraph. According to the data screening paragraph, collect the call frequency of the data screening paragraph, the real-time average upload speed of the user side, and the frequency of data distribution of the data screening paragraph in different regions;
[0094] Specifically, the data screening paragraph contains multiple data tags, and each data tag contains multiple data;
[0095] The acquisition logic of the real-time average upload speed of the user side for the data called by the data screening paragraph is to obtain the number of users calling within the current data screening paragraph, accumulate the real-time upload speeds of all calling users, and calculate the ratio with the number of users calling within the current data screening paragraph to obtain the real-time average upload speed of the user side for the data called by the data screening paragraph;
[0096] Among them, when the user scans the code or performs other interaction operations, the client application will monitor the upload rate in real time. Specifically, during the implementation of the upload module on the client side, the upload speed is collected regularly, for example, the upload volume and the time taken are recorded once per second, and the real-time upload speed is stored in the local cache. The specific real-time acquisition method is not limited here and will not be elaborated;
[0097] The frequency of data distribution of the data screening paragraph in different regions reflects the activity of the data screening paragraph in terms of geographical location. The acquisition logic is to extract the geographical location data corresponding to the data screening paragraph from the records of the data source, calculate the ratio of the data quantity corresponding to different locations to the total data quantity of different data screening paragraphs, and then accumulate the data distribution frequencies in each region to obtain the frequency of data distribution of the data screening paragraph in different regions;
[0098] It should be noted that for the geographical location selection and quantity, by analyzing the activity of users in each geographical location, geographical locations with a larger number of users or higher activity can be preferentially selected, etc., which will not be elaborated here;
[0099] S5: According to the data screening paragraphs, call the real-time average upload speed of the data user side and the frequency of data distribution of the data screening paragraphs in different regions, and use fuzzy logic to determine the sorting result of the data screening paragraph calls;
[0100] For example, "Fast", "Slow", "Moderate" for the real-time average upload speed of the data user side called by the data screening paragraphs, and "High", "Low", "Medium" for the frequency of data distribution of the data screening paragraphs in different regions;
[0101] Formulate a set of fuzzy rules to describe the influence of different input variables on the output variable. The definition of the rules can be based on professional knowledge or obtained through data analysis and experiments. For example:
[0102] Mark the real-time average upload speed of the data user side called by the data screening paragraphs as X, the frequency of data distribution of the data screening paragraphs in different regions as U, and the sorting result of the data screening paragraph calls as C_results;
[0103] Then it can be defined as: Rule 1: IF (X is Fast) AND (U is High) THEN (C_results is High) Rule 2: IF (U is Slow) AND (U is Low) THEN (C_results is Low) ...
[0104] Perform fuzzy inference according to the fuzzy rules to determine the sorting result of the data screening paragraph calls;
[0105] It should be noted that the division of fuzzy sets can be adjusted according to the actual situation. For example, although this embodiment takes three fuzzy sets as an example, in fact, the real-time average upload speed of the data user side called by the data screening paragraphs and the frequency of data distribution of the data screening paragraphs in different regions can be divided into more than three sets to facilitate more accurate adjustment according to different data labels;
[0106] Furthermore, for the judgment of high, medium, and low of the real-time average upload speed of the data user side called by the data screening paragraphs and the frequency of data distribution of the data screening paragraphs in different regions, thresholds can be set according to the actual situation for judgment. For example, when the real-time average upload speed of the data user side called by the data screening paragraphs exceeds 75%, it is labeled as "Fast", and when the frequency of data distribution of the data screening paragraphs in different regions is higher than 70%, it is labeled as "High", etc., which will not be elaborated here;
[0107] The present invention obtains data screening paragraphs, and based on the data screening paragraphs, collects the call frequency of the data screening paragraphs, the real-time average upload speed of the user side, and the data distribution frequency of the data screening paragraphs in different regions to formulate a set of fuzzy rules for fuzzy inference, determines the call sorting result of the data screening paragraphs, reduces the computational pressure of subsequent data analysis, improves service efficiency, and enhances the user experience.
[0108] Embodiment 3
[0109] Please refer to Figure 2 , a data acquisition system for a device service management platform, including a data acquisition module, a data processing module, a screening and analysis module, and a paragraph sorting module;
[0110] The data acquisition module is used to collect data platform feature information and data service feature information based on the device update log database, perform data processing to obtain data call frequency, data fluctuation value, data validity period duration, and data area coverage range value, and send them to the data processing module;
[0111] The data processing module is used to obtain the data call frequency, data fluctuation value, data validity period duration, and data area coverage range value, establish a data analysis model, perform logistic regression calculation to obtain a screening evaluation coefficient, and send it to the screening and analysis module;
[0112] The screening and analysis module is used to obtain the screening evaluation coefficient, compare it with a preset screening threshold, obtain the device data paragraphs to be preferentially screened based on the device data paragraph results marked as priority paragraphs in the comparison result, and then determine the data screening paragraphs according to the data similarity of each preferentially screened device data paragraph, and send them to the paragraph sorting module;
[0113] The paragraph sorting module is used to obtain the data screening paragraphs, and based on the data screening paragraphs, collect the call frequency of the data screening paragraphs, the real-time average upload speed of the user side, and the data distribution frequency of the data screening paragraphs in different regions, and use fuzzy logic to determine the call sorting result of the data screening paragraphs.
[0114] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the real situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0115] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0116] It should be understood that in various embodiments of the present application, the order numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0117] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0118] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0119] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0120] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0121] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0122] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0123] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A data collection method for a device service management platform, characterized in that: include: S1: According to the device update log database, collect data platform characteristic information and data service characteristic information, perform data processing, and obtain data call frequency, data fluctuation value, data validity period and data area coverage value; S2: Obtain data call frequency, data fluctuation value, data validity period, and data area coverage value, establish a data analysis model, perform logistic regression calculation, and obtain the screening evaluation coefficient; S3: Obtaining a screening evaluation coefficient and comparing it with a preset screening threshold, obtaining a device data segment for priority screening according to the device data segment results marked as priority segments in the comparison results, and then determining a data screening segment according to the data similarity of each device data segment for priority screening; S4: Obtain the data screening section, and according to the data screening section, collect the real-time average upload speed of the user end of the data screening section call frequency and the frequency of data distribution of the data screening section in different regions; S5: According to the real-time average upload speed of the data filtering section calling data user end and the frequency of data distribution in different regions of the data filtering section, the fuzzy logic is used to determine the data filtering section calling sorting result.
2. A data collection method for a device service management platform according to claim 1, characterized in that: Data platform characteristic information includes data call frequency and data fluctuation value; data service characteristic information includes data validity period and data area coverage value; By determining the unit time for analysis, recording each data call event, and calculating the total number of data calls within the set unit time, the data call frequency is obtained by calculating the ratio of the total number of data calls within the unit time to the unit time length. ; Where i is the i-th unit time, and g is the g-th data label; Obtain the time series data of operating parameters with different data tags from the equipment operation data records, and obtain the data fluctuation value by calculating the standard deviation of the operating parameter value relative to its mean ; Obtain the data validity period end time and current time from the data information database, and subtract the current time from the data validity period end time to get the data validity period length ; By providing the data latitude and longitude coordinates of each data label point for calculation, the geographical location distribution coordinate area of the data label is obtained, and the ratio calculation is performed with the coverable area within the time range to obtain the data area coverage value. .
3. A data collection method for a device service management platform according to claim 2, characterized in that: Substituting the data call frequency, data fluctuation value, data validity period and data area coverage value into the logistic regression calculation, the specific formula is expressed as follows: ; In the formula, is the result of logistic regression calculation, that is, the screening evaluation coefficient, e is the natural base, y is the linear combination term of the logistic regression model, and the specific y is set as: ; In the formula, is the bias term, , , as well as They are the regression coefficients of data call frequency, data fluctuation value, data validity period and data area coverage value.
4. The data collection method for a device service management platform according to claim 3, characterized in that: After obtaining the screening evaluation coefficient, the screening evaluation coefficient is compared and analyzed with the continuously iterated screening threshold; If the screening evaluation coefficient is greater than or equal to the screening threshold, the current device data segment is marked as a priority segment and a screening signal is generated; If the screening evaluation coefficient is less than the screening threshold, the current device data segment is marked as a screened-out segment and an end signal is generated.
5. A data collection method for a device service management platform according to claim 4, characterized in that: The current device data section marked as a priority section is recorded as the device data section for priority screening; The three dimensions of value, distribution and structure in the prioritized device data segments are vectorized, and the value similarity, time series similarity and feature vector similarity are calculated using the cosine similarity formula.
6. A data collection method for a device service management platform according to claim 5, characterized in that: Substitute the numerical similarity, time series similarity and feature vector similarity into the weighted formula to calculate the data similarity of each device data segment that is prioritized for screening; Compare the data similarity of the device data segments that are prioritized for screening with the preset similarity threshold. If the data similarity of the device data segments that are prioritized for screening is greater than the similarity threshold, the device data segments with the corresponding comparison similarity are screened out. Otherwise, the device data segments with the corresponding comparison similarity are retained until the comparison is completed. The device data segments retained by the comparison results are collected to obtain data screening segments.
7. A data collection method for a device service management platform according to claim 6, characterized in that: The data screening paragraph contains multiple data labels, and the data labels contain multiple data; By obtaining the number of calling users in the current data screening section, the real-time upload speeds of all calling users are accumulated, and the ratio is calculated with the number of calling users in the current data screening section to obtain the real-time average upload speed of the data user end calling the data screening section; The geographic location data corresponding to the data filtering paragraphs are extracted from the records of the data source, and the ratio of the number of data corresponding to different locations to the total amount of data in different data filtering paragraphs is calculated. Then, the data distribution frequency in each area is counted and accumulated to obtain the frequency of data distribution of the data filtering paragraphs in different areas.
8. The data collection method for a device service management platform according to claim 7, characterized in that: The real-time average upload speed of the data user end calling the data screening section and the frequency of data distribution in different regions of the data screening section are divided into different fuzzy sets; Define the sorting results of the data screening paragraph call as output variables and divide them into fuzzy sets; Formulate fuzzy rules to describe the impact of the real-time average upload speed of the data filtering section calling data user end and the frequency of data distribution in different regions of the data filtering section on the data filtering section calling sorting results; Perform fuzzy reasoning based on fuzzy rules to determine the data screening paragraph call sorting results.
9. A data collection system for a device service management platform, used to implement a data collection method for a device service management platform according to any one of claims 1 to 8, characterized in that: It includes data collection module, data processing module, screening and analysis module and paragraph sorting module; The data collection module is used to collect data platform characteristic information and data service characteristic information based on the device update log database, perform data processing, obtain data call frequency, data fluctuation value, data validity period and data area coverage value, and send them to the data processing module; The data processing module is used to obtain data call frequency, data fluctuation value, data validity period and data area coverage value, establish a data analysis model, perform logistic regression calculation, obtain the screening evaluation coefficient, and send it to the screening analysis module; The screening analysis module is used to obtain the screening evaluation coefficient and compare it with the preset screening threshold value, obtain the device data paragraphs with priority screening according to the device data paragraph results marked as priority paragraphs in the comparison results, and then determine the data screening paragraphs according to the data similarity of each device data paragraph with priority screening, and send them to the paragraph sorting module; The paragraph sorting module is used to obtain data filtering paragraphs. According to the data filtering paragraphs, the real-time average upload speed of the user end of the data filtering paragraph call frequency and the frequency of data distribution of the data filtering paragraphs in different regions are collected, and fuzzy logic is used to determine the data filtering paragraph call sorting results.
Citation Information
Patent Citations
Effective data screening method, readable storage medium and terminal
CN109542927A
Data processing system suitable for accounting financial management
CN119088830A
Customer behavior analysis method and system based on retail data
CN119624511A
Traditional Chinese medicinal material radix angelicae pubescentis quality evaluation method based on data analysis
CN119648042A
Calculation device, calculation method, and calculation program
JP2017054554A