A data collection method and system for a device service management platform
By establishing a data analysis model in the equipment service management platform, performing logistic regression calculations and fuzzy logic sorting, the problems of low data collection efficiency and wasted computing resources are solved, and efficient and accurate data filtering and sorting are achieved.
Patent Information
- Application Number
- CN202510324745.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-03-19
AI Technical Summary
In existing technologies, the data collection of equipment service management platforms cannot be effectively filtered, resulting in low data collection efficiency, wasted computing resources, and reduced accuracy.
By collecting data platform and service feature information from the device update log database, a data analysis model is established, logistic regression calculation is performed, screening evaluation coefficients are obtained, and compared with preset thresholds. Combined with data similarity and fuzzy logic, data screening segments are determined, and multi-dimensional analysis and ranking are performed.
It improves the accuracy and efficiency of data collection and filtering, reduces the waste of computing resources, and enhances the user experience.
Smart Images

Figure CN120196616B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data collection, more particularly, the present application relates to a data collection method and system for a device service management platform. BACKGROUND
[0002] With the rapid development of the Internet of Things technology, the demand for intelligent devices in various industries is increasing year by year. As a core tool for realizing remote monitoring, operation and maintenance management, and intelligent decision-making of devices, the device service management platform has been widely used in industrial manufacturing, energy management, smart city and other fields. The core of the device service management platform is to grasp the real-time running state of distributed devices, which cannot be achieved without efficient and accurate data collection function.
[0003] The prior art has the following disadvantages:
[0004] Currently, data collection refers to obtaining running parameters and state information from distributed devices and uploading them to the management platform to support subsequent analysis, storage and decision-making. However, in actual application, the data collection of each device cannot provide an effective intelligent screening, thereby reducing the data collection efficiency. At the same time, some old data of device maintenance will consume more computing resources when used in comprehensive calculation, which reduces the accuracy of platform data processing. Therefore, a data collection method and system for a device service management platform are proposed.
[0005] The above information disclosed in the background section is only used to enhance the understanding of the background of the present disclosure, and therefore it can include information that does not constitute prior art known to those of ordinary skill in the art. SUMMARY
[0006] In order to overcome the above-mentioned defects of the prior art, the embodiments of the present application provide a data collection method and system for a device service management platform, which solves the problems raised in the above background technology by using different product inspection methods.
[0007] To achieve the above-mentioned purpose, the present application provides the following technical scheme, a data collection method for a device service management platform, comprising:
[0008] S1: According to the equipment update log database, collect data platform feature information and data service feature information, and perform data processing to obtain data call frequency, data fluctuation value, data validity period length and data area coverage range value;
[0009] S2: Obtain the data call frequency, data fluctuation value, data validity period length and data area coverage range value, establish a data analysis model, perform logistic regression calculation, and obtain a screening evaluation coefficient;
[0010] S3: Obtain the screening evaluation coefficient and compare it with the preset screening threshold, obtain the device data paragraph result marked as a priority paragraph according to the comparison result, and determine the data screening paragraph according to the data similarity of each priority screening device data paragraph;
[0011] S4: Obtain the data screening paragraph, collect the data screening paragraph calling frequency user end real-time average upload speed and the data screening paragraph in different area data distribution frequency according to the data screening paragraph;
[0012] S5: According to the data screening paragraph calling data user end real-time average upload speed and the data screening paragraph in different area data distribution frequency, use fuzzy logic to determine the data screening paragraph calling sorting result.
[0013] In a preferred embodiment, the data platform feature information includes data calling frequency and data fluctuation value; the data service feature information includes data validity period length and data area coverage value;
[0014] By determining the unit time of analysis, record the event of each data call, in the set unit time, count the total number of data call occurrences, calculate the ratio of the total number of data calls in the unit time to the unit time length, and obtain the data calling frequency ; Wherein, i is the i th unit time, g is the g th data label;
[0015] Obtain the running parameter time series data of different data labels from the device running data record, calculate the standard deviation of the running parameter value relative to its mean value, and obtain the data fluctuation value ;
[0016] Obtain the data validity period end time and the current time from the data information database, subtract the current time from the data validity period end time to obtain the data validity period length ;
[0017] By providing the data latitude and longitude coordinates of each data label point, the geographical position distribution coordinate area of the data label is obtained, and the ratio of the area to the coverable area in the time range is calculated to obtain the data area coverage value .
[0018] In a preferred embodiment, the data calling frequency, data fluctuation value, data validity period length and data area coverage value are substituted into the logistic regression calculation specific formula as follows:
[0019] ;
[0020] In the formula, The result of the logistic regression is calculated, that is, the screening evaluation coefficient, e is the natural base, y is the linear combination term of the logistic regression model, and y can be set as:
[0021] ;
[0022] In the formula, is the bias term, , , and are the regression coefficients of the data call frequency, the data fluctuation value, the data validity period length and the data area coverage value respectively.
[0023] In a preferred embodiment, after obtaining the screening evaluation coefficient, the screening evaluation coefficient is compared and analyzed with the constantly iterated screening threshold value;
[0024] If the screening evaluation coefficient is greater than or equal to the screening threshold value, the current device data paragraph is marked as a priority paragraph, and a screening signal is generated;
[0025] If the screening evaluation coefficient is less than the screening threshold value, the current device data paragraph is marked as a screened-out paragraph, and an end signal is generated.
[0026] In a preferred embodiment, the current device data paragraph marked as a priority paragraph is recorded as a priority screened device data paragraph;
[0027] The three dimensions of value, distribution and structure in the priority screened device data paragraph are vectorized, and the value similarity, time series similarity and feature vector similarity are calculated by the cosine similarity formula.
[0028] In a preferred embodiment, the value similarity, time series similarity and feature vector similarity are substituted into the weighted formula to calculate the data similarity of each priority screened device data paragraph;
[0029] The data similarity of the priority screened device data paragraph is compared with the preset similarity threshold value, if the data similarity of the priority screened device data paragraph is greater than the similarity threshold value, the corresponding device data paragraph with the compared similarity is screened out, otherwise, the corresponding device data paragraph with the compared similarity is retained, until the comparison ends;
[0030] The device data paragraphs retained by the comparison result are collected to obtain a data screening paragraph.
[0031] In a preferred embodiment, the data screening paragraph contains multiple data tags, and the data tags contain multiple data.
[0032] The number of called users in the current data screening paragraph is obtained, the real-time upload speeds of all called users are accumulated, and the real-time average upload speed of the data user end of the data screening paragraph is obtained by ratio calculation of the number of called users in the current data screening paragraph.
[0033] The geographical position data corresponding to the data screening paragraph is extracted from the records of the data source, the data quantity corresponding to different positions is ratio calculated with the total data quantity of different data screening paragraphs, and the data distribution frequency in each region is accumulated to obtain the data distribution frequency of the data screening paragraph in different regions.
[0034] In a preferred embodiment, the real-time average upload speed of the data user end of the data screening paragraph and the data distribution frequency of the data screening paragraph in different regions are respectively divided into different fuzzy sets.
[0035] The data screening paragraph calling sorting result is defined as an output variable and divided into a fuzzy set.
[0036] Fuzzy rules are formulated to describe the influence of the real-time average upload speed of the data user end of the data screening paragraph and the data distribution frequency of the data screening paragraph in different regions on the data screening paragraph calling sorting result.
[0037] According to the fuzzy rules, the data screening paragraph calling sorting result is determined by fuzzy reasoning.
[0038] A data collection system for a device service management platform includes a data collection module, a data processing module, a screening analysis module, and a paragraph sorting module.
[0039] The data collection module is used to collect data platform feature information and data service feature information according to the device update log database, perform data processing, obtain data calling frequency, data fluctuation value, data effective period length, and data region coverage range value, and send them to the data processing module.
[0040] The data processing module is used to obtain the data calling frequency, data fluctuation value, data effective period length, and data region coverage range value, establish a data analysis model, perform logistic regression calculation, obtain a screening evaluation coefficient, and send it to the screening analysis module.
[0041] The screening analysis module is used to obtain the screening evaluation coefficient and compare it with a preset screening threshold, obtain the device data paragraph result marked as a priority paragraph according to the comparison result, determine the data screening paragraph according to the data similarity of each priority screened device data paragraph, and send it to the paragraph sorting module.
[0042] The paragraph ordering module is used for obtaining a data screening paragraph, collecting a data screening paragraph calling frequency user end real-time average upload speed and a data screening paragraph in different area data distribution frequency, and using fuzzy logic to determine a data screening paragraph calling ordering result.
[0043] Technical effects and advantages of the present application:
[0044] 1. The present application collects data calling frequency, data fluctuation value, data validity period length and data area coverage value, establishes a data analysis model, performs logic regression calculation, obtains a screening evaluation coefficient, compares with a preset screening threshold value, obtains a prior screening device data paragraph according to the comparison result, and determines a data screening paragraph according to the data similarity of each prior screening device data paragraph, performs multi-dimensional comprehensive analysis, improves screening accuracy, improves data collection efficiency, and avoids waste of computing resources.
[0045] 2. The present application obtains a data screening paragraph, collects a data screening paragraph calling frequency user end real-time average upload speed and a data screening paragraph in different area data distribution frequency, formulates a group of fuzzy rules for fuzzy reasoning, determines a data screening paragraph calling ordering result, reduces the calculation pressure of subsequent data analysis, improves service efficiency, and enhances user experience. BRIEF DESCRIPTION OF DRAWINGS
[0046] Figure 1 The present application is a method flow chart of a device service management platform-oriented data collection method.
[0047] Figure 2 The present application is a module schematic diagram of a device service management platform-oriented data collection system. DETAILED DESCRIPTION
[0048] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0049] The present application collects device data feature information and device service feature information according to a device update log database, analyzes from two aspects of data quality and service period characteristics, performs logic regression calculation through data analysis model analysis, outputs a screening evaluation coefficient, determines a data screening paragraph through the screening evaluation coefficient, and determines a data collection result through a fuzzy Bayesian model according to data importance score and data selection frequency in the data screening paragraph analysis paragraph.
[0050] The device service period belongs to a warranty period after the user owns the device, and in this warranty period, data collection is performed when the customer scans the code and needs to upload the required service, and the collected data needs to be analyzed in advance to collect the range from the database;
[0051] Embodiment 1
[0052] Please refer to Figure 1 A data collection method for a device service management platform, the specific operation process is as follows:
[0053] S1: According to the device update log database, collect data platform feature information and data service feature information, and perform data processing to obtain data call frequency, data fluctuation value, data validity period length and data area coverage range value;
[0054] Among them, the data processing includes data type conversion, missing value processing, data standardization, and feature extraction operation;
[0055] Among them, data type conversion refers to converting different types of data (such as strings, dates, numerical values, etc.) according to the source format of the data and the target application requirements, for example, converting device logs or data records in string form to floating-point or integer data;
[0056] The missing value processing is used to solve the null value problem caused by device failure, network transmission problem or incomplete record in the device log or data record, and the missing value is processed by using the deletion method or the mean filling method;
[0057] Data standardization standardizes each item of collected data according to a unified scale, and normalizes data call frequency, data fluctuation value, data validity period length and data area coverage range value;
[0058] Specifically, the above data operation method is prior art, which will not be described here;
[0059] Among them, the data platform feature information includes data call frequency and data fluctuation value; the data service feature information includes data validity period length and data area coverage range value;
[0060] Data call frequency refers to the number of data calls in a set unit of time, which measures the data service response ability, stability and efficiency, and its acquisition logic is to determine the analysis unit of time, record each data call event, and in a set unit of time, count the total number of data call events. The total number of data calls in a unit of time is compared with the unit of time to obtain the data call frequency ; Wherein i is the ith unit of time, and g is the gth data tag;
[0061] It should be noted that the unit time can be set to "24 hours", "36 hours" and "72 hours" in length, and the specific length is determined by the experimental personnel according to the historical data calling frequency and the historical service period length, which is not described here;
[0062] It should be noted that the data label refers to the platform pre-processing all data in the database, dividing into multiple groups, and adding corresponding data labels, and the specific format of the data label is not limited;
[0063] The data fluctuation value refers to the change amplitude of the operating parameters (such as temperature, pressure, current, etc.) of the equipment within a unit time, which is used to evaluate the data screening of the equipment, and its acquisition logic is to obtain the time series data of the operating parameters of different data labels in the equipment operating data record, and the standard deviation of the operating parameter value relative to its mean value is calculated to obtain the data fluctuation value ;
[0064] Among them, the device service management platform divides the data into multiple labels and labels them as g, that is, the gth data label, and the specific division rule can be based on the classification of the experimental personnel according to the original data or according to the similarity between the data, and the specific classification method is not limited. The classification method is obtained by the experimental personnel according to the specific embodiment, which is not described here;
[0065] Specifically, the above operating parameters are not limited and can be temperature, pressure, current, etc. In this embodiment, only the operating parameters that best represent the operating parameters are selected, for example, for refrigeration equipment, the fluctuation value of its refrigeration efficiency data is obtained, for heating equipment, the fluctuation value of its heating temperature is obtained, etc. The selection of the specific operating parameter is set by the experimental personnel according to the specific device application scene and the device running characteristics, which is not described here;
[0066] Among them, the formula for calculating the standard deviation of the operating parameter value relative to its mean value is as follows:
[0067] ;
[0068] In the formula, The standard deviation of the gth data in the ith unit time, that is, the data fluctuation value, P is the total number of data labels, is the operating parameter value of the gth data, is the mean value of the operating parameter;
[0069] The data validity period length refers to the time left from the current time to the end of the data validity period, which is an important parameter for measuring the current life cycle state of the data, for predicting the filterability and validity period length of the data, and its acquisition logic is to obtain the end time of the data validity period and the current time in the data information database, and to obtain the data validity period length by subtracting the current time from the end time of the data validity period ;
[0070] Specifically, the data validity period length system compares the end time of the data validity period with the current time by default, and if the end time of the data validity period is greater than the current time, it means that the current data is within the validity period, otherwise, the data validity period is ended, and the data is deleted;
[0071] The data area coverage range value refers to the geographical area range that can be covered by different data tags in a corresponding unit of time, and is usually used to measure the availability and coverage of the data platform in different geographical positions. Its acquisition logic is to calculate the geographical position distribution coordinate area of the data tag by providing the latitude and longitude coordinates of each data tag point, and to calculate the ratio of the area of the data tag to the area of the coverable area in the time range to obtain the data area coverage range value ;
[0072] The latitude and longitude coordinates of each data tag point refer to the specific geographical position associated with each data tag, and the corresponding latitude and longitude points are obtained. The area of the data coverage range is calculated by calculating the minimum rectangular boundary of these points. The range of the boundary box is determined by the maximum and minimum latitude and longitude values, forming a rectangular area. The specific area calculation can use spherical triangle formula calculation, etc., which will not be described here;
[0073] S2: Obtain the data call frequency, data fluctuation value, data validity period length and data area coverage range value, establish a data analysis model, perform logistic regression calculation, and obtain a screening evaluation coefficient;
[0074] The data analysis model refers to a logistic regression calculation model, which generates a screening evaluation coefficient through logistic regression calculation;
[0075] The data call frequency, data fluctuation value, data validity period length and data area coverage range value are substituted into the logistic regression calculation formula as follows:
[0076] ;
[0077] In the formula, is the logistic regression calculation result, i.e. the screening evaluation coefficient, e is the natural base, y is the linear combination term of the logistic regression model, and y can be set as follows:
[0078] ;
[0079] In the formula, is a bias term, , , and are regression coefficients of the data call frequency, the data fluctuation value, the data validity period length and the data area coverage value respectively;
[0080] The data call frequency, the data fluctuation value, the data validity period length and the data area coverage value are all direct data expressions of the current device data paragraph screening selection;
[0081] As can be seen from the formula, the higher the data call frequency, the data fluctuation value and the data area coverage value, the more valuable the current device data paragraph in terms of demand, dynamics, applicability and the like, and the higher the screening evaluation coefficient, and vice versa, the higher the data validity period length, the stronger the timeliness of the current device data paragraph, the lower the priority, and the lower the screening evaluation coefficient;
[0082] S3: Obtain the screening evaluation coefficient and compare it with the preset screening threshold, obtain the device data paragraph for priority screening according to the device data paragraph result marked as a priority paragraph in the comparison result, and determine the data screening paragraph according to the data similarity of each device data paragraph for priority screening;
[0083] The acquisition logic of the screening threshold is to collect the priority classification set of historical device data paragraphs, then divide the data set into a training set and a test set, set an evaluation index and a clustering algorithm, train the model on the training set and evaluate the model performance on the test set in each iteration of cross-validation, and then adjust the screening threshold according to the performance of the validation set, so the screening threshold is constantly updated;
[0084] In the present application, the clustering algorithm is a kind of unsupervised learning algorithm, which is used to divide the priority of the device data paragraphs in the data set into groups or clusters with labels; the common one is K-means clustering, which divides the weighted data in the data set into K clusters, so that the distance between each device data paragraph priority and the center point (centroid) of its belonging cluster is minimized, and finally the priority of the device data paragraph distribution is measured by the Euclidean distance, so as to set the screening threshold;
[0085] After obtaining the screening evaluation coefficient, the screening evaluation coefficient is compared and analyzed with the constantly updated screening threshold;
[0086] If the screening evaluation coefficient is greater than or equal to the screening threshold, the current device data paragraph is marked as a priority paragraph, and a screening signal is generated;
[0087] If the screening evaluation coefficient is less than the screening threshold value, the current device data paragraph is marked as a screening-out paragraph, and an end signal is generated;
[0088] The current device data paragraph marked as a priority paragraph is recorded as a priority screened device data paragraph;
[0089] The three dimensions of values, distributions, and structures in the priority screened device data paragraph are vectorized, and the value similarity, time series similarity, and feature vector similarity are calculated through a cosine similarity formula;
[0090] Specifically, the value similarity, time series similarity, and feature vector similarity are analyzed in three dimensions. In practice, the experimenters can set more dense data similarity according to actual application to more accurately express the data similarity of each priority screened device data paragraph, and perform operations to improve screening accuracy, etc., which will not be described here.
[0091] It should be noted that the cosine similarity is used to calculate the similarity formula in this example, but in actual application, the Euclidean distance can also be used to determine the value similarity, and the dynamic time warping method can also be used to determine the time series similarity, etc. The method of calculating the similarity formula is not limited, but is set according to the calculation model preset by the experimenters, which will not be described here.
[0092] The value similarity, time series similarity, and feature vector similarity are substituted into the weighted formula to calculate the data similarity of each priority screened device data paragraph;
[0093] The similarity of the priority screened device data paragraph data is compared with the preset similarity threshold value. If the similarity of the priority screened device data paragraph data is greater than the similarity threshold value, the corresponding device data paragraph with the compared similarity is screened out, otherwise, the corresponding device data paragraph with the compared similarity is retained, until the comparison ends.
[0094] The device data paragraphs retained in the comparison result are collected to obtain a data screening paragraph;
[0095] The present application collects data call frequency, data fluctuation value, data validity period length, and data area coverage value, establishes a data analysis model, performs logistic regression calculation, obtains a screening evaluation coefficient, compares it with a preset screening threshold value, obtains priority screened device data paragraphs according to the comparison result, determines a data screening paragraph according to the data similarity of each priority screened device data paragraph, performs multi-dimensional comprehensive analysis, improves screening accuracy, improves data collection efficiency, and avoids waste of computing resources.
[0096] Example 2
[0097] The embodiment 1 of the present application mainly illustrates the collection data calling frequency, data fluctuation value, data validity period length and data area coverage value, establishes a data analysis model, performs a logistic regression calculation, obtains a screening evaluation coefficient, and compares with a preset screening threshold value, according to the comparison result, obtains a prior screening device data paragraph, and then determines the operation strategy of the data screening paragraph according to the data similarity of each prior screening device data paragraph. However, in the embodiment 1, only the data screening paragraph is obtained, which provides an accurate method for data screening, but the calling of the data screening paragraph is not sorted, obviously, this will cause the user to still collect more data after scanning the code, which will cause the calculation pressure of subsequent data analysis, delay the service efficiency and reduce the user experience. In view of the above problem, the embodiment 2 of the present application is further refined;
[0098] S4: Obtain the data screening paragraph, and according to the data screening paragraph, collect the data screening paragraph calling frequency user end real-time average upload speed and the data screening paragraph in different area data distribution frequency;
[0099] Specifically, the data screening paragraph contains multiple data tags, and the data tag contains multiple data.
[0100] The acquisition logic of the data screening paragraph calling data user end real-time average upload speed is to obtain the number of calling users in the current data screening paragraph, accumulate the real-time upload speeds of all calling users, and calculate the ratio of the number of calling users in the current data screening paragraph to obtain the data screening paragraph calling data user end real-time average upload speed.
[0101] Wherein, the client application will monitor the upload rate in real time when the user scans the code or other interactive operation, specifically, in the process of implementing the upload module in the client, the upload speed is collected regularly, for example, the upload amount and time consumption are recorded once per second, the real-time upload speed is stored in the local cache, and the real-time collection method is not limited, which will not be described here.
[0102] The data screening paragraph in different area data distribution frequency reflects the activity degree of the data screening paragraph in the geographical position, and its acquisition logic is to extract the geographical position data corresponding to the data screening paragraph from the record of the data source, calculate the ratio of the data amount corresponding to different positions to the total amount of data screening paragraph data, and then accumulate the data distribution frequency in each region to obtain the data screening paragraph in different area data distribution frequency.
[0103] It should be noted that the geographical position selection and quantity can be selected by analyzing the activity degree of the user in each geographical position, and the geographical position with more users or higher activity degree is preferred, and the like, which will not be described here.
[0104] S5: According to the data filtering paragraph call data user end real-time average upload speed and the data filtering paragraph in different regions data distribution frequency, using fuzzy logic to determine the data filtering paragraph call sorting results;
[0105] For example, "Fast", "Slow", "Moderate" for data filtering paragraph call data user end real-time average upload speed, "High", "Low", "Medium" for data filtering paragraph in different regions data distribution frequency;
[0106] A set of fuzzy rules are developed to describe the influence of different input variables on the output variable. The definition of the rules can be based on professional knowledge, or obtained through data analysis and experiments. For example:
[0107] The data filtering paragraph call data user end real-time average upload speed is marked as X, the data filtering paragraph in different regions data distribution frequency is marked as U, and the data filtering paragraph call sorting results is marked as C_results;
[0108] Then it can be defined:
[0109] Rule 1: IF (X is Fast) AND (U is High) THEN (C_results is High)
[0110] Rule 2: IF (U is Slow) AND (U is Low) THEN (C_results is Low) ...
[0111] According to the fuzzy rules, fuzzy reasoning is carried out to determine the data filtering paragraph call sorting results;
[0112] It should be noted that the division of fuzzy sets can be adjusted according to actual conditions. For example, although this embodiment takes three fuzzy sets as an example, in fact, the data filtering paragraph call data user end real-time average upload speed and the data filtering paragraph in different regions data distribution frequency can be divided into more than three sets to facilitate better precision adjustment according to different data tags;
[0113] Further, for the data filtering paragraph to call the data user end real-time average upload speed and the data filtering paragraph in different area data distribution frequency high and low degree judgment, can be set according to the actual situation Threshold value is judged, for example, in the data filtering paragraph to call the data user end real-time average upload speed exceeds 75%, it is marked as "Fast", the data filtering paragraph in different area data distribution frequency is higher than 70% when it is marked as "High", etc., not described here;
[0114] The application obtains the data filtering paragraph, collects the data filtering paragraph calling frequency user end real-time average upload speed and the data filtering paragraph in different area data distribution frequency degree according to the data filtering paragraph, formulates a set of fuzzy rules for fuzzy reasoning, determines the data filtering paragraph calling sorting result, reduces the calculation pressure of subsequent data analysis, improves service efficiency, and enhances user experience.
[0115] Embodiment 3
[0116] Please refer to Figure 2 A data acquisition system for a device service management platform, comprising a data acquisition module, a data processing module, a screening analysis module and a paragraph sorting module;
[0117] The data acquisition module is used for collecting data platform feature information and data service feature information according to the device update log database, performing data processing, obtaining data calling frequency, data fluctuation value, data effective period length and data area coverage value, and sending to the data processing module;
[0118] The data processing module is used for obtaining the data calling frequency, the data fluctuation value, the data effective period length and the data area coverage value, establishing a data analysis model, performing logistic regression calculation, obtaining a screening evaluation coefficient, and sending to the screening analysis module;
[0119] The screening analysis module is used for obtaining the screening evaluation coefficient, comparing with a preset screening threshold value, obtaining the device data paragraph result marked as a priority paragraph in the comparison result to obtain the priority screened device data paragraph, and determining the data filtering paragraph according to the data similarity of each priority screened device data paragraph, and sending to the paragraph sorting module;
[0120] The paragraph sorting module is used for obtaining the data filtering paragraph, collecting the data filtering paragraph calling frequency user end real-time average upload speed and the data filtering paragraph in different area data distribution frequency according to the data filtering paragraph, and determining the data filtering paragraph calling sorting result by using fuzzy logic.
[0121] The above formulas are all de-dimensioned to calculate the numerical values, the formulas are obtained by collecting a large amount of data to simulate a formula of the most recent real situation, and the preset parameters in the formulas are set by a person skilled in the art according to the actual situation.
[0122] The above embodiments can be implemented wholly or partially by software, hardware, firmware, or any other combination. When implemented by software, the above embodiments can be implemented wholly or partially in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are wholly or partially generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another, for example, the computer instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center through a wired or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. containing one or more available medium collections. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state disk.
[0123] It should be understood that the size of the sequence number of each process described above in various embodiments of the present application does not mean the order of execution, and the execution order of each process should be determined by its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0124] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0125] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the above-described system, device, and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be described here.
[0126] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the described device embodiments are merely schematic. The division of the units is merely logical function division. There can be other division manners in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.
[0127] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0128] In addition, each functional unit in the various embodiments of the present application can be integrated into a processing unit, or each unit can be a physically separate unit, or two or more units can be integrated into one unit.
[0129] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0130] The above description is merely a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A data collection method for a device service management platform, characterized by: The method comprises the following steps: S1: According to the device update log database, collect data platform feature information and data service feature information, perform data processing, obtain data call frequency, data fluctuation value, data validity period length and data area coverage range value; S2: Obtain the data call frequency, data fluctuation value, data validity period length and data area coverage range value, establish a data analysis model, perform logical regression calculation, and obtain a screening evaluation coefficient; S3: Obtain the screening evaluation coefficient, compare it with the preset screening threshold, obtain the priority screening device data paragraph according to the device data paragraph result marked as the priority paragraph in the comparison result, and determine the data screening paragraph according to the data similarity of each priority screening device data paragraph; S4: Obtain the data screening paragraph, collect the data screening paragraph call frequency user end real-time average upload speed and the data screening paragraph in different area data distribution frequency; S5: According to the data screening paragraph call data user end real-time average upload speed and the data screening paragraph in different area data distribution frequency, use fuzzy logic to determine the data screening paragraph call sorting result; The data platform feature information includes data call frequency and data fluctuation value; the data service feature information includes data validity period length and data area coverage range value; By determining the unit time of analysis, recording the event of each data call, in the set unit time, counting the total number of data call occurrences, and calculating the ratio of the total number of data calls in the unit time to the unit time length, the data call frequency is obtained ; wherein i is the i th unit time, and g is the g th data tag The running parameter time series data of different data tags is obtained through the device running data record, and the data fluctuation value is obtained by calculating the standard deviation of the running parameter value relative to the mean value ; By obtaining the end time of the validity period of the data in the data information database and the current time, subtracting the current time from the end time of the validity period of the data to obtain the data validity period length ; The geographical position distribution coordinate area of the data label is obtained by providing the data longitude and latitude coordinates of each data label point, and the data area coverage range value is obtained by ratio calculation between the data area coverage range in the time range and the covered area ; Substitute the data call frequency, data fluctuation value, data validity period length and data area coverage range value into the logical regression calculation specific formula, which is expressed as follows: ; In the formula, is the result of the logistic regression, i.e. the screening evaluation coefficient, e is the natural base, and y is the linear combination term of the logistic regression model. Specifically, y can be set as: ; In the formula, is a bias term, , , and are regression coefficients of the data calling frequency, the data fluctuation value, the data validity period length, and the data area coverage range value, respectively.
2. The data collection method for device service management platform according to claim 1, characterized in that: After obtaining the screening evaluation coefficient, compare the screening evaluation coefficient with the continuously iterated screening threshold; If the screening evaluation coefficient is greater than or equal to the screening threshold, mark the current device data paragraph as the priority paragraph, and generate a screening signal; If the screening evaluation coefficient is less than the screening threshold, mark the current device data paragraph as the screening-out paragraph, and generate an end signal.
3. The data collection method for device service management platform according to claim 2, characterized in that: Mark the current device data paragraph marked as the priority paragraph as the priority screening device data paragraph; Vectorize the value, distribution and structure dimensions in the priority screening device data paragraph, and calculate the value similarity, time series similarity and feature vector similarity through the cosine similarity formula.
4. The data collection method for device service management platform according to claim 3, characterized in that: Substitute the value similarity, time series similarity and feature vector similarity into the weighted formula to calculate the data similarity of each priority screening device data paragraph; Compare the data similarity of the priority screening device data paragraph with the preset similarity threshold, if the data similarity of the priority screening device data paragraph is greater than the similarity threshold, screen out the device data paragraph corresponding to the comparison similarity, otherwise, retain the device data paragraph corresponding to the comparison similarity, until the comparison is completed; Collect the device data paragraphs retained in the comparison result to obtain the data screening paragraph.
5. The data collection method for device service management platform according to claim 4, characterized in that: The data screening paragraph contains multiple data tags, and each data tag contains multiple data. The real-time average upload speed of the data user end of the data screening paragraph is obtained by acquiring the number of called users in the current data screening paragraph, accumulating the real-time upload speeds of all called users, and performing ratio calculation on the number of called users in the current data screening paragraph; The geographical position data corresponding to the data screening paragraph is extracted from the record of the data source, the data amount corresponding to different positions is ratio calculated with the total amount of data in different data screening paragraphs, and the data distribution frequency in each region is accumulated to obtain the data distribution frequency of the data screening paragraph in different regions.
6. The data collection method for device service management platform according to claim 5, characterized in that: The real-time average upload speed of the data user end of the data screening paragraph and the data distribution frequency of the data screening paragraph in different regions are divided into different fuzzy sets, respectively. The sorting result of the data screening paragraph is defined as an output variable and divided into a fuzzy set. Fuzzy rules are formulated to describe the influence of the real-time average upload speed of the data user end of the data screening paragraph and the data distribution frequency of the data screening paragraph in different regions on the sorting result of the data screening paragraph. The sorting result of the data screening paragraph is determined according to the fuzzy rules.
7. A data collection system for a device service management platform, which is used to implement the data collection method for a device service management platform according to any one of claims 1-6, characterized in that: The data acquisition module, the data processing module, the screening analysis module, and the paragraph sorting module are included. The data acquisition module is used to collect data platform feature information and data service feature information according to the equipment update log database, perform data processing, obtain data call frequency, data fluctuation value, data effective period length, and data region coverage range value, and send them to the data processing module. The data processing module is used to obtain the data call frequency, the data fluctuation value, the data effective period length, and the data region coverage range value, establish a data analysis model, perform logistic regression calculation, obtain a screening evaluation coefficient, and send it to the screening analysis module. The screening analysis module is used to obtain the screening evaluation coefficient, compare it with a preset screening threshold, obtain the priority screening equipment data paragraph according to the equipment data paragraph result marked as a priority paragraph in the comparison result, determine the data screening paragraph according to the data similarity of each priority screening equipment data paragraph, and send it to the paragraph sorting module. The paragraph sorting module is used to obtain the data screening paragraph, collect the real-time average upload speed of the data user end of the data screening paragraph and the data distribution frequency of the data screening paragraph in different regions, and determine the sorting result of the data screening paragraph using fuzzy logic.
Citation Information
Patent Citations
Effective data screening method, readable storage medium and terminal
CN109542927A
Data processing system suitable for accounting financial management
CN119088830A