Method and device for correcting terminal model library, electronic device and medium
By identifying non-standard terminal models in user agent data and establishing a mapping relationship with standard models, the problem of inaccurate terminal model library under self-registration method is solved, and the accuracy of terminal model library is improved.
Patent Information
- Application Number
- CN202111588647.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-23
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-12-23
AI Technical Summary
In the prior art, the registration and reporting terminal model library established by the self-registration method is inaccurate and cannot truly obtain the user's terminal data, resulting in insufficient accuracy of the library.
By obtaining user agent data in user behavior data, identifying non-standard terminal models based on the identification algorithm, establishing a mapping relationship between non-standard terminal models and standard terminal models, and then correcting the terminal models in the registered and reported terminal model library.
Improve the accuracy of the terminal model library registered and reported, ensuring the authenticity and consistency of terminal model data.
Smart Images

Figure CN114265829B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular, to a method and device for correcting a terminal model library, an electronic device, and a medium. Background Art
[0002] An operator can establish a registered reported terminal model library based on the terminal information reported by a user during self-registration. The self-registration method is that when the user turns on or off the device or inserts or removes the SIM card (Subscriber Identity Module), a registration request is initiated in the form of a short message to obtain the user's terminal model, device identification code, SIM card number, etc. Due to system configuration or serial number forgery, the self-registration method cannot truly obtain the user's terminal data, resulting in an inaccurate registered reported terminal model library.
[0003] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0004] The present disclosure provides a method and device for correcting a terminal model library, an electronic device, and a medium, which can realize intelligent correction of the registered reported terminal model and improve the accuracy of the registered reported terminal model library. The technical solution of the present disclosure is as follows:
[0005] According to one aspect of the embodiments of the present disclosure, a method for correcting a terminal model library is provided, including: obtaining user agent data in user behavior data, and identifying non-standard terminal models in the user agent data based on an identification algorithm for the user agent data; generating a mapping relationship between the non-standard terminal models and standard terminal models according to the registered reported terminal model library; and correcting the terminal models in the registered reported terminal model library according to the mapping relationship between the non-standard terminal models and standard terminal models.
[0006] In an embodiment of the present disclosure, identifying non-standard terminal models in the user agent data based on an identification algorithm for the user agent data includes: labeling category tags of the user agent data; obtaining user agent data of each category according to the labeled category tags of the user agent data; determining a data processing order of the user agent data of each category according to the top-down order of a pre-constructed category tree; and performing identification processing on the user agent data of each category according to the data processing order to obtain non-standard terminal models in the user agent data.
[0007] In one embodiment of the present disclosure, the category tree is pre-constructed according to the following method: obtain sample data, and label the category labels of the sample data; according to the labeled category labels of the sample data, obtain the sample data of each category; based on the modified empirical entropy and the modified conditional entropy, calculate the sample data of each category, and sequentially extract the sample data of each category; construct the category tree according to the extraction order of the sample data of each category.
[0008] In one embodiment of the present disclosure, based on the modified empirical entropy and the modified conditional entropy, calculating the sample data of each category and sequentially extracting the sample data of each category includes: calculating the information gain of the sample data of each category with respect to the sample data based on the modified empirical entropy and the modified conditional entropy; determining the category corresponding to the maximum information gain as the target category, and extracting the sample data of the target category from the sample data.
[0009] In one embodiment of the present disclosure, after extracting the sample data of the target category from the sample data, the method further includes: obtaining the remaining sample data; calculating the information gain of the sample data of each category in the remaining sample data with respect to the remaining sample data based on the modified empirical entropy and the modified conditional entropy; determining the category corresponding to the maximum information gain as the new target category, and extracting the sample data of the new target category from the remaining sample data.
[0010] In one embodiment of the present disclosure, the modified empirical entropy and the modified conditional entropy include a user-defined modification coefficient; the user-defined modification coefficient is obtained by statistically analyzing the abnormal data in the sample data; and the abnormal data includes data with an identified abnormal terminal model and data with an abnormal category.
[0011] In one embodiment of the present disclosure, after identifying the non-standard terminal model in the user agent data based on the identification algorithm of the user agent data, the method further includes: obtaining the identification result, and verifying the category tree according to the identification result.
[0012] In an embodiment of the present disclosure, generating a mapping relationship between the non-standard terminal model and the standard terminal model according to the registered reported terminal model library includes: associating the non-standard terminal model with the registered reported terminal models in the registered reported terminal model library to obtain one or more registered reported terminal models associated with the non-standard terminal model; selecting a target terminal model from the one or more registered reported terminal models according to the number of users of the one or more registered reported terminal models; if the number of users of the target terminal model is greater than a preset user number threshold and the proportion of the number of users of the target terminal model in the number of users of the non-standard terminal model is greater than a preset proportion threshold, determining that the target terminal model is the standard terminal model corresponding to the non-standard terminal model, and generating a mapping relationship between the non-standard terminal model and the target terminal model.
[0013] In an embodiment of the present disclosure, correcting the terminal models in the registered reported terminal model library according to the mapping relationship between the non-standard terminal model and the standard terminal model includes: obtaining the user identifier of the non-standard terminal model; querying the registered reported terminal model of the user identifier from the registered reported terminal model library according to the user identifier; if the registered reported terminal model of the user identifier does not exist in the registered reported terminal model library, inserting the standard terminal model corresponding to the non-standard terminal model into the registered reported terminal model library; if the standard terminal model corresponding to the non-standard terminal model is inconsistent with the registered reported terminal model of the user identifier, replacing the registered reported terminal model of the user identifier with the standard terminal model corresponding to the non-standard terminal model.
[0014] On the other hand, according to an embodiment of the present disclosure, a device for correcting a terminal model library is provided, including: an identification module, configured to obtain user agent data in user behavior data and identify non-standard terminal models in the user agent data based on an identification algorithm of the user agent data; a generation module, configured to generate a mapping relationship between the non-standard terminal model and the standard terminal model according to the registered reported terminal model library; a correction module, configured to correct the terminal models in the registered reported terminal model library according to the mapping relationship between the non-standard terminal model and the standard terminal model.
[0015] In an embodiment of the present disclosure, the identification module is further configured to: label a category label of the user agent data; obtain user agent data of each category according to the labeled category label of the user agent data; determine a data processing order of the user agent data of each category according to a top-down order of a pre-constructed category tree; and perform identification processing on the user agent data of each category according to the data processing order to obtain non-standard terminal models in the user agent data.
[0016] In one embodiment of the present disclosure, the device further includes a construction module, which is used to pre-construct a category tree according to the following method: obtain sample data, and label the category labels of the sample data; according to the labeled category labels of the sample data, obtain the sample data of each category; based on the corrected empirical entropy and the corrected conditional entropy, calculate the sample data of each category, and sequentially extract the sample data of each category; construct the category tree according to the extraction order of the sample data of each category.
[0017] In one embodiment of the present disclosure, the construction module is further used to: calculate the information gain of the sample data of each category with respect to the sample data based on the corrected empirical entropy and the corrected conditional entropy; determine the category corresponding to the maximum information gain as the target category, and extract the sample data of the target category from the sample data.
[0018] In one embodiment of the present disclosure, the construction module is further used to: obtain the remaining sample data; calculate the information gain of the sample data of each category in the remaining sample data with respect to the remaining sample data based on the corrected empirical entropy and the corrected conditional entropy; determine the category corresponding to the maximum information gain as the new target category, and extract the sample data of the new target category from the remaining sample data.
[0019] In one embodiment of the present disclosure, the corrected empirical entropy and the corrected conditional entropy include a custom correction coefficient; the custom correction coefficient is obtained by statistically analyzing the abnormal data in the sample data; and the abnormal data includes data with an identified abnormal terminal model and data with an abnormal category.
[0020] In one embodiment of the present disclosure, the device further includes a verification module, which is used to: obtain the recognition result and verify the category tree according to the recognition result.
[0021] In one embodiment of the present disclosure, the generation module is further used to: associate the non-standard terminal model with the registered and reported terminal models in the registered and reported terminal model library to obtain one or more registered and reported terminal models associated with the non-standard terminal model; select a target terminal model from the one or more registered and reported terminal models according to the number of users of the one or more registered and reported terminal models; if the number of users of the target terminal model is greater than a preset user number threshold and the proportion of the number of users of the target terminal model in the number of users of the non-standard terminal model is greater than a preset proportion threshold, determine the target terminal model as the standard terminal model corresponding to the non-standard terminal model, and generate a mapping relationship between the non-standard terminal model and the target terminal model.
[0022] In one embodiment of the present disclosure, the correction module is further configured to: obtain the user identifier of the non-standard terminal model; query the registered reported terminal model of the user identifier from the registered reported terminal model library according to the user identifier; if the registered reported terminal model of the user identifier does not exist in the registered reported terminal model library, insert the standard terminal model corresponding to the non-standard terminal model into the registered reported terminal model library; if the standard terminal model corresponding to the non-standard terminal model is inconsistent with the registered reported terminal model of the user identifier, replace the registered reported terminal model of the user identifier with the standard terminal model corresponding to the non-standard terminal model.
[0023] According to another aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the instructions to implement the method for correcting the terminal model library as described above.
[0024] According to another aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the method for correcting the terminal model library as described above.
[0025] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects: starting from the perspective that the user agent data contains the terminal model, through the recognition algorithm of the user agent data, the non-standard terminal model in the user agent data is recognized, and then with the help of the existing registered reported terminal model library, the mapping relationship between the non-standard terminal model and the standard terminal model is established, and then the registered reported terminal model library can be corrected according to the mapping relationship, improving the accuracy of the registered reported terminal model library.
[0026] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure and do not constitute an improper limitation of the present disclosure.
[0028] Figure 1 is a flowchart of a method for correcting a terminal model library shown according to an exemplary embodiment;
[0029] Figure 2 is a flowchart of identifying a non-standard terminal model in user agent data shown according to an exemplary embodiment;
[0030] Figure 3is a schematic diagram of a category tree shown according to an exemplary embodiment;
[0031] Figure 4 is a flowchart of constructing a category tree shown according to an exemplary embodiment;
[0032] Figure 5 is a flowchart of constructing a category tree shown according to another exemplary embodiment;
[0033] Figure 6 is a flowchart of generating a mapping relationship shown according to an exemplary embodiment;
[0034] Figure 7 is a flowchart of a method for correcting a terminal model library shown according to another exemplary embodiment;
[0035] Figure 8 is a block diagram of a device for correcting a terminal model library shown according to an exemplary embodiment;
[0036] Figure 9 is a block diagram of an electronic device for correcting a terminal model library shown according to an exemplary embodiment. Detailed implementation manners
[0037] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0038] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such used data may be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0039] It should be noted that the user information involved in the present disclosure, including but not limited to user device information, user personal information, etc., are all information authorized by the user or fully authorized by all parties.
[0040] The method provided by the embodiments of the present disclosure can be executed by any type of electronic device, such as a server or a terminal device, or the interaction between a server and a terminal device. The terminal device and the server can be directly or indirectly connected through wired or wireless communication methods, and the present application does not make any restrictions here.
[0041] Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0042] The terminal device can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto.
[0043] Figure 1 It is a flowchart of a method for correcting a terminal model library shown according to an exemplary embodiment. As Figure 1 shown, the method for correcting the terminal model library includes the following steps.
[0044] Step S110, obtain the user agent data in the user behavior data, and based on the recognition algorithm of the user agent data, recognize the non-standard terminal models in the user agent data.
[0045] Step S120, generate a mapping relationship between the non-standard terminal models and the standard terminal models according to the registered and reported terminal model library.
[0046] Step S130, correct the terminal models in the registered and reported terminal model library according to the mapping relationship between the non-standard terminal models and the standard terminal models.
[0047] Among them, the user behavior data can be the data generated when the user uses the terminal to access the Internet, that is, the user's Internet access behavior log, that is, the user's DPI (Deep Packet Inspection) data. Considering that the user behavior data can include the user agent data, and the user agent data contains non-standard terminal models, the user behavior data can be collected, then the user agent data in the user behavior data can be obtained, and then based on the recognition algorithm of the user agent data, the non-standard terminal models in the user agent data can be recognized.
[0048] Since the identified terminal model is in a non-standard form, in step S102, a mapping relationship between the non-standard terminal model and the standard terminal model can be generated with the help of the registered and reported terminal model library. It should be noted that the terminal models stored in the registered and reported terminal model library are the terminal models self-registered and reported by users, and their forms are standard. However, due to system configuration or serial number forgery, the self-registration method cannot truly obtain the user's terminal model, which leads to the inaccuracy of the registered and reported terminal model library. For example, the terminal model reported by the user for mobile phone P1 is OB-PBAM00, but the actual terminal model of mobile phone P1 is HW-HMA AL00, and the terminal model of mobile phone P1 identified through user agent data is HMA-AL00. Among them, both OB-PBAM00 and HW-HMA AL00 are standard terminal models, and HMA-AL00 is a non-standard terminal model. In step S102, a mapping relationship between the non-standard terminal model HMA-AL00 and the standard terminal model HW-HMA AL00 can be established.
[0049] In step S103, the terminal models in the registered and reported terminal model library can be corrected according to the mapping relationship between the non-standard terminal model and the standard terminal model. Continuing with the above example, after establishing the mapping relationship between the non-standard terminal model HMA-AL00 and the standard terminal model HW-HMA AL00, the terminal model OB-PBAM00 of P1 in the registered and reported terminal model library can be replaced with HW-HMA AL00.
[0050] The method for correcting the terminal model library provided by the embodiments of the present disclosure can start from the perspective of the terminal model included in the user agent data, identify the non-standard terminal model in the user agent data through the identification algorithm of the user agent data, and then establish a mapping relationship between the non-standard terminal model and the standard terminal model with the help of the existing registered and reported terminal model library, so as to correct the registered and reported terminal model library according to the mapping relationship and improve the accuracy of the registered and reported terminal model library.
[0051] Identifying the non-standard terminal model from the user agent data is an important part of the embodiments of the present disclosure. Figure 2 is a flowchart showing the identification of non-standard terminal models in user agent data according to an exemplary embodiment. As Figure 2 shown, identifying the non-standard terminal model in the user agent data may include the following steps.
[0052] Step S201, label the category tags of the user agent data.
[0053] The user agent data can be data for a period of time, such as collecting user behavior data for one day or one hour. The user agent data is obtained from the collected user behavior data, and then category labels are assigned to each user agent data. The category labels are the categories to which the user agent data belongs, such as general category, non-general category, linux category, non-linux category, weibo category, non-weibo category, and network disk category.
[0054] Step S202: According to the category labels of the labeled user agent data, obtain the user agent data of each category.
[0055] After assigning category labels to the user agent data, the user agent data can be classified according to the category labels to obtain the user agent data of each category. For example, there are 10,000 user agent data. Category labels are assigned to these 10,000 user agent data, and it is statistically obtained that there are 10 categories. Then these 10,000 user agent data can be divided into 10 categories to obtain the user agent data of each category.
[0056] Step S203: According to the top-down order of the pre-constructed category tree, determine the data processing order of the user agent data of each category.
[0057] Step S204: According to the data processing order, perform identification processing on the user agent data of each category to obtain non-standard terminal models in the user agent data.
[0058] After obtaining the user agent data of each category, the data processing order of the user agent data of each category can be determined according to the top-down order of the pre-constructed category tree. That is, it is determined which category of user agent data to process first. Figure 3 It is a schematic diagram of a category tree shown according to an exemplary embodiment. Refer to Figure 3 , the top-down order of the category tree is: general category, ting category, xiaomi category, linux category, weibo category, okhttp category / network disk category / IOS category. That is to say, the user agent data of the general category can be processed first to obtain the non-standard terminal models in the user agent data of the general category. Considering that the user agent data of the same user may include multiple categories, after the identification processing of the user agent data of each category, it can be determined whether the non-standard terminal models of all users have been identified. If so, there is no need to perform identification processing on the user agent data of the remaining categories.
[0059] For example, according to Figure 3The category tree shown identifies and processes 20,000 user agent data of 1,000 users. First, category labels are assigned to the 20,000 user agent data to obtain user agent data for each category. Then, the user agent data of the general category is first identified and processed to obtain non-standard terminal models. Next, based on all the obtained non-standard terminal models (i.e., the non-standard terminal models obtained from the user agent data of the general category), it is determined whether the non-standard terminal models of these 1,000 users have been identified. If so, the user agent data of other categories does not need to be identified and processed. If not, the user agent data of the ting category can be identified and processed to obtain non-standard terminal models, and then based on all the obtained non-standard terminal models (i.e., the non-standard terminal models obtained from the user agent data of the general category and the ting category), it is determined whether the non-standard terminal models of these 1,000 users have been identified. If so, the user agent data of other categories does not need to be identified and processed. If not, the user agent data of the next category can be identified and processed, and so on, until the non-standard terminal models of these 1,000 users are identified. That is to say, after identifying and processing the user agent data of each category to obtain non-standard terminal models, it can be determined whether the non-standard terminal models of all users have been identified. If so, the user agent data of the remaining categories does not need to be identified and processed. If not, the user agent data of the next category can be identified and processed. Table 1 is the recognition result table of the user agent data, and the non-standard terminal models obtained by recognizing the user agent data can be obtained from Table 1.
[0060] Table 1 is the recognition result table of the user agent data
[0061]
[0062]
[0063] Figure 4 is a flowchart of constructing a category tree shown according to an exemplary embodiment. As Figure 4 shown, the category tree is constructed according to the following method.
[0064] Step S401, obtain sample data and label the category labels of the sample data.
[0065] Step S402, according to the labeled category labels of the sample data, obtain the sample data of each category.
[0066] Among them, the sample data can be user agent data for which non-standard terminal models have been identified. To improve the accuracy of the category tree, a certain number of sample data need to be obtained, such as 1 million sample data. After obtaining the sample data, category labels are assigned to each sample data, and then the sample data are classified according to the category labels to obtain sample data for each category.
[0067] Step S403: Calculate the sample data for each category based on the corrected empirical entropy and the corrected conditional entropy, and sequentially extract the sample data for each category.
[0068] Step S404: Construct a category tree according to the extraction order of the sample data for each category.
[0069] Among them, the corrected empirical entropy and the corrected conditional entropy can include a custom correction coefficient. The custom correction coefficient is obtained by statistically analyzing the abnormal data in the sample data; and, the abnormal data includes data with identified terminal models as outliers and data with categories as outliers. In the process of constructing the category tree, the corrected empirical entropy and the corrected conditional entropy can be used. Next, the derivation process of the corrected empirical entropy and the corrected conditional entropy will be described. The formula for the empirical entropy is:
[0070]
[0071] where C k is the number of sample data for category k, D is the total number of sample data, n is the category, is the probability of the sample data for category k appearing. Since the non-standard terminal models identified from the data may be empty or outliers, it is necessary to consider the probability of successful data identification, that is, the probability that the sample data can successfully identify non-standard terminal models. In addition, there are also data in the sample data for which the category label annotation fails, and these are also abnormal data. That is to say, the data with the identified terminal model as an outlier is the data for which the identification of the non-standard terminal model fails, and the data with the category as an outlier is the data for which the category label annotation fails.
[0072] Suppose there are m k abnormal data with failed identification of non-standard terminal models in C k , and C n is the data for which the category label annotation fails in the sample data. Then the probability of success is Abstract α as the correction coefficient, then the empirical entropy is corrected to the empirical entropy in the abnormal case:
[0073]
[0074] where Substitute the correction coefficient and expand to get:
[0075]
[0076] Therefore, in the case of an anomaly, the conditional entropy of feature A corresponding to the sample data of a certain category with respect to the sample data D is corrected to:
[0077]
[0078] where A represents the feature of the sample data of a certain category, Di represents the sample data of the i-th category, |D i | represents the number of sample data of the i-th category, |D| represents the total number of sample data, and the conditional entropy represents the uncertainty of the sample data D under the condition of the known random variable A.
[0079] That is to say, in step S403, the empirical entropy and conditional entropy of the sample data of each category can be calculated, and then according to the calculated results, the sample data of each category are extracted in turn. In the embodiments of the present disclosure, the purpose of extracting the sample data of each category is to construct a category tree, and the category tree is used to determine the data processing order of the user agent data of each category during the recognition process. Therefore, the main purpose of extracting the sample data of each category is to determine the extraction order of each category.
[0080] In an exemplary embodiment, based on the corrected empirical entropy and the corrected conditional entropy, the sample data of each category are calculated, and the sample data of each category are extracted in turn, which may include: calculating the information gain of the sample data of each category with respect to the sample data based on the corrected empirical entropy and the corrected conditional entropy; determining the category corresponding to the maximum information gain as the target category, and extracting the sample data of the target category from the sample data.
[0081] where the formula for information gain is: g(D,A) = H modified (D) - H modified (D|A), that is, for the sample data A of each category, the information gain of the sample data A of this category with respect to the sample data D is calculated. Considering that the larger the information gain, the more and more important the information brought, after calculating the information gain of the sample data of each category with respect to the sample data, the maximum information gain is selected, and the category corresponding to the maximum information gain is determined as the target category. When constructing the category tree, the target category can be constructed first, so that when identifying the user agent data in the top-down order of the category tree, the user agent data of the target category can be identified preferentially.
[0082] In addition, after determining the target category, the sample data of the target category can be segmented out from the sample data, which facilitates subsequent analysis of the remaining sample data. Therefore, in an exemplary embodiment, after extracting the sample data of the target category from the sample data, the remaining sample data can be obtained; based on the modified empirical entropy and the modified conditional entropy, the information gain of the sample data of each category in the remaining sample data with respect to the remaining sample data is calculated; the category corresponding to the maximum information gain is determined as the new target category, and the sample data of the new target category is extracted from the remaining sample data.
[0083] Specifically, after obtaining the remaining sample data, based on the modified empirical entropy and the modified conditional entropy, the information gain of the sample data of each category (i.e., the sample data of other categories except the sample data of the target category) in the remaining sample data with respect to the remaining sample data can be calculated. Then, the maximum information gain is selected, and the category corresponding to the maximum information gain is determined as the new target category. When constructing the category tree, the target category can be constructed first, and then the new target category. In this way, when identifying the user agent data in the order from top to bottom of the category tree, the user agent data of the target category can be identified first, and then the user agent data of the new target category.
[0084] It should be noted that after determining the new target category, the sample data of the new target category can be segmented out from the remaining sample data to obtain the new remaining sample data. If the new remaining sample data still includes other categories, that is, categories other than the target category and the new target category, then the new remaining sample data needs to be analyzed continuously until the sample data of all categories are extracted to obtain the order of all categories, and then the category tree is constructed according to the order of all categories.
[0085] Figure 5 is a flowchart of constructing a category tree shown according to another exemplary embodiment. As Figure 5 shown, the process of constructing a category tree may include the following steps.
[0086] Step S501, obtain sample data and label the category labels of the sample data;
[0087] Step S502, according to the labeled category labels of the sample data, obtain the sample data of each category;
[0088] Step S503, based on the modified empirical entropy and the modified conditional entropy, calculate the information gain of the sample data of each category with respect to the sample data;
[0089] Step S504, determine the category corresponding to the maximum information gain as the target category;
[0090] Step S505: Extract the sample data of the target category from the sample data;
[0091] Step S506: Determine whether all categories have been traversed. If so, execute Step S508; if not, execute Step S507;
[0092] Step S507: Update the sample data and the sample data of each category in the sample data;
[0093] Step S508: Construct a category tree.
[0094] Among them, the execution order of Step S505 can be adjusted according to the actual situation and can be executed before Step S507.
[0095] In an exemplary embodiment, after identifying non-standard terminal models in user agent data based on an identification algorithm for user agent data, the method for correcting the terminal model library may further include: obtaining the identification result and verifying the category tree according to the identification result.
[0096] After obtaining the identification result, it can first be determined whether the identification result meets a preset condition, such as the identification accuracy reaching a set threshold. The specific condition can be set according to the actual situation and is not limited herein. If the identification result does not meet the preset condition, the category tree can be verified. Specifically, after obtaining the user agent data, the data with failed annotation category labels in the user agent data can be obtained; after obtaining the identification result, the recognition success probability of the user agent data of each category can be obtained. Therefore, based on the modified empirical entropy and the modified conditional entropy, the user agent data of each category can be calculated, and the user agent data of each category can be extracted in turn, that is, the priority order of each category can be obtained. Then, the obtained priority order of each category is compared with the top-down order of the category tree. If they are inconsistent, the sample data can be recollected and then the category tree can be constructed.
[0097] In the process of identifying the user agent data to obtain the non-standard terminal model in the embodiment of the present disclosure, the construction of a category tree and identification can be included. In the process of constructing the category tree, considering the probability of successful data identification and the data in the sample data that fail to label the category label, the empirical entropy and conditional entropy are corrected to the empirical entropy and conditional entropy under abnormal conditions, and then the information gain of the sample data of each category is calculated, and the extraction order of the sample data of each category is determined by using the characteristic that the greater the information gain, the more important the information is, and then the category tree is constructed according to the extraction order. In the identification process, the category priority can be determined according to the top-down order of the category tree, and the processing priority order of the user agent data of each category is obtained according to the principle of the lowest possible degree of data confusion and the shortest possible calculation time, and finally the non-standard terminal model is obtained. In addition, in the identification process, after the identification and processing of the user agent data of each category, if it is determined that the non-standard terminal model of all users has been identified, it is not necessary to identify and process the user agent data of the remaining categories, thereby improving the identification efficiency. In addition, after obtaining the identification result, the category tree can be verified using the identification result, which can ensure the accuracy of the category tree, and further ensure the accuracy of the identified non-standard terminal model.
[0098] Figure 6 FIG. 1 is a flow chart showing a method of generating a mapping relationship according to an exemplary embodiment. Figure 6 As shown, generating a mapping relationship may include the following steps.
[0099] Step S601: associate the non-standard terminal model with the registered and reported terminal models in the registered and reported terminal model library to obtain one or more registered and reported terminal models associated with the non-standard terminal model.
[0100] The terminal model obtained by identifying the user proxy data is not a standard terminal model, such as PBAM00, HMA-AL00, vivo X9, while the standard terminal models are OB-PBAM00, HW-HMA AL00, VIV-VIVO X9. Considering that the registered and reported terminal models are standard terminal models, in order to obtain a statistically significant terminal model mapping relationship in the existing network, the non-standard terminal model obtained by identifying the user proxy data can be associated with the registered and reported terminal model in the registered and reported terminal model library according to the MDN (Mobile Directory Number, which is the number that the calling user needs to dial when the mobile user of this network is called) user number.
[0101] It should be noted that a non-standard terminal model can be associated with one or more registered reported terminal models. For example, by identifying the user agent data of User 1 and User 2, their non-standard terminal models are both PBAM00. In the registered reported terminal model library, the terminal model registered and reported by User 1 is OB-PBAM00, and the terminal model registered and reported by User 2 is OB-PBEM00. Then, the registered reported terminal models associated with the non-standard terminal model PBAM00 are OB-PBAM00 and OB-PBEM00.
[0102] Step S602: Select a target terminal model from one or more registered reported terminal models according to the number of users of one or more registered reported terminal models.
[0103] Specifically, the number of users of one or more registered reported terminal models can be sorted from high to low, and the registered reported terminal model ranked first is determined as the target terminal model.
[0104] Step S603: If the number of users of the target terminal model is greater than the preset user number threshold and the proportion of the number of users of the target terminal model in the number of users of the non-standard terminal model is greater than the preset proportion threshold, then determine that the target terminal model is the standard terminal model corresponding to the non-standard terminal model, and generate a mapping relationship between the non-standard terminal model and the target terminal model.
[0105] If the target terminal model meets the following conditions, then it can be determined that the target terminal model is the standard terminal model corresponding to the non-standard terminal model, and a mapping relationship between the non-standard terminal model and the target terminal model is generated. Condition (1) is that the number of users is greater than the preset user number threshold. That is, the number of users who register and report this target terminal model in the registered reported terminal model library is greater than the preset user number threshold. Among them, the user number threshold is preferably 3000, and can be adjusted according to the actual situation specifically. Condition (2) is that the proportion of the number of users in the number of users of the non-standard terminal model is greater than the preset proportion threshold. Among the number of users of the non-standard terminal model, the proportion of the number of users who register and report this target terminal model is greater than the preset proportion threshold. Among them, the preset proportion threshold is preferably 80%, and can be adjusted according to the actual situation specifically.
[0106] For example, the non-standard terminal model is PBAM00, and the registered and reported terminal models associated with it are OB-PBAM00 and OB-PBEM00. After statistics, if the number of users who register and report OB-PBAM00 in the registered and reported terminal model library is greater than 3000, and the proportion of the number of users who register and report OB-PBAM00 among the users of PBAM00 is greater than 80%, then OB-PBAM00 is considered as the standard terminal model of PBAM00, and a mapping relationship is generated. For the convenience of understanding, a mapping relationship table shown in Table 2 is provided. It can be obtained from Table 2 the non-standard terminal model identified through user agent data, the registered and reported terminal model corresponding to the non-standard terminal model, and the number of times the registered and reported terminal model appears in the registered and reported terminal model library. The top1 in Table 2 can be understood as the registered and reported terminal model with the largest number of users among the users of the non-standard terminal model.
[0107] Table 2 Mapping Relationship Table
[0108]
[0109]
[0110] In the embodiments of the present disclosure, by using the registered and reported terminal model library, a mapping relationship between the non-standard terminal model and the standard terminal model is generated, and subsequently, the registered and reported terminal models in the registered and reported terminal model library can be corrected according to the mapping relationship.
[0111] In addition, in the embodiments of the present disclosure, each time the user agent data is analyzed, a mapping relationship between the non-standard terminal model and the standard terminal model in the user agent data can be generated, and then the mapping relationship is used for correction. The generated mapping relationship can also be stored. After the non-standard terminal model in the user agent data is identified, the standard terminal model corresponding to the non-standard terminal model can be queried by using the stored mapping relationship, and subsequent correction can be directly performed by using the queried standard terminal model. Of course, in order to improve the accuracy of the registered and reported terminal model library, the stored mapping relationship can be updated regularly.
[0112] In an exemplary embodiment, correcting the terminal models in the registered reported terminal model library according to the mapping relationship between the non-standard terminal model and the standard terminal model may include: obtaining the user identifier of the non-standard terminal model; querying the registered reported terminal model of the user identifier from the registered reported terminal model library according to the user identifier; if the registered reported terminal model of the user identifier does not exist in the registered reported terminal model library, inserting the standard terminal model corresponding to the non-standard terminal model into the registered reported terminal model library; if the standard terminal model corresponding to the non-standard terminal model is inconsistent with the registered reported terminal model of the user identifier, replacing the registered reported terminal model of the user identifier with the standard terminal model corresponding to the non-standard terminal model.
[0113] After identifying the non-standard terminal model in the user agent data, the specific user identifier, such as the MDN user number, can be obtained, and then the registered reported terminal model corresponding to the user identifier can be queried in the registered reported terminal model library. Moreover, after identifying the non-standard terminal model, the standard terminal model corresponding to the non-standard terminal model can be determined according to the mapping relationship. If the registered reported terminal model corresponding to the user identifier does not exist in the registered reported terminal model library, then the standard terminal model corresponding to the non-standard terminal model can be inserted into the registered reported terminal model library. If the registered reported terminal model corresponding to the user identifier in the registered reported terminal model library is inconsistent with the standard terminal model corresponding to the non-standard terminal model, then the registered reported terminal model can be replaced with the standard terminal model.
[0114] In addition, after identifying the non-standard terminal model from the user agent data, it can be first determined whether the time when the non-standard terminal model appears exceeds a preset time threshold. If so, further processing can be performed on it; otherwise, no processing is performed first. The advantage of doing this is to avoid the small number of registered reported terminal models corresponding to the non-standard terminal model in the registered reported database.
[0115] Figure 7 is a flowchart of a method for correcting a terminal model library according to another exemplary embodiment. As Figure 7 shown, the method for correcting the terminal model library includes the following steps.
[0116] Step S701, obtaining the user agent data in the user behavior data;
[0117] Step S702, labeling the category label of the user agent data;
[0118] Step S703, obtaining the user agent data of each category according to the labeled category label of the user agent data;
[0119] Step S704: Determine the data processing order of the user agent data for each category according to the top-down order of the pre-constructed category tree;
[0120] Step S705: Identify and process the user agent data for each category according to the data processing order to obtain non-standard terminal models in the user agent data;
[0121] Step S706: Associate the non-standard terminal models with the registered and reported terminal models in the registered and reported terminal model library to obtain one or more registered and reported terminal models associated with the non-standard terminal models;
[0122] Step S707: Select a target terminal model from one or more registered and reported terminal models according to the number of users of one or more registered and reported terminal models;
[0123] Step S708: If the number of users of the target terminal model is greater than the preset user quantity threshold and the proportion of the number of users of the target terminal model to the number of users of the non-standard terminal model is greater than the preset proportion threshold, determine that the target terminal model is the standard terminal model corresponding to the non-standard terminal model, and generate a mapping relationship between the non-standard terminal model and the target terminal model;
[0124] Step S709: Obtain the user identifier of the non-standard terminal model;
[0125] Step S710: Query the registered and reported terminal model of the user identifier from the registered and reported terminal model library according to the user identifier;
[0126] Step S711: If the registered and reported terminal model of the user identifier does not exist in the registered and reported terminal model library, insert the standard terminal model corresponding to the non-standard terminal model into the registered and reported terminal model library;
[0127] Step S712: If the standard terminal model corresponding to the non-standard terminal model is inconsistent with the registered and reported terminal model of the user identifier, replace the registered and reported terminal model of the user identifier with the standard terminal model corresponding to the non-standard terminal model.
[0128] Among them, the category tree can be as described above Figure 4 or Figure 5The method shown is constructed and will not be elaborated here. In step S705, after identifying and processing the user agent data of each category, if it is determined that the non-standard terminal models of all users have been identified, there is no need to identify and process the user agent data of the remaining categories. After obtaining the non-standard terminal models in the user agent data through step S705, the category tree can be verified. How to verify has been described in detail above and will not be elaborated here. In step S708, if it is determined that the target terminal model does not meet the conditions, it cannot be considered that the target terminal model is the standard terminal model corresponding to the non-standard terminal model. Then, the registration and reporting terminal model library may not be corrected using this non-standard terminal model for the time being. Subsequently, when the standard terminal model corresponding to this non-standard terminal model exists in the registration and reporting terminal model library, the registration and reporting terminal model library can be corrected using this non-standard terminal model.
[0129] The method for correcting the terminal model library provided by the embodiments of the present disclosure can start from the perspective of the terminal model included in the user agent data, identify the non-standard terminal models in the user agent data through the identification algorithm of the user agent data, and then establish the mapping relationship between the non-standard terminal models and the standard terminal models with the help of the existing registration and reporting terminal model library. Furthermore, the registration and reporting terminal model library can be corrected according to the mapping relationship, improving the accuracy of the registration and reporting terminal model library.
[0130] Figure 8 is a block diagram of a device for correcting a terminal model library shown according to an exemplary embodiment. Refer to Figure 8 and the device may include: an identification module 810, a generation module 820, a correction module 830, a construction module 840, and a verification module 850.
[0131] The identification module 810 can be used to: obtain the user agent data in the user behavior data, and identify the non-standard terminal models in the user agent data based on the identification algorithm of the user agent data.
[0132] The generation module 820 can be used to: generate the mapping relationship between the non-standard terminal models and the standard terminal models according to the registration and reporting terminal model library.
[0133] The correction module 830 can be used to: correct the terminal models in the registration and reporting terminal model library according to the mapping relationship between the non-standard terminal models and the standard terminal models.
[0134] In an exemplary embodiment, the recognition module 810 may further be configured to: label the category tags of the user agent data; obtain the user agent data of each category according to the labeled category tags of the user agent data; determine the data processing order of the user agent data of each category according to the top-down order of the pre-constructed category tree; and perform recognition processing on the user agent data of each category according to the data processing order to obtain non-standard terminal models in the user agent data.
[0135] In an exemplary embodiment, the construction module 840 may be configured to pre-construct a category tree according to the following method: obtain sample data, and label the category tags of the sample data; obtain the sample data of each category according to the labeled category tags of the sample data; calculate the sample data of each category based on the modified empirical entropy and the modified conditional entropy, and sequentially extract the sample data of each category; and construct a category tree according to the extraction order of the sample data of each category.
[0136] In an exemplary embodiment, the construction module 840 may further be configured to: calculate the information gain of the sample data of each category with respect to the sample data based on the modified empirical entropy and the modified conditional entropy; determine the category corresponding to the maximum information gain as the target category, and extract the sample data of the target category from the sample data.
[0137] In an exemplary embodiment, the construction module 840 may further be configured to: obtain the remaining sample data; calculate the information gain of the sample data of each category in the remaining sample data with respect to the remaining sample data based on the modified empirical entropy and the modified conditional entropy; determine the category corresponding to the maximum information gain as the new target category, and extract the sample data of the new target category from the remaining sample data.
[0138] In an exemplary embodiment, the modified empirical entropy and the modified conditional entropy may include a custom modification coefficient; the custom modification coefficient is obtained by statistically analyzing the abnormal data in the sample data; and the abnormal data may include data with an identified terminal model as an outlier and data with a category as an outlier.
[0139] In an exemplary embodiment, the verification module 850 may be configured to: obtain the recognition result and verify the category tree according to the recognition result.
[0140] In an exemplary embodiment, the generating module 820 may further be configured to: associate a non-standard terminal model with the registered and reported terminal models in the registered and reported terminal model library to obtain one or more registered and reported terminal models associated with the non-standard terminal model; select a target terminal model from the one or more registered and reported terminal models according to the number of users of the one or more registered and reported terminal models; if the number of users of the target terminal model is greater than a preset user quantity threshold and the proportion of the number of users of the target terminal model to the number of users of the non-standard terminal model is greater than a preset proportion threshold, determine that the target terminal model is the standard terminal model corresponding to the non-standard terminal model, and generate a mapping relationship between the non-standard terminal model and the target terminal model.
[0141] In an exemplary embodiment, the correcting module 830 may further be configured to: obtain the user identifier of the non-standard terminal model; query the registered and reported terminal model of the user identifier from the registered and reported terminal model library according to the user identifier; if the registered and reported terminal model of the user identifier does not exist in the registered and reported terminal model library, insert the standard terminal model corresponding to the non-standard terminal model into the registered and reported terminal model library; if the standard terminal model corresponding to the non-standard terminal model is inconsistent with the registered and reported terminal model of the user identifier, replace the registered and reported terminal model of the user identifier with the standard terminal model corresponding to the non-standard terminal model.
[0142] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0143] Figure 9 It is a structural block diagram of an electronic device for correcting a terminal model library shown according to an exemplary embodiment. It should be noted that the illustrated electronic device is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.
[0144] Next, refer to Figure 9 to describe the electronic device 900 according to this embodiment of the present invention. Figure 9 The illustrated electronic device 900 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.
[0145] As Figure 9 shown, the electronic device 900 is presented in the form of a general-purpose computing device. The components of the electronic device 900 may include, but are not limited to: the at least one processing unit 910 described above, the at least one storage unit 920 described above, and a bus 930 connecting different system components (including the storage unit 920 and the processing unit 910).
[0146] Among them, the storage unit stores program codes, and the program codes can be executed by the processing unit 910, so that the processing unit 910 executes the steps according to various exemplary embodiments of the present invention described in the "Exemplary Method" section above of this specification. For example, the processing unit 910 can execute steps such as Figure 1 shown in, step S110, obtain user agent data in the user behavior data, and identify non-standard terminal models in the user agent data based on the identification algorithm of the user agent data; step S120, generate a mapping relationship between the non-standard terminal model and the standard terminal model according to the registered reported terminal model library; step S130, correct the terminal models in the registered reported terminal model library according to the mapping relationship between the non-standard terminal model and the standard terminal model.
[0147] The storage unit 920 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 9201 and / or a cache storage unit 9202, and may further include a read-only storage unit (ROM) 9203.
[0148] The storage unit 920 may further include a program / utility 9204 having a set (at least one) of program modules 9205. Such program modules 9205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.
[0149] The bus 930 may represent one or more of several types of bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any bus structure in a variety of bus structures.
[0150] The electronic device 900 can also communicate with one or more external devices 1000 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 900, and / or communicate with any device that enables the electronic device 900 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 950. Moreover, the electronic device 900 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 940. As shown in the figure, the network adapter 940 communicates with other modules of the electronic device 900 through the bus 930. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 900, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0151] In an exemplary embodiment of the present disclosure, there is also provided a computer-readable storage medium, on which a program product capable of implementing the above method in this specification is stored. In some possible implementation manners, various aspects of the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the above "Exemplary Method" section of this specification.
[0152] The program product for implementing the above method according to an embodiment of the present invention can be a portable compact disc read-only memory (CD-ROM) and includes program code, and can run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, the readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.
[0153] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0154] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable signal medium may also be any readable medium other than a readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.
[0155] The program code contained on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0156] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computing device, partially on the user's device, as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0157] It should be noted that although several modules or units of a device for performing actions are mentioned in the foregoing detailed description, such a division is not mandatory. In fact, according to embodiments of the present disclosure, the features and functions of two or more of the above-described modules or units may be embodied in one module or unit. Conversely, the features and functions of one module or unit described above may be further divided and embodied by multiple modules or units.
[0158] Furthermore, although the steps of the methods in the present disclosure are described in a specific order in the drawings, this does not require or imply that the steps must be performed in that specific order, or that all of the shown steps must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.
[0159] Those skilled in the art can easily understand from the description of the above embodiments that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (such as a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0160] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only illustrative, and the true scope and spirit of the present disclosure are pointed out by the appended claims.
[0161] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A method for correcting a terminal model library, characterized in that, it includes: Obtain user agent data in user behavior data, and based on an identification algorithm for user agent data, identify non-standard terminal models in the user agent data; Generate a mapping relationship between the non-standard terminal models and standard terminal models according to the registered and reported terminal model library; Correct the terminal models in the registered and reported terminal model library according to the mapping relationship between the non-standard terminal models and standard terminal models; Among them, generating the mapping relationship between the non-standard terminal models and standard terminal models according to the registered and reported terminal model library includes: associating the non-standard terminal models with the registered and reported terminal models in the registered and reported terminal model library to obtain one or more registered and reported terminal models associated with the non-standard terminal models; according to the number of users of the one or more registered and reported terminal models, select a target terminal model from the one or more registered and reported terminal models; if the number of users of the target terminal model is greater than a preset user quantity threshold and the proportion of the number of users of the target terminal model in the number of users of the non-standard terminal model is greater than a preset proportion threshold, then determine that the target terminal model is the standard terminal model corresponding to the non-standard terminal model, and generate a mapping relationship between the non-standard terminal model and the target terminal model.
2. The method according to claim 1, characterized in that, Based on the identification algorithm for user agent data, identifying non-standard terminal models in the user agent data includes: Labeling the category tags of the user agent data; Obtaining user agent data of each category according to the labeled category tags of the user agent data; Determining the data processing order of the user agent data of each category according to the top-down order of the pre-constructed category tree; Performing identification processing on the user agent data of each category according to the data processing order to obtain non-standard terminal models in the user agent data.
3. The method according to claim 2, characterized in that, The category tree is pre-constructed according to the following method: Obtain sample data and label the category tags of the sample data; Obtain sample data of each category according to the labeled category tags of the sample data; Perform calculations on the sample data of each category based on the modified empirical entropy and modified conditional entropy, and sequentially extract the sample data of each category; Construct the category tree according to the extraction order of the sample data of each category.
4. The method according to claim 3, characterized in that, Performing calculations on the sample data of each category based on the modified empirical entropy and modified conditional entropy, and sequentially extracting the sample data of each category includes: Calculating the information gain of the sample data of each category for the sample data based on the modified empirical entropy and modified conditional entropy; Determine the category corresponding to the maximum information gain as the target category, and extract the sample data of the target category from the sample data.
5. The method according to claim 4, characterized in that, After extracting the sample data of the target category from the sample data, the method further includes: Obtaining the remaining sample data; Calculating the information gain of the sample data of each category in the remaining sample data with respect to the remaining sample data based on the modified empirical entropy and the modified conditional entropy; Determining the category corresponding to the maximum information gain as the new target category, and extracting the sample data of the new target category from the remaining sample data.
6. The method according to any one of claims 3 to 5, wherein, the modified empirical entropy and the modified conditional entropy include a user-defined modification coefficient; the user-defined modification coefficient is obtained by statistically analyzing the abnormal data in the sample data; and, the abnormal data includes data with an identified terminal model as an outlier and data with a category as an outlier.
7. The method according to claim 2, wherein, after identifying the non-standard terminal model in the user agent data based on the identification algorithm of the user agent data, the method further includes: Obtaining the identification result and verifying the category tree according to the identification result.
8. The method according to claim 1, wherein, correcting the terminal models in the registered reported terminal model library according to the mapping relationship between the non-standard terminal model and the standard terminal model, including: Obtaining the user identifier of the non-standard terminal model; Querying the registered reported terminal model of the user identifier from the registered reported terminal model library according to the user identifier; If the registered reported terminal model of the user identifier does not exist in the registered reported terminal model library, inserting the standard terminal model corresponding to the non-standard terminal model into the registered reported terminal model library; If the standard terminal model corresponding to the non-standard terminal model is inconsistent with the registered reported terminal model of the user identifier, replacing the registered reported terminal model of the user identifier with the standard terminal model corresponding to the non-standard terminal model.
9. An apparatus for correcting a terminal model library, wherein, comprising: An identification module, configured to obtain user agent data in user behavior data and identify non-standard terminal models in the user agent data based on an identification algorithm of the user agent data; A generation module, configured to generate a mapping relationship between the non-standard terminal model and the standard terminal model according to the registered reported terminal model library; A correction module, configured to correct the terminal models in the registered reported terminal model library according to the mapping relationship between the non-standard terminal model and the standard terminal model; The generating module is further configured to associate the non-standard terminal model with the registered and reported terminal models in the registered and reported terminal model library to obtain one or more registered and reported terminal models associated with the non-standard terminal model; select a target terminal model from the one or more registered and reported terminal models according to the number of users of the one or more registered and reported terminal models; if the number of users of the target terminal model is greater than a preset user quantity threshold and the proportion of the number of users of the target terminal model to the number of users of the non-standard terminal model is greater than a preset proportion threshold, determine that the target terminal model is the standard terminal model corresponding to the non-standard terminal model, and generate a mapping relationship between the non-standard terminal model and the target terminal model.
10. An electronic device, characterized in that, it includes: a processor; a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the instructions to implement the instruction processing method according to any one of claims 1 to 8.
11. A computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enable the electronic device to execute the instruction processing method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Method and apparatus for identifying terminal version number of target terminal
CN109117172A
BOM model matching device and method, electronic equipment and storage medium
CN111061770A